AI Workflow Evidence Protocol
Create a comparable, auditable record of one AI-assisted workflow experiment.
Practical AI, proven in use
Use versioned methods to define the baseline, compare approaches, record failures and human decisions, and measure what happened after use. Private work stays local; public claims begin only after verification.
One defined endeavor
Develop, validate, and disseminate privacy-preserving methods that help U.S. organizations, educators, workforce programs, professionals, and technical teams test generative-AI workflows, measure real outcomes, and implement appropriate human oversight.
Choose the work
The public evidence loop
The platform preserves the chain from contribution and version through independent adoption, measured outcome, outside review, and external citation.
Working specifications
Create a comparable, auditable record of one AI-assisted workflow experiment.
Share enough evidence to review a workflow claim without publishing private project inputs.
Decide whether a set of outcome reports is comparable and private enough to publish as a benchmark.
Editorial reference protocols
| Protocol | Track | Version | Review scope | Action |
|---|---|---|---|---|
| Run a controlled support-triage pilotTest whether AI can reduce triage time while preserving routing accuracy and human approval. | business | 1.0 | Protocol structure, privacy boundary, and measurement method | Open → |
| Build portfolio proof in a weekly evidence cycleCreate verifiable proof of skill without inventing job outcomes or exposing employer information. | careers | 1.0 | Evidence rubric, privacy boundary, and claim language | Open → |
| Design an assessment with declared AI rolesUse AI in an assessment without hiding its role or weakening the evidence of learning. | teaching | 1.0 | Protocol structure, learner privacy, accessibility, and measurement method | Open → |
| Gate a RAG change with a fixed evaluation setDecide whether a RAG change is safe to ship using reproducible evidence instead of a demo impression. | builders | 1.0 | Evaluation structure, version traceability, and rollback controls | Open → |
Benchmarks
No outcome benchmark is public today. Each cohort needs at least 10 reviewed reports from at least three independent organizations, with no majority contributor.
September field tests