Test the ordinary operation
Five Tests That Separate Real AI Implementation From an AI Demo
A five-test readiness rubric for evidence, variation, human judgment, failure recovery, and measurable downstream acceptance.
By WhichAI. Published 2026-07-12. Updated 2026-07-12.
Methodology: Editorial synthesis of workflow design patterns and implementation constraints. Public control references provide context, not proof of a deployment or legal advice. Where a versioned evidence pack appears, its evidence class, method, and limitations govern what the artifact can support. Read the full method. Report a correction.
Built for
Buyers and sponsors evaluating vendors, pilots, and internal demonstrations
The decision
Decide whether a demonstration has enough operating evidence to enter a controlled pilot.
Answer first
A demonstration proves a prepared example can produce output. An implementation must pass source, variation, judgment, failure, and downstream acceptance tests using representative local cases.
WhichAI Solutions diagnostic
Bring this operating problem to the diagnostic
Use WhichAI Solutions when a vendor or internal team is asking to move from demonstration to operational use and the evaluation must cover several systems, risk owners, or staffing implications.
The capacity leak
What the team is doing before anyone calls it a systems problem
Headcount pressure rarely starts with one giant task. It starts when ordinary work is split across inboxes, tabs, handoffs, and undocumented judgment calls. These are the signals to map first.
The demonstration uses selected clean inputs and excludes missing or conflicting evidence.
The presenter resolves failures off-screen or restarts without recording recovery work.
Reviewers judge visual quality but do not test source traceability or downstream acceptance.
Leadership is asked to scale before access, ownership, incident, and rollback responsibilities exist.
The implementation
The system should prepare the decision, not pretend the decision disappeared
A complete implementation connects the intake, context, transformation, review, and record. The output of one stage becomes the controlled input to the next. A human owns the exceptions and the final consequence.
| Stage | Current drag | System responsibility | Human responsibility | Evidence kept |
|---|---|---|---|---|
| 1. Source test | Output looks plausible but its factual basis is unclear. | Require material fields and claims to link to the exact input source and transformation version. | A domain reviewer checks whether the evidence supports the intended use. | Source map, transformation version, unsupported fields, and reviewer disposition. |
| 2. Variation test | Only ideal cases appear in the demo. | Run representative ordinary, incomplete, conflicting, unusual, and out-of-scope cases from a frozen test set. | Operators validate that the set reflects real production variation. | Test-set manifest, case labels, results, corrections, and exclusions. |
| 3. Judgment test | The system quietly makes or implies consequential decisions. | Separate preparation from judgment and show the reviewer the evidence, uncertainty, rule, and permitted actions. | The accountable human approves, rejects, corrects, or escalates. | Decision boundary, reviewer actions, rationale, and overrides. |
| 4. Failure test | Vendor, permission, duplicate, and downstream errors are not demonstrated. | Trigger known failures and verify detection, containment, ownership, duplicate safety, recovery, and rollback. | Incident owners execute the runbook and confirm restoration. | Failure injections, alerts, assignments, recovery times, and rollback proof. |
| 5. Acceptance test | Success ends when output is generated. | Measure whether the next process accepts the result without material repair and whether total human work changes. | The downstream owner defines acceptance and signs the pilot decision. | Acceptance result, correction, cycle time, human touch, and decision memo. |
What the human keeps
The goal is not zero humans. It is zero avoidable preparation around the judgment only a responsible owner should make.
- Operators select a representative frozen test set and expose ordinary variation.
- Domain reviewers verify evidence, retain consequential judgment, and record corrections.
- System and business owners execute failure recovery and approve the pilot decision.
Controls before volume
A workflow is not ready because the happy path worked once. It is ready when access, review, fallback, and evidence are explicit.
- Freeze test cases before evaluation and include incomplete, conflicting, unusual, and failed examples.
- Require source-level evidence for material output and a named reviewer for consequential use.
- Test duplicate safety, vendor failure, permission failure, downstream rejection, and rollback.
- End the test at accepted downstream action, not generated output.
The scorecard
Measure capacity, not activity
A system can produce more messages and still make the operation worse. Measure movement through the workflow, the quality of review, and the load that still reaches a person.
Source support rate
Share of material output fields linked to sufficient source evidence.
Variation pass rate
Accepted cases by ordinary, incomplete, conflicting, unusual, and out-of-scope category.
Failure recovery
Detection and recovery time for each injected critical failure.
Downstream acceptance
Share of cases accepted by the next process without material repair.
What a fake implementation looks like here
These patterns create an AI demo while leaving the labor, risk, and accountability in the same place.
- Letting the vendor choose only successful examples after seeing the evaluation criteria.
- Scoring fluency or appearance without verifying sources and downstream use.
- Treating a manual restart as an adequate recovery design.
- Using a passed demonstration as proof of scaled reliability or outcomes for customers.
Two ways to act
Use the path that matches the decision
WhichAI Solutions
The workflow is becoming a company problem.
Use WhichAI Solutions when a vendor or internal team is asking to move from demonstration to operational use and the evaluation must cover several systems, risk owners, or staffing implications.
Bring one bottleneck. We map the work under it, separate consequential judgment from mechanical drag, and decide whether the next move is a hire, a tool, or a rebuild.
See company solutionsTask-specific workflow brief
Plan this recurring task.
Start with this task draft, then complete the three-question brief:
Create a five-test evaluation for one AI workflow: source support, representative variation, human judgment boundaries, failure recovery, and downstream acceptance. Define frozen cases, evidence, reviewer actions, injected failures, metrics, and a bounded pilot decision.
Choose a paid plan after reviewing your brief. WhichAI creates a plan and does not set up tools or accounts.
Start the briefQuestions
What operators ask before they build
What are the five tests?
Source support, real-case variation, human judgment boundaries, failure recovery, and downstream acceptance with total work measured.
Does passing mean the workflow is ready to scale?
No. It supports a bounded pilot decision. Scale requires evidence across production volume, variation, reviewers, incidents, changes, and recovery.
Who should choose the test set?
The buyer's operators and domain reviewers should freeze a representative set before evaluation. Include cases the current process finds difficult.
Primary references
Controls should come from the specific operating environment
These are broad public control references, not article-specific evidence, vendor endorsements, or legal advice. Validate the current rules, contracts, system configuration, and organization-specific risk before deployment.
National Institute of Standards and Technology
AI Risk Management Framework
A voluntary framework for mapping, measuring, managing, and governing AI risk.
Accessed 2026-07-14
Federal Trade Commission
Operation AI Comply
Enforcement examples showing why AI performance and substitution claims need evidence.
Accessed 2026-07-14
Apply a public research tool
Use the artifact before the next operating decision.
Keep mapping
Related implementation guides
More in Real implementation
From Tool Stack to Operating Capacity: The Missing Implementation Layer
A stack-to-owner-to-metric map for turning selected products into one controlled workflow with source evidence, review, and recovery.
Explore more Real implementation guidesMore in Real implementation
The Exception Path Is the Real Workflow
An exception operating model for detecting missing, conflicting, unusual, or failed cases and routing them with evidence to a named owner.
Explore more Real implementation guides