Design the failure path first
The Exception Path Is the Real Workflow
An exception operating model for detecting missing, conflicting, unusual, or failed cases and routing them with evidence to a named owner.
By WhichAI. Published 2026-07-12. Updated 2026-07-12.
Methodology: Editorial synthesis of workflow design patterns and implementation constraints. Public control references provide context, not proof of a deployment or legal advice. Where a versioned evidence pack appears, its evidence class, method, and limitations govern what the artifact can support. Read the full method. Report a correction.
Built for
Operations teams whose clean pilot cases hide ordinary edge work
The decision
Decide how exceptions are detected, classified, enriched, assigned, recovered, and used to improve the system.
Answer first
The happy path proves that a component can work. The exception path determines whether the operation can depend on it. Every failed or ambiguous case needs a reason, evidence, owner, deadline, and recovery action.
WhichAI Solutions diagnostic
Bring this operating problem to the diagnostic
Use WhichAI Solutions when exceptions already consume specialists or managers, failures cross system boundaries, or a missed recovery could affect customers or material operations.
The capacity leak
What the team is doing before anyone calls it a systems problem
Headcount pressure rarely starts with one giant task. It starts when ordinary work is split across inboxes, tabs, handoffs, and undocumented judgment calls. These are the signals to map first.
Exceptions return to a shared inbox with no structured reason or priority.
Staff reconstruct source context and prior attempts before they can resolve a failed case.
Retries create duplicate records or downstream actions because case identity is unstable.
Recurring failure patterns remain hidden in notes and never change preparation or routing rules.
The implementation
The system should prepare the decision, not pretend the decision disappeared
A complete implementation connects the intake, context, transformation, review, and record. The output of one stage becomes the controlled input to the next. A human owns the exceptions and the final consequence.
| Stage | Current drag | System responsibility | Human responsibility | Evidence kept |
|---|---|---|---|---|
| 1. Exception taxonomy | Anything outside the happy path is labeled failed or manual. | Define reason codes for missing input, conflicting evidence, unsupported format, policy ambiguity, integration failure, duplicate, and consequential review. | Exception owners approve categories, severity, and required response. | Taxonomy version, examples, severity, owner, and service target. |
| 2. Detection boundary | Failures are noticed only when a person spots a bad result. | Add explicit completeness, conflict, format, permission, duplicate, and downstream-write checks at each stage. | Humans validate detection on representative edge cases and identify false alarms. | Check result, rule version, test cases, false positives, and missed cases. |
| 3. Exception packet | The resolver receives an error message without business context. | Bundle source records, attempted transformations, check results, prior retries, affected action, and suggested next questions. | The resolver determines the correct action rather than accepting an inferred fix. | Case identifier, source links, attempts, error details, and reviewer decision. |
| 4. Owned recovery | Exceptions wait in shared queues or are retried informally. | Assign by reason and severity, set a response target, prevent duplicate action, and record resolve, return, reject, or escalate. | Named owners approve recovery and consequential downstream action. | Assignment, age, action, rationale, downstream reference, and closure time. |
| 5. Failure learning | Resolved exceptions disappear from view. | Review volume and recovery time by cause, then change inputs, rules, training, or scope under version control. | Owners approve changes and decide when a case type should leave the automated path. | Trend report, root cause, approved change, regression result, and rollback plan. |
What the human keeps
The goal is not zero humans. It is zero avoidable preparation around the judgment only a responsible owner should make.
- Business owners define exception severity, response targets, and consequential escalation.
- Resolvers interpret ambiguous cases, communicate, and authorize recovery actions.
- Workflow owners analyze recurring causes and approve controlled changes or scope reduction.
Controls before volume
A workflow is not ready because the happy path worked once. It is ready when access, review, fallback, and evidence are explicit.
- Give every exception a stable case identifier, reason code, severity, owner, and age.
- Prevent retries from creating duplicate downstream records or actions.
- Preserve source evidence and every attempted transformation or recovery.
- Stop the affected workflow path when a critical control or recovery threshold is breached.
The scorecard
Measure capacity, not activity
A system can produce more messages and still make the operation worse. Measure movement through the workflow, the quality of review, and the load that still reaches a person.
Exception rate by cause
Exceptions per one hundred accepted cases, separated by reason and severity.
Recovery cycle time
Median detection-to-closure time by exception type.
Repeat failure rate
Share of resolved cases that fail again for the same cause or re-enter the queue.
Human recovery load
Resolver minutes per one hundred cases, including context reconstruction and communication.
What a fake implementation looks like here
These patterns create an AI demo while leaving the labor, risk, and accountability in the same place.
- Using one generic manual queue for every failure type.
- Retrying failed writes without duplicate protection.
- Hiding missing or conflicting input behind a plausible default.
- Optimizing happy-path speed while exception volume consumes more human time.
Two ways to act
Use the path that matches the decision
WhichAI Solutions
The workflow is becoming a company problem.
Use WhichAI Solutions when exceptions already consume specialists or managers, failures cross system boundaries, or a missed recovery could affect customers or material operations.
Bring one bottleneck. We map the work under it, separate consequential judgment from mechanical drag, and decide whether the next move is a hire, a tool, or a rebuild.
See company solutionsTask-specific workflow brief
Plan this recurring task.
Start with this task draft, then complete the three-question brief:
Design exception handling for one recurring AI-assisted workflow. Create a reason taxonomy, detection checks, source-linked exception packet, severity and ownership rules, duplicate-safe recovery, service targets, learning loop, stop thresholds, and four measures.
Choose a paid plan after reviewing your brief. WhichAI creates a plan and does not set up tools or accounts.
Start the briefQuestions
What operators ask before they build
What is an exception in an AI workflow?
Any case that cannot safely follow the ordinary preparation and review path because evidence is missing, conflicting, unsupported, consequential, duplicated, or affected by a technical failure.
Should the system suggest a recovery?
It may present approved options and evidence, but the responsible person should choose when the case is ambiguous or consequential. Preserve the selected action and rationale.
When should a case type leave the automated path?
When exception frequency, severity, recovery time, or missed issues exceed the approved threshold and a redesign cannot restore dependable control.
Primary references
Controls should come from the specific operating environment
These are broad public control references, not article-specific evidence, vendor endorsements, or legal advice. Validate the current rules, contracts, system configuration, and organization-specific risk before deployment.
National Institute of Standards and Technology
AI Risk Management Framework
A voluntary framework for mapping, measuring, managing, and governing AI risk.
Accessed 2026-07-14
Cybersecurity and Infrastructure Security Agency
Secure by Design
Security principles for making systems safer by default and reducing avoidable customer burden.
Accessed 2026-07-14
Keep mapping
Related implementation guides
More in Real implementation
Human Review Is Not a Failure of Automation
A risk-tiered review design for directing human attention to consequential, ambiguous, low-confidence, and exception cases without recreating the original workload.
Explore more Real implementation guidesMore in Real implementation
Five Tests That Separate Real AI Implementation From an AI Demo
A five-test readiness rubric for evidence, variation, human judgment, failure recovery, and measurable downstream acceptance.
Explore more Real implementation guidesUse this evidence with
Continue into an inspectable flagship
Record exception work in the removal ledger
The happy path cannot establish capacity when exceptions dominate the operating load.
Open the evidenceInspect a freight-specific exception test pack
The synthetic cases make normal, missing, mismatch, unreadable, duplicate, and field-exception paths inspectable.
Open the evidence