Pre-automation baseline
Build the Baseline Before You Automate Anything
A baseline-building workflow for process boundaries, volume, time, quality, exceptions, cost, and evidence before an automation pilot changes the operation.
By WhichAI. Published 2026-07-12. Updated 2026-07-12.
Methodology: Editorial synthesis of workflow design patterns and implementation constraints. Public control references provide context, not proof of a deployment or legal advice. Where a versioned evidence pack appears, its evidence class, method, and limitations govern what the artifact can support. Read the full method. Report a correction.
Built for
COOs, process owners, finance partners, and teams preparing an AI or automation initiative
The decision
Decide what must be measured before a workflow change so later claims can be tested against local evidence.
Answer first
Freeze a representative picture of the current process before changing tools, staffing, routing, or policy. Measure work units, stage time, quality, exceptions, demand, and cost with source records that can survive later scrutiny.
WhichAI Solutions diagnostic
Bring this operating problem to the diagnostic
Use Solutions when the workflow crosses teams, the baseline is disputed, data is fragmented, or the automation decision is tied to meaningful headcount, service, or operating spend.
The capacity leak
What the team is doing before anyone calls it a systems problem
Headcount pressure rarely starts with one giant task. It starts when ordinary work is split across inboxes, tabs, handoffs, and undocumented judgment calls. These are the signals to map first.
The current process is described with anecdotes instead of a defined start, finish, owner, and work unit.
Volume counts completed outputs but not rejected inputs, rework, abandoned cases, or queue age.
Quality and exceptions are discussed informally without a stable taxonomy or sampling method.
The automation pilot begins before current labor, tool, delay, and failure costs are captured.
The implementation
The system should prepare the decision, not pretend the decision disappeared
A complete implementation connects the intake, context, transformation, review, and record. The output of one stage becomes the controlled input to the next. A human owns the exceptions and the final consequence.
| Stage | Current drag | System responsibility | Human responsibility | Evidence kept |
|---|---|---|---|---|
| 1. Process boundary | Different teams describe different versions of the same workflow. | Define trigger, valid input, stages, owners, completion state, and excluded work. | Approve the boundary and identify consequences that require separate measurement. | Process map, unit definition, owners, exclusions, and approval. |
| 2. Demand and queue | Workload is represented by completed cases alone. | Capture arrivals, completions, rejections, backlog, age, seasonality, and case mix. | Explain exceptional periods and choose the representative window. | Source counts, queue snapshots, period, case tags, and notes. |
| 3. Time and cost | Labor estimates come from role assumptions rather than observed activity. | Sample stage effort, wait time, handoffs, tools, vendor cost, and rework. | Validate samples and distinguish active labor from elapsed delay. | Work IDs, timestamps, effort logs, rates, costs, and reviewer. |
| 4. Quality and exceptions | Errors and nonstandard cases are hidden inside general rework. | Define quality criteria and an exception taxonomy, then sample outcomes consistently. | Approve severity, consequence, and the cases requiring specialist judgment. | Rubric, sample, outcome, exception code, severity, and disposition. |
| 5. Baseline freeze | Metrics change during the pilot and overwrite the original comparison point. | Publish an immutable baseline snapshot with sources, assumptions, gaps, and confidence limits. | Sign off on the baseline and decide what the pilot must prove. | Snapshot, source links, assumptions, approvers, and pilot criteria. |
What the human keeps
The goal is not zero humans. It is zero avoidable preparation around the judgment only a responsible owner should make.
- Process owners define the real workflow boundary, completion state, quality bar, and exception consequences.
- Operators validate observed work and identify hidden rework, queue, and handoff activity.
- Finance and leadership approve cost assumptions and the evidence required for a later investment decision.
Controls before volume
A workflow is not ready because the happy path worked once. It is ready when access, review, fallback, and evidence are explicit.
- Freeze source extracts and definitions before the pilot changes behavior or system records.
- Include failed, rejected, reworked, and abandoned units in the workload picture.
- Record sample selection and case mix so the after period can be compared honestly.
- Label missing data and assumptions rather than replacing them with industry averages presented as local fact.
The scorecard
Measure capacity, not activity
A system can produce more messages and still make the operation worse. Measure movement through the workflow, the quality of review, and the load that still reaches a person.
Demand profile
Arrivals, completions, backlog, age, rejections, and case mix during the approved baseline period.
Total effort per unit
Human preparation, execution, review, correction, exception, and handoff effort per representative unit.
Quality and rework
Share of completed units meeting the approved quality bar without material correction or repeated processing.
Exception burden
Rate, severity, owner time, and cycle impact of defined nonstandard cases.
What a fake implementation looks like here
These patterns create an AI demo while leaving the labor, risk, and accountability in the same place.
- Starting the pilot and then reconstructing the old process from memory.
- Using completed volume while excluding rejected, reworked, and abandoned work.
- Applying generic salary or productivity benchmarks as if they were measured local costs.
- Changing definitions between baseline and pilot so improvement cannot be interpreted.
Two ways to act
Use the path that matches the decision
WhichAI Solutions
The workflow is becoming a company problem.
Use Solutions when the workflow crosses teams, the baseline is disputed, data is fragmented, or the automation decision is tied to meaningful headcount, service, or operating spend.
Bring one bottleneck. We map the work under it, separate consequential judgment from mechanical drag, and decide whether the next move is a hire, a tool, or a rebuild.
See company solutionsTask-specific workflow brief
Plan this recurring task.
Start with this task draft, then complete the three-question brief:
Build a pre-automation baseline for one workflow with an explicit process boundary, demand and queue snapshot, stage time and cost sample, quality rubric, exception taxonomy, source records, assumptions, and immutable signoff.
Choose a paid plan after reviewing your brief. WhichAI creates a plan and does not set up tools or accounts.
Start the briefQuestions
What operators ask before they build
How long should a baseline period be?
Choose a period that captures representative demand and case mix. Document seasonality, unusual events, and sample limits rather than relying on a universal duration.
What if the current process has poor data?
Record what can be observed, run a bounded manual sample, identify missing fields, and treat uncertainty as part of the pilot decision instead of inventing precision.
Should we use industry benchmarks?
Benchmarks can provide context, but they should not replace local workflow volume, effort, quality, exception, cost, and queue evidence.
Primary references
Controls should come from the specific operating environment
These are broad public control references, not article-specific evidence, vendor endorsements, or legal advice. Validate the current rules, contracts, system configuration, and organization-specific risk before deployment.
National Institute of Standards and Technology
AI Risk Management Framework
A voluntary framework for mapping, measuring, managing, and governing AI risk.
Accessed 2026-07-14
Federal Trade Commission
Operation AI Comply
Enforcement examples showing why AI performance and substitution claims need evidence.
Accessed 2026-07-14
Keep mapping
Related implementation guides
More in Measurement and proof
Exception Rate Is the AI Metric Most Teams Ignore
A workflow measurement design for defining, capturing, aging, resolving, and learning from the cases an AI system cannot handle cleanly.
Explore more Measurement and proof guidesMore in Measurement and proof
How to Measure Time Saved by an AI Workflow
A practical measurement design for comparing human work before and after an AI workflow without ignoring review, corrections, exceptions, or displaced labor.
Explore more Measurement and proof guidesUse this evidence with