Pre-automation baseline

Build the Baseline Before You Automate Anything

A baseline-building workflow for process boundaries, volume, time, quality, exceptions, cost, and evidence before an automation pilot changes the operation.

By WhichAI. Published 2026-07-12. Updated 2026-07-12.

Methodology: Editorial synthesis of workflow design patterns and implementation constraints. Public control references provide context, not proof of a deployment or legal advice. Where a versioned evidence pack appears, its evidence class, method, and limitations govern what the artifact can support. Read the full method. Report a correction.

Built for

COOs, process owners, finance partners, and teams preparing an AI or automation initiative

The decision

Decide what must be measured before a workflow change so later claims can be tested against local evidence.

Answer first

Freeze a representative picture of the current process before changing tools, staffing, routing, or policy. Measure work units, stage time, quality, exceptions, demand, and cost with source records that can survive later scrutiny.

WhichAI Solutions diagnostic

Bring this operating problem to the diagnostic

Use Solutions when the workflow crosses teams, the baseline is disputed, data is fragmented, or the automation decision is tied to meaningful headcount, service, or operating spend.

Open the diagnostic

The capacity leak

What the team is doing before anyone calls it a systems problem

Headcount pressure rarely starts with one giant task. It starts when ordinary work is split across inboxes, tabs, handoffs, and undocumented judgment calls. These are the signals to map first.

SIGNAL 01

The current process is described with anecdotes instead of a defined start, finish, owner, and work unit.

SIGNAL 02

Volume counts completed outputs but not rejected inputs, rework, abandoned cases, or queue age.

SIGNAL 03

Quality and exceptions are discussed informally without a stable taxonomy or sampling method.

SIGNAL 04

The automation pilot begins before current labor, tool, delay, and failure costs are captured.

The implementation

The system should prepare the decision, not pretend the decision disappeared

A complete implementation connects the intake, context, transformation, review, and record. The output of one stage becomes the controlled input to the next. A human owns the exceptions and the final consequence.

StageCurrent dragSystem responsibilityHuman responsibilityEvidence kept
1. Process boundaryDifferent teams describe different versions of the same workflow.Define trigger, valid input, stages, owners, completion state, and excluded work.Approve the boundary and identify consequences that require separate measurement.Process map, unit definition, owners, exclusions, and approval.
2. Demand and queueWorkload is represented by completed cases alone.Capture arrivals, completions, rejections, backlog, age, seasonality, and case mix.Explain exceptional periods and choose the representative window.Source counts, queue snapshots, period, case tags, and notes.
3. Time and costLabor estimates come from role assumptions rather than observed activity.Sample stage effort, wait time, handoffs, tools, vendor cost, and rework.Validate samples and distinguish active labor from elapsed delay.Work IDs, timestamps, effort logs, rates, costs, and reviewer.
4. Quality and exceptionsErrors and nonstandard cases are hidden inside general rework.Define quality criteria and an exception taxonomy, then sample outcomes consistently.Approve severity, consequence, and the cases requiring specialist judgment.Rubric, sample, outcome, exception code, severity, and disposition.
5. Baseline freezeMetrics change during the pilot and overwrite the original comparison point.Publish an immutable baseline snapshot with sources, assumptions, gaps, and confidence limits.Sign off on the baseline and decide what the pilot must prove.Snapshot, source links, assumptions, approvers, and pilot criteria.

What the human keeps

The goal is not zero humans. It is zero avoidable preparation around the judgment only a responsible owner should make.

  • Process owners define the real workflow boundary, completion state, quality bar, and exception consequences.
  • Operators validate observed work and identify hidden rework, queue, and handoff activity.
  • Finance and leadership approve cost assumptions and the evidence required for a later investment decision.

Controls before volume

A workflow is not ready because the happy path worked once. It is ready when access, review, fallback, and evidence are explicit.

  • Freeze source extracts and definitions before the pilot changes behavior or system records.
  • Include failed, rejected, reworked, and abandoned units in the workload picture.
  • Record sample selection and case mix so the after period can be compared honestly.
  • Label missing data and assumptions rather than replacing them with industry averages presented as local fact.

The scorecard

Measure capacity, not activity

A system can produce more messages and still make the operation worse. Measure movement through the workflow, the quality of review, and the load that still reaches a person.

Demand profile

Arrivals, completions, backlog, age, rejections, and case mix during the approved baseline period.

Total effort per unit

Human preparation, execution, review, correction, exception, and handoff effort per representative unit.

Quality and rework

Share of completed units meeting the approved quality bar without material correction or repeated processing.

Exception burden

Rate, severity, owner time, and cycle impact of defined nonstandard cases.

What a fake implementation looks like here

These patterns create an AI demo while leaving the labor, risk, and accountability in the same place.

  • Starting the pilot and then reconstructing the old process from memory.
  • Using completed volume while excluding rejected, reworked, and abandoned work.
  • Applying generic salary or productivity benchmarks as if they were measured local costs.
  • Changing definitions between baseline and pilot so improvement cannot be interpreted.

Two ways to act

Use the path that matches the decision

Questions

What operators ask before they build

How long should a baseline period be?

Choose a period that captures representative demand and case mix. Document seasonality, unusual events, and sample limits rather than relying on a universal duration.

What if the current process has poor data?

Record what can be observed, run a bounded manual sample, identify missing fields, and treat uncertainty as part of the pilot decision instead of inventing precision.

Should we use industry benchmarks?

Benchmarks can provide context, but they should not replace local workflow volume, effort, quality, exception, cost, and queue evidence.

Primary references

Controls should come from the specific operating environment

These are broad public control references, not article-specific evidence, vendor endorsements, or legal advice. Validate the current rules, contracts, system configuration, and organization-specific risk before deployment.

Keep mapping

Related implementation guides

Use this evidence with