Evaluate a new AI workflow as a replacement for one narrow old workflow, not as a general productivity claim. Start with a task that recurs often enough to compare, record how the old method performs, and test the alternative on comparable work. The relevant result is the end-to-end outcome from receiving the work to an accepted result, including prompting, checking, revisions, handoffs, and later correction, not how quickly a model produces its first draft.
Measure time, direct cost, quality, errors, rework, review burden, consistency, cognitive load, privacy and security exposure, and a downstream outcome that matters to the work. Do not trade an unacceptable security exposure, legal risk, or quality failure for a lower average time. Research has found useful AI gains in one real customer-support setting and a slowdown in a controlled study of experienced developers working on familiar mature codebases, which is exactly why task-specific measurement matters. NBER customer-support study and METR developer study
Run a short crossover experiment with matched work samples, make the keep or stop rule before you see the results, and review every eligible task rather than only memorable wins. Keep the workflow when it improves representative outcomes and clears the risk gates. Modify it when the value is real but the protocol is weak, restrict it to low-risk stages when human verification remains expensive, and stop it when the gain disappears after rework, review, coordination, or risk is counted.
Define better before choosing a tool
“Better” is not a model leaderboard position or a faster first response. It is an improvement in a specific piece of work under the constraints that matter to the person or team doing it. A research assistant who produces a plausible answer in two minutes but requires twenty minutes of source checking may be slower overall. A drafting tool that reduces blank-page anxiety may be worth keeping even if its clock time is neutral. A coding assistant that saves implementation time but increases production defects is worse for the organization that must repair the defects.
Write a one-sentence success statement before running a trial. For example: “For standard customer briefings, reduce median time from assigned request to editor-approved briefing without lowering source accuracy, exposing confidential material, or increasing editor time.” This sentence prevents a trial from quietly becoming a contest to find any favorable metric.
Then choose one narrow workflow. Good starting candidates have a repeatable input, an observable finished state, and enough volume to collect a handful of comparable examples. Examples include producing a first draft from approved notes, turning meeting recordings into action items, triaging routine support tickets, making a research brief from supplied sources, or writing a test plan from a stable specification. Do not begin with a high-stakes decision, a complex multi-month project, or every type of work at once.
The evidence does not support one universal productivity number. In a large customer-support deployment, access to a generative assistant increased issues resolved per hour on average, with a much larger effect for novice and lower-skilled workers and minimal effect for experienced and highly skilled workers. NBER study By contrast, METR's randomized study of 16 experienced open-source developers completing 246 tasks in repositories they knew well found that allowing early-2025 AI tools increased task completion time, despite participants expecting a speedup. METR study Those results are not contradictory. They show that task context, user experience, tool fit, and quality requirements can change the outcome.
Capture the old workflow baseline
Before changing anything, record five to ten ordinary examples of the existing process. If recent records exist, use them. If they do not, do a short baseline week before starting the AI trial. Do not reconstruct times from memory after adopting the new workflow because memory will favor the new and unusual method.
For each eligible item, record the following from the moment work begins until the result is accepted or sent:
- Task type, complexity band, risk level, input quality, and who performed it.
- Start and finish time, including searching, prompting, waiting, editing, checking, review, handoffs, and correction.
- Direct cost, such as API usage, subscription allocation, compute, external tools, and paid review.
- Output quality against a short rubric, plus material errors, rejected items, and required rework.
- Reviewer minutes and the number of clarification cycles with colleagues, clients, or approvers.
- A lightweight cognitive-load rating immediately after the task.
- The downstream result that matters, such as acceptance on first review, support-ticket resolution, test pass rate, customer response, correction rate, or use by the intended audience.
The baseline must represent the work you actually receive. If the old process was used for routine tasks in one week and the AI process is used for unusually clear, enjoyable, low-risk tasks in the next, the comparison is invalid. Classify work before looking at outcomes and include abandoned or escalated tasks in the log.
Use an end-to-end scorecard
Use a small scorecard rather than a complicated formula that hides important tradeoffs. Record the raw measures, then rate each row as better, unchanged or uncertain, or worse. An individual can do this in a spreadsheet. A small team can add a second reviewer for quality and risk-sensitive work.
| Measure | What to record | Why model speed alone misses it | Suggested decision role |
|---|---|---|---|
| End-to-end time | Median minutes from work start to accepted output | Prompting, waiting, editing, verification, and handoffs can exceed generation time | High |
| Direct cost | Tool, API, compute, and paid review cost per accepted item | A cheap prompt can create expensive review or rework | High |
| Output quality | A task-specific rubric, preferably scored blind when practical | Fluent text can be inaccurate, incomplete, poorly structured, or off-brand | Gate plus high |
| Errors and rework | Factual errors, bugs, unsupported claims, rewrites, and later corrections | Defects often appear after the first draft is counted as done | Gate plus high |
| Review burden | Reviewer minutes, escalation count, and clarification cycles | Work may be shifted from author to reviewer instead of removed | High |
| Consistency | Variation in quality and adherence across comparable tasks | A strong average can conceal an unacceptable tail of failures | Medium |
| Cognitive load | Brief rating of mental demand, effort, frustration, and perceived control | A workflow can be faster but exhausting, or slower but more sustainable | Medium |
| Privacy and security exposure | Data types entered, provider settings, access path, and policy exceptions | A useful output does not compensate for an unauthorized disclosure | Non-negotiable gate |
| Downstream outcome | Acceptance, resolution, test result, adoption, or another business result | Local drafting speed may not improve the actual goal | High where measurable |
For cognitive load, do not needlessly invent a new survey. NASA's Task Load Index is a well-established subjective workload assessment with mental demand, temporal demand, performance, effort, and frustration among its dimensions. A small team can use a short consistent subset after each task, then compare the pattern rather than treating it as a precise clinical diagnosis. NASA TLX
Use gates before totals. A new workflow fails the trial if it produces an unapproved data transfer, creates a high-severity safety or security issue, or falls below the agreed quality floor. For acceptable-risk work, a simple rule is enough: call a measure better only when the observed improvement is meaningful to the team and does not create a worse outcome in a high-priority row. Keep the original measurements available so a persuasive-looking average cannot hide a serious failure.
This approach follows the broader principle behind NIST's AI Risk Management Framework: measure both quantitative and qualitative impacts in the deployment context, document tradeoffs, and monitor them as the system and its risks evolve. The framework also calls out privacy and security as distinct evaluation concerns. NIST AI RMF Core
Run a practical two-week crossover experiment
The aim is not a publishable clinical trial. It is a fair, repeatable decision for an individual or small team. A crossover lets the same person, or the same kind of work, appear in both conditions so that a naturally fast writer or an unusually strong reviewer does not decide the result by themselves.
Before day one
Choose one task type and write the protocol on one page. Define what counts as “old workflow” and “AI workflow,” the approved tools and data, the quality rubric, the risk gates, the time boundary, and the decision rule. Freeze the model, tool configuration, prompt template, and review requirement during the two weeks unless a safety issue requires a change. Give everyone a brief practice session, but exclude practice tasks from the results so learning the interface is not mistaken for ongoing value.
Create comparable task pairs or small bundles using criteria known before work begins: subject matter, complexity, expected length, risk tier, input completeness, and required review. A research brief based on twelve supplied documents should be paired with another brief of similar source count and ambiguity, not with a one-paragraph email. For a team, make the pairs visible and assign them before results are known.
Week one
For each matched pair, randomly assign one item to the old workflow and one to the AI workflow. If several people participate, randomly assign half to begin with AI and half to begin with the old method for a comparable bundle. Record every eligible item, including tasks that are paused, escalated, or abandoned. Do not let people switch conditions midway because a task looks difficult. If an exception is necessary, log the reason and keep it in the analysis.
Week two
Cross over the conditions. The people or task bundles that began with AI begin with the old workflow, and the others do the opposite. Keep the same data boundary and review standard. This reduces the risk that a calendar event, one person's speed, or a single easy category is mistaken for a tool effect.
At the end, compare medians and ranges within each task class instead of only a grand total. Read two or three paired examples where the result disagrees with the average. Ask what actually happened: Did the AI reduce drafting time but add source verification? Did it produce a more consistent structure? Did people use it only for easy work? Did it make the work feel lighter while delaying the final approval? The explanation is often more actionable than a single percentage.
Handle learning and carryover honestly
People usually get better at a new tool while they are testing it. That is a real adoption effect, but it can also bias a short experiment. Counterbalance the order, use a brief uncounted practice period, and record the task sequence number. If week two is much faster in both conditions, the likely explanation may be familiarity with the task or process rather than the AI workflow.
Do not assume that novelty means value. The new method may feel unusually easy because it removes the blank page, even if reviewers later pay the cost. Conversely, a new workflow may look slow in the first week but become valuable once a template, retrieval set, or quality checklist is refined. Label the two-week result as an adoption-stage result and schedule a follow-up on routine work before scaling it widely.
Example with a research or writing workflow
Example
Hypothetical setup: a two-person policy team produces eight 700-word research briefings over two weeks from approved internal reports and public sources. The old workflow is manual outline, search, drafting, citation checking, and editor review. The proposed workflow lets the writer ask an approved AI tool to extract claims from the supplied source pack, suggest an outline, and produce a draft. The writer must verify every factual claim and citation against the source pack before editorial review. No confidential source is sent to an unapproved service.
The team pairs the briefings by topic difficulty and source volume. Within each pair, one uses the old workflow and one the proposed workflow in week one. In week two, the writers switch conditions for comparable topics. They log elapsed time, tool cost, writer check time, editor minutes, unsupported citations, material factual corrections, and first-pass approval. The editor scores clarity, completeness, traceability to sources, and suitability for the intended audience without being told which condition produced the draft when blinding is practical.
The useful conclusion might be mixed. If AI-assisted drafts reduce first-draft time but create unsupported citations that double verification time, the team should not call the workflow faster. It might restrict AI to outline generation and source organization, require the source pack in every prompt, and preserve manual drafting for high-stakes claims. If the paired results show equal or better source accuracy, lower total time, and no additional editor burden, it can keep the workflow for that class of briefings while continuing to monitor it.
Account for the costs people hide from themselves
Cherry-picked wins and changing task mix
Teams naturally remember the excellent output that saved an afternoon and forget the ordinary tasks where prompting added little. Log all eligible work. Mark why an item was excluded, which condition it used, and whether it required an exception. Separate routine, ambiguous, novel, and high-risk tasks before comparing them. An AI workflow may be a good fit for routine transformations and a poor fit for original research, or vice versa.
Hidden coordination cost
Count the work around the model: sharing prompts, preparing context, gaining access, resolving output ownership, explaining an answer to a colleague, reviewing generated changes, keeping templates current, and repairing automation when its assumptions break. A workflow that makes one author faster while creating an unplanned reviewer queue has not necessarily improved team throughput.
Automation complacency
Fluent output can make weak reasoning or missing evidence harder to notice. Set a verification rule that matches the harm of an error. For research, check every citation and material claim. For code, run tests, inspect changes, and apply the normal security review. For customer communication, preserve an approval path for exceptions. Avoid using the model's confidence or prose quality as evidence that the answer is correct.
Make the reviewer independent where possible. Blind review of a sample, predefined checklists, and adversarial examples reduce the chance that the creator is simply defending a favorite tool. NIST specifically recommends documenting test methods, relevant risks, and independent review where appropriate to reduce internal bias and conflicts of interest. NIST AI RMF Core
Speed that damages quality
Speed is harmful when it encourages sending work before it is checked, floods colleagues with low-quality options, displaces deliberate analysis, or expands output volume beyond the team's ability to review it. Protect the quality floor with an explicit release criterion. If a faster workflow causes more late corrections, customer confusion, test failures, or reputational damage, record that downstream harm even when the initial task timer looks excellent.
Decide keep, modify, restrict, or stop
| Decision | Evidence pattern | Next action |
|---|---|---|
| Keep | Representative tasks clear the risk gates and show meaningful improvement in end-to-end time, quality, cost, review burden, or downstream outcome | Document the protocol, train users, retain monitoring, and expand only to similar task classes |
| Modify | Some benefits are real, but prompt setup, source grounding, review, or handoffs prevent a reliable gain | Improve templates, retrieval inputs, guardrails, and review checkpoints, then rerun a smaller comparison |
| Restrict | The workflow helps only in low-risk stages or for certain task classes | Limit it to tasks such as brainstorming, outlining, transformation, or test-case generation, with explicit exclusions |
| Stop | No meaningful net benefit, quality declines, privacy or security gates fail, or review costs erase the gain | Remove it from the workflow, preserve the lessons, and revisit only after a material capability or process change |
This is a decision, not a verdict on AI as a category. A team may keep an AI workflow for first drafts, restrict it for fact claims, and stop it for sensitive data. The correct scope is the smallest one supported by the evidence collected.
When to re-evaluate
Run the two-week comparison when adopting a materially new workflow, then maintain a lightweight operating log. For a high-volume process, review the scorecard monthly and rerun a representative comparison every three to six months. For a lower-volume personal workflow, review after enough new examples accumulate or at the next meaningful change.
Re-evaluate immediately when any of these change: the model or tool version, pricing or rate limits, provider data handling or retention terms, enabled connectors and permissions, the prompt template or retrieval corpus, the kind of work being routed to the workflow, the review policy, a material error, or the legal, privacy, security, or customer impact of the output. NIST treats risk management as iterative because context, capability, risk, and impact evolve in use. NIST AI RMF Core
Keep the baseline records and the decision log. They make it possible to tell whether a new model genuinely improves the workflow or merely creates a fresh round of enthusiasm. They also make a later rollback straightforward when a model, price, policy, or work mix changes.
Evidence
Sources used for this answer.
Question signals show what people need. Primary documentation supports the answer. Both remain visible.
- 01How has your AI workflow changed over the past year?Reddit · question signal · checked 1 Sept 2026
- 02NBER customer-support studynber.org · primary evidence · checked 1 Sept 2026
- 03METR developer studymetr.org · primary evidence · checked 1 Sept 2026
- 04Observed discussionreddit.com · primary evidence · checked 1 Sept 2026
- 05NASA TLXnasa.gov · primary evidence · checked 1 Sept 2026
- 06NIST AI RMF Coreairc.nist.gov · primary evidence · checked 1 Sept 2026