AI tools have produced measurable gains in some business settings and slowdowns in others. A customer-support study reported about 14% more issues resolved per hour, while a study of experienced developers using early-2025 tools found tasks took 19% longer. The studies involved different users, tools, and work; neither establishes a general return from autonomous agents. Customer-support study and developer experiment
For an agent in your workflow, measure the completed work, including checking, corrections, and exceptions. A faster draft can still leave the team with more work overall. Compare similar tasks with and without the agent, using the same quality standard, and expand only where the results show a benefit worth the cost and risk.
What counts as a productivity gain
Productivity is the value of accepted output relative to all inputs required to produce it. In a business workflow, the output should be a completed, correct, policy-compliant unit that the next person, customer, or system can use. Inputs include not only agent runtime and the operator's prompt, but also data preparation, integration, licensing, monitoring, review, corrections, incident handling, and the time of people who receive the work downstream.
That definition prevents two common errors. First, an agent may create a draft in two minutes yet add twenty minutes of expert verification. Second, it may remove time from one queue while moving the same work into escalations, complaints, back-office reconciliation, or customer self-service. Neither is a business gain unless the end-to-end result improves.
Use both a workflow metric and a business metric. The workflow metric shows whether the team can complete valid units more efficiently. The business metric shows whether that improvement changes a result the organization cares about, such as resolution quality, customer retention, cash collection, fraud loss, conversion, delivery reliability, or capacity released for higher-value work. Do not add these into a single opaque score. A workflow that improves throughput while materially harming quality should fail even if a dashboard still shows a positive average.
The available evidence supports this cautious approach. An OECD review of experimental research finds meaningful gains in some tasks, but concludes that effects depend on the task and the user's experience and that long-term firm effects remain under-studied. OECD review of generative AI productivity evidence Evidence for fully autonomous, tool-using agents in production is thinner still, so leaders should not infer enterprise return from benchmark scores or chatbot adoption.
Define a task unit before selecting a tool
A task unit is the smallest result that counts as complete for the workflow. It must include its acceptance condition. Good units are concrete: a customer case resolved without reopening within the defined period; an invoice matched and posted with required evidence; a claims packet complete and passed to an authorized decision maker; a compliant product listing published; or a software change accepted through the normal review and test process.
Avoid units such as prompts sent, documents generated, messages drafted, tickets touched, model tokens, or agent steps completed. They are activity measures. They can rise while actual output, customer experience, or compliance falls. If work items vary greatly in difficulty, create a complexity class before the pilot begins or use a weighted workload measure agreed by operations and finance. Do not let the agent receive only the easiest cases and then compare its rate with the team's normal mix.
For each unit, record a start event, an end event, the acceptance test, the source system of record, the severity or complexity band, and whether an exception occurred. The same definitions must apply to the comparison group. If a pilot cannot produce this instrumentation, it is not ready to make a productivity claim.
The minimum scorecard for an agent pilot
| Dimension | What to measure | Why it matters | A practical definition |
|---|---|---|---|
| Cycle time | Median and upper-percentile time from intake to accepted completion | Average time can hide a slower tail that harms customers | Measure elapsed time for each completed task unit and report the 50th and 90th percentiles by complexity band |
| Quality | Acceptance rate, defect rate, and severity of defects | More output is not better if errors reach customers or decisions | Use a pre-defined rubric, independent sampling, and error categories that distinguish cosmetic defects from material ones |
| Rework | Reopens, corrections, reviewer edits, returns from downstream teams | An apparent local gain may simply defer work | Count time and causes from first completion through the agreed stabilization window |
| Exceptions | Escalations, tool failures, policy blocks, manual fallbacks, and unsafe actions | Agents are most fragile at the boundary cases that consume operations time | Report the rate, type, queue age, and resolution effort for every exception |
| Adoption | Eligible work actually routed through the agent and retained use over time | A tool cannot create workflow value if trained users bypass it | Measure eligible volume, opt-outs, repeat use, and stated reasons for non-use |
| Oversight | Human review time, interventions, overrides, and supervisor load | Review is an input, not free quality assurance | Include reviewer minutes and the share of cases needing expert correction or escalation |
| Cost | People, model and vendor fees, infrastructure, integration, controls, and remediation | Token price is not total cost | Divide fully loaded cost by accepted, compliant task units and compare it with the baseline |
| Business outcome | The result the workflow was created to deliver | Local efficiency may not create value | Select one or two outcomes with an expected causal link, such as repeat-contact rate, cash collected, conversion, or loss prevented |
Read the scorecard as a set of trade-offs. For example, a lower median cycle time is not a success if the 90th percentile, defect severity, or reviewer burden worsens. Likewise, a more expensive agent may be rational if it raises accepted quality or revenue enough to justify the cost. The pilot owner should write down these trade-offs before seeing results.
Measure the whole cost of an agent
Use a fully loaded cost calculation for each accepted unit. Include employee handling and review time, model and software costs, integration and orchestration maintenance, retrieval or data costs, security and compliance work, evaluation runs, incident response, vendor management, and any customer remediation. Include setup costs too, spreading them over a stated period and expected number of completed cases.
Then ask a separate capacity question. If the agent saves staff time, what actually happens to it? Capacity becomes a financial gain only if the firm avoids external spend, reduces overtime or hiring, increases profitable volume, shortens a revenue-bearing cycle, or redeploys people to work with measurable value. Time saved but not redeployed is still operationally useful, but it should not be reported as cash savings.
Why effects vary so much
Agents are not one technology or one intervention. Their impact changes with model capability, prompt and tool design, retrieval quality, authorization design, task complexity, operator skill, process maturity, data quality, customer tolerance for error, and the amount of hidden context a human already knows. An agent can be excellent at collecting information from a few stable systems and poor at deciding what an ambiguous case means.
Worker and task heterogeneity are especially important. In the customer-support field study, AI assistance raised issues resolved per hour by about 14% overall, with approximately 35% improvement for novice and lower-skilled workers and little to no benefit for the most experienced and skilled workers. NBER summary of the study That pattern can be valuable because it narrows a performance gap, but it also means the average does not predict the impact for a senior specialist.
The opposite result is possible in high-context work. METR's randomized experiment involved 16 experienced open-source developers completing 246 real issues in repositories they knew well. With early-2025 AI tools available, they took 19% longer, despite expecting a speedup. The researchers explicitly caution that this result does not establish that AI slows most developers or other domains. METR study and limitations It does establish why perception, a benchmark, or a vendor demonstration cannot substitute for an experiment in the intended workflow.
Positive firm-level results also depend on the starting point. A forthcoming Columbia Business School working paper reports randomized experiments across seven online-retail workflows, with sales effects ranging from 0% to 16.3% depending on the application's marginal contribution to existing practice. That variation is the lesson. It should not be converted into a forecast for a different firm. Columbia Business School working paper
Segment every result at least by task type, complexity, user experience, customer or product segment where relevant, and exception status. Look for who benefits, who slows down, and who bears risk. A single average can hide an agent that helps routine work but makes the rare cases that matter most slower or less safe.
Design a controlled pilot that can answer the question
Start with a bounded workflow
Choose a recurring workflow with enough comparable cases to evaluate. Identify the person responsible, the result that counts as complete, and the actions excluded from the trial. That gives the comparison a clear scope.
Write a one-page pilot charter before building. It should name the task unit, eligible and excluded cases, intended agent role, authorizations, source systems, human review points, baseline period, comparison method, measures, quality floor, risk owner, expected capacity use, and stop rules. Treat the agent's scaffolding, integrations, and monitoring as part of the intervention. The source discussion behind this question correctly highlights that real workflows require more than a course demonstration, especially where deterministic behavior and rigorous testing are necessary. Original DeepLearning.AI discussion
For high-consequence workflow steps, use the agent for retrieval, document extraction, explanation, triage, and preparation, not for an unbounded final decision. A consumer-credit accept, reject, or refer outcome should remain in a validated deterministic rules or approved statistical-decision system, with the agent able to gather documentation, flag missing fields, summarize a reason code, or prepare an underwriter's packet. The agent may make the workflow more productive without becoming the decision engine.
Choose a comparison that withstands scrutiny
Random assignment of eligible cases, users, or teams to agent-assisted and usual-work conditions is the clearest design when feasible. It protects the comparison from cherry-picking easy cases and from coincidental changes in demand. Stratify by complexity or risk if those strongly affect outcomes, and keep the normal quality and escalation policies in place for both groups.
When randomization is not feasible, use a phased rollout with a contemporaneous holdout queue, matched teams, or a stepped introduction. Record volume, complexity, staffing, policy changes, seasonality, outages, and customer mix. A simple before-and-after comparison is weak evidence when a new policy, training wave, or seasonal demand shift happens at the same time as the agent launch.
Pre-register the primary measures and decision thresholds in the charter. This does not require academic formality. It means leaders cannot declare victory because one favorable number appeared after trying many dashboards. NIST recommends measuring system performance or assurance criteria in conditions similar to deployment, documenting limitations, and sharing pre-deployment test results with people who have release authority. NIST AI 600-1
Include learning and adoption in the design
Agents can improve as prompts, retrieval, tools, and staff practices improve. They can also look good at first because a small group of enthusiasts handles clean cases. Separate a short onboarding period from the measurement window, track version changes, and do not mix radically different agent configurations under one result.
Measure adoption honestly. Record which eligible cases were routed to the agent, when users overrode it, when they bypassed it, and why. Low use may indicate a poor interface, slow performance, missing context, lack of trust, or a workflow where the agent adds less than promised. Do not punish staff for reporting failure. Their overrides and comments are often the quickest way to find a flawed task definition or data connection.
Account for work displaced rather than removed
An agent can shift work in at least five directions: from an operator to a reviewer; from a front-line queue to an exception queue; from internal staff to customers who must correct an answer; from the current period to later rework; or from human effort to vendor, infrastructure, and governance spend. Map these handoffs before the pilot and include them in the outcome model.
Review load deserves special attention. If a reviewer must inspect every generated response line by line, the agent is functioning as a drafting tool, not an autonomous workflow. That may still be useful, but the cost and quality claims should say so. Conversely, reducing review requires evidence that the task, sources, controls, and error detection justify it. NIST notes that pre-deployment tests and generic benchmarks may not reflect the context and impacts of real deployment. NIST AI 600-1
Track downstream consequences for long enough to see them. For support work, that can include repeat contacts, customer satisfaction, transfer rate, and complaints. For claims or finance operations, it can include reversals, audit findings, recovery costs, and time to final resolution. For software work, it can include review acceptance, rollback, defects, maintenance burden, and time spent integrating the change. The correct observation window depends on the workflow's error latency.
Set stop rules before deployment
Stop rules prevent a team from continuing a pilot because it has consumed budget or generated excitement. Set them before live use, attach each to an owner with authority to pause the system, and make the response operational rather than aspirational.
| Trigger | Example pre-agreed rule | Required response |
|---|---|---|
| Safety, security, privacy, or compliance breach | Any confirmed material breach, or an action outside the agent's authorization | Disable the affected capability, preserve evidence, contain impact, and conduct incident review before resuming |
| Quality floor missed | Accepted quality falls below the agreed non-inferiority threshold, or material-defect rate rises above the limit | Pause expansion, route work to the established process, analyze error categories, and retest the fix |
| Exceptions overwhelm operations | Manual fallback or escalation rate exceeds the staffed capacity or queue-age limit | Reduce scope, correct routing or tool failures, and reassess the business case |
| Economics fail | Fully loaded cost per accepted unit is worse than baseline without a justified business-outcome gain | Stop or redesign; do not extend on the basis of a lower token cost alone |
| Adoption fails | Eligible users or cases bypass the agent enough that the planned benefit cannot materialize | Interview users, observe the workflow, and either redesign or discontinue rather than mandate use |
| Unmeasurable outcome | Logs, acceptance data, or comparison group cannot support a credible conclusion | Do not make a productivity claim; fix instrumentation before continuing |
The exact thresholds should be selected by the workflow owner, finance, risk, and operations before the pilot. A stop is not necessarily a permanent failure. It may reveal that the agent belongs on a narrower subtask, needs a better knowledge source, or should be replaced by a deterministic automation. The important discipline is to distinguish a learning outcome from a production success.
Example
Hypothetical commercial-credit intake workflow
In this hypothetical example, a lender receives applications that require identity documents, bank statements, and external credit-agency checks. The final credit decision is governed by approved policy, validated models, and human underwriting rules. Staff want to spend less time locating documents, identifying missing fields, recording evidence, and preparing cases for review. The credit decision remains with the existing process.
Action. A tightly scoped agent reads only authorized documents, extracts fields into a schema, compares the fields with a deterministic completeness checklist, requests missing items using approved language, and creates an underwriter packet with source links. It cannot change the approved decision rules, call an acceptance endpoint, or send a final decision. The pilot compares the agent-assisted intake queue with the usual queue across the same application-risk bands.
Judge the result by whether the packet is complete and its evidence is correct enough for the authorized decision maker to use. Measure time to a complete packet, extraction accuracy, missing-document repeat contacts, underwriter correction time, exception rate, queue age, cost per accepted packet, and eventual decision turnaround. If completeness rises but underwriters spend more time correcting evidence, the workflow has not yet produced a net gain.
How to start the evaluation
- Identify a workflow where delay, rework, or lack of capacity has a visible business cost.
- Define the accepted task unit, complexity classes, quality rubric, downstream outcome, and full cost boundary.
- Establish the baseline with real historical work and include the people who handle exceptions and rework.
- Limit the agent's authority, connect it to authoritative data, and preserve deterministic systems for deterministic decisions.
- Run an instrumented comparison, preferably randomized or with a contemporaneous holdout, and monitor safety and quality continuously.
- Segment results, calculate fully loaded economics, and verify where freed capacity actually went.
- Scale only the tasks and groups that clear the pre-agreed quality, risk, adoption, and value thresholds. Redesign or stop the rest.
Evidence
Sources used for this answer.
Question signals show what people need. Primary documentation supports the answer. Both remain visible.
- 01How suitable is Agentic AI for real business world workflows and are the productivity gains promised real?DeepLearning.AI Community · question signal · checked 4 Sept 2026
- 02Customer-support studynber.org · primary evidence · checked 4 Sept 2026
- 03developer experimentmetr.org · primary evidence · checked 4 Sept 2026
- 04OECD review of generative AI productivity evidenceoecd.org · primary evidence · checked 4 Sept 2026
- 05NBER summary of the studynber.org · primary evidence · checked 4 Sept 2026
- 06Columbia Business School working paperbusiness.columbia.edu · primary evidence · checked 4 Sept 2026
- 07NIST AI 600-1nvlpubs.nist.gov · primary evidence · checked 4 Sept 2026