Observe the process with the people who do the work before choosing an AI solution. Define the user outcome, map routine cases and exceptions, and identify why delays or errors occur. Remove unnecessary steps and address unclear rules, missing ownership, and poor inputs.
Measure the improved process before adding AI: time, cost, rework, errors, and customer outcomes. Then identify any remaining task where handling varied language or documents could benefit from AI. A simpler form, clearer policy, or ordinary automation may already solve much of the problem.
Pilot the AI-assisted version on a limited set of work and compare it with the redesigned alternative. Include review and correction effort in the result. Expand only when the complete process improves and its failure handling is workable.
The problem behind the question
The observed Ask HN question asks whether organizations are automating with AI instead of analysing and designing first. Original question That concern is well founded. A process can be inefficient because it asks for information nobody uses, routes work through unnecessary approvals, has no clear owner, contains contradictory rules, or compensates for a defect upstream. An AI layer can make those defects move faster, less visibly, and at a larger scale.
This does not mean AI should be considered only after a lengthy documentation exercise. It means the organization needs enough evidence to make a credible choice among several interventions: eliminate work, simplify it, standardize it, change a policy or interface, improve a data source, automate deterministically, add AI assistance, or leave expert judgment in place. The default should be to seek the smallest change that reliably improves the user’s outcome.
The sequence aligns with established process and risk-management practice. Value-stream mapping starts from the actual material and information flow and distinguishes current state from a desired future state. Lean Enterprise Institute on value-stream mapping NIST’s AI Risk Management Framework similarly calls for understanding the context, intended purpose, affected people, risks, benefits, and the need for an AI solution before deciding to deploy one. NIST AI RMF Core, Map function
Define the outcome
A task is an activity performed by the organization. An outcome is the condition the user needs to achieve. Confusing the two is the first route to bad automation.
For example, “review every invoice email” describes a task. “Pay valid invoices on time, prevent duplicate or unauthorized payments, and give suppliers a clear status” describes the outcome. The second wording makes it easier to see that some emails may not need review at all, that invoice submission may be redesigned, and that payment release needs controls that an email-reading agent should not bypass.
Write a short outcome statement before selecting a tool:
A named user needs to achieve a specified result in a stated context, with an acceptable time, quality, cost, and risk boundary.
The statement should name the user, their goal, the triggering event, and the harm if the process gets it wrong. Government service guidance makes a similar distinction: user needs should describe the person’s problem, not a proposed solution, and should be based on research rather than assumptions. GOV.UK guidance on user needs
Questions leaders should answer first
- Who receives the benefit or harm: an external customer, employee, partner, regulator, or another team?
- What starts the work, and what counts as a completed outcome from that person’s point of view?
- Which steps create value, satisfy a necessary control, or are required by a contract or policy? Which steps are merely habits, duplicate entry, or compensation for another defect?
- What decision is actually being made, and what evidence, authority, and accountability does it require?
- What happens in a rare but material case? What is the cost of a false approval, missed escalation, wrong communication, or delayed response?
- What would make the organization decide that no automation is the right outcome?
If these questions produce different answers from different leaders, workers, or customer groups, the process is not yet ready for AI automation. That disagreement is useful discovery, not a reason to ask a model to paper over it.
Discover the real process
The documented procedure is often a simplified ideal. The real process includes spreadsheet workarounds, informal messages, repeated clarifications, manual corrections, queue jumping, and tacit knowledge. Design from the real process, otherwise the AI will be trained or prompted for the process people wish existed.
Use three kinds of evidence
Walk a case end to end. Observe a representative worker complete a normal case and an exception. Record every handoff, system, decision, wait, re-entry of data, and correction. Ask what makes the step necessary and what would happen if it were skipped.
Analyze operational traces. Case identifiers, activity names, timestamps, owner or queue, status changes, outcomes, and rework markers can reveal actual paths, waiting time, loopbacks, and variation. Process mining is one way to reconstruct a process from event data. A usable event log associates each event with a case, activity, and time; it may also include resource, cost, and other attributes. Process Mining event-data guidance
Listen to affected people. Interview frontline workers, supervisors, support, compliance, data owners, customers, and teams downstream. People doing the work know the exceptions that logs cannot explain, such as when a supposedly complete field is routinely wrong or when escalation rules create avoidable delay. Involving them also exposes whether a proposed improvement shifts effort or risk onto somebody else.
Use sampling deliberately. Analyze high-volume routine cases, but also sample expensive, delayed, reworked, escalated, and customer-complaint cases. An average can hide the small share of cases that consume most effort or cause the most harm. Conversely, do not treat every observed variation as waste. A variation may be an appropriate response to a different customer need or risk.
Map the current state
A current-state map can be a simple table before it becomes a formal diagram. For every step, capture:
| Field | What to record | Why it matters |
|---|---|---|
| Trigger and input | Who or what starts the step, data supplied, source system | Reveals missing, duplicated, or low-quality inputs |
| Activity and decision | What the worker or system does and which rule is applied | Separates judgment from clerical work |
| Owner and handoff | Role, queue, team, and required authority | Locates waiting, ambiguity, and accountability gaps |
| Time | Work time, queue time, deadline, and variation | Shows whether effort or waiting is the real bottleneck |
| Output and recipient | Decision, record, message, or transaction produced | Tests whether the output helps the next person or user |
| Rework and exception | Corrections, loopbacks, overrides, complaints, and cause | Identifies unstable steps that are risky to scale |
| Control and evidence | Required approval, audit record, privacy rule, or reconciliation | Keeps simplification from removing a necessary safeguard |
Do not begin with a generic process map built in a workshop. Start with real cases, then have the workshop reconcile the observed evidence with the intended policy. Process-discovery methods can help visualise actual behavior, but their output is only as complete as the selected data. Event logs can omit offline conversations, unrecorded judgments, and activities in systems outside the analysis scope. Process Mining on discovery limits and observed behavior
Remove, simplify, and redesign before adding AI
Once the current state is visible, challenge every step. The order matters because it prevents an organization from paying an AI system to preserve work it should delete.
| Intervention | Test | Typical result |
|---|---|---|
| Remove | Does this step change a decision, reduce a real risk, or meet a requirement? | Delete duplicate reports, duplicate data entry, unused approvals, and notifications nobody reads |
| Simplify | Can a user provide the information once, can a rule be clarified, or can an interface prevent the error? | Shorter forms, fewer categories, plain-language policy, validated fields |
| Standardize | Are routine cases being handled differently without a justified reason? | Defined inputs, ownership, decision rules, service targets, and exception taxonomy |
| Fix upstream | Is this work compensating for a source-system, product, or policy defect? | Better data capture, clearer entitlement, changed handoff, removal of a recurring exception |
| Deterministic automation | Is the desired behavior fully specified and stable? | Rules engine, workflow, calculation, integration, or template with predictable execution |
| AI assistance | Is there useful ambiguity or unstructured input after the preceding changes? | Extraction, classification, summarization, retrieval-supported drafting, or recommendation with bounded authority |
The table is not a ban on AI. It is a decision sequence. AI is a good candidate when the remaining work requires handling variable language, documents, images, or patterns and the organization can define how a person or system should verify and act on its output. A deterministic rule is normally better when the policy is clear and the input is structured. A redesigned form is better when the problem is that key information was never captured.
Consider the cost of exceptions. A process that is 90 percent routine but has 10 percent high-risk cases may need a design that separates the paths. Automate a narrow, verifiable part of the routine work, then route the exceptions with context and a clear owner. Do not set an AI agent loose across both paths simply because one model can technically read all the inputs.
Establish a baseline that makes comparison possible
Without a baseline, a pilot can appear successful because people are excited, demand happens to be low, or the agent produces a polished output. Leaders need to compare a changed process with the relevant alternative: often the redesigned manual or rules-based process, not the old inefficient one.
Measure a small set of metrics before changing anything. The metrics should cover the user outcome, operations, quality, risk, and economics.
| Dimension | Example measures | Common trap |
|---|---|---|
| User outcome | Resolution rate, time to a usable result, customer effort, complaint rate | Counting fast replies that do not solve the problem |
| Flow | End-to-end lead time, queue age, touch time, rework rate, handoffs | Reporting average completion while the long tail gets worse |
| Quality | Correct decision rate, missing-data rate, override rate, downstream defect rate | Treating fluent model output as correct output |
| Control | Unauthorized actions, audit completeness, privacy or policy exceptions, reconciliation breaks | Measuring only the model and not the full workflow |
| Economics | Cost per successful outcome, model and tool usage, human review time, change and maintenance cost | Ignoring the labor displaced into checking, repairing, and escalation |
| Workforce | Training time, workload distribution, worker-reported friction, adoption, safety incidents | Calling a burden transfer a productivity gain |
Record the population and window used for each baseline. A pilot that excludes the difficult cases or runs during an unusually quiet month cannot justify a broad rollout. Segment results by work type, channel, language, customer group, data source, and exception severity where these differences materially affect outcomes. NIST’s framework recommends measuring performance and assurance criteria in conditions similar to the intended deployment setting and documenting the methods and metrics used. NIST AI RMF Core, Measure function
Use a counterfactual, not a slogan
Before the pilot, make the comparison explicit:
- Current process: how it performs now, including its defects and hidden effort.
- Redesigned non-AI process: the best credible improvement using removed steps, clarified rules, changed forms, or deterministic workflow.
- AI-assisted process: the exact added capability, authority, review path, and operational cost.
If the redesigned non-AI version produces nearly all the benefit, choose it. It is often easier to test, explain, audit, maintain, and recover. If AI adds a material improvement after the redesign, the comparison clarifies what it is actually contributing.
Analyze failure modes before scaling
Teams often evaluate an AI system by asking whether it can perform the happy path. Leaders should instead ask how the combined process fails, who notices, who can stop it, and whether the harm is reversible.
For each important step, identify the failure mode, cause, detection method, consequence, owner, and response. This can be done in a lightweight failure-mode analysis without producing a large risk register. The aim is to find the few failures that change the design.
| Failure mode | Likely cause | Design response |
|---|---|---|
| Wrong classification routes a critical case to a low-priority queue | Ambiguous language, weak examples, missing context | Confidence threshold, critical-keyword guardrail, sampled review, direct escalation path |
| Plausible but unsupported answer is sent to a customer | Weak retrieval, prompt injection, no source check, excessive autonomy | Approved source set, citations or evidence links, constrained templates, human review for high-impact replies |
| Agent takes an action twice | Retry after a timeout or unclear state | Idempotency key, durable action log, reconciliation before retry |
| Sensitive information enters an unapproved tool | Broad access, copied conversation, unclear data policy | Data classification, approved environment, access controls, minimization, audit logs |
| Workers override the tool but the pattern is ignored | Feedback is not captured or reviewed | Structured override reason, review cadence, owner for process and model change |
| Automation meets a local speed target but worsens the customer outcome | Narrow metric and no downstream measure | End-to-end measures, customer feedback, rollout gate tied to the outcome |
NIST frames AI risk management as continuous and iterative, with governance, mapping, measurement, and management rather than a one-time approval. It also calls for documented roles, human-AI oversight, testing before deployment and regular testing in operation. NIST AI RMF Core This is directly relevant to process automation because a system can be technically accurate while unsafe in the policy, data, or handoff context surrounding it.
Check data and control readiness
AI should not be asked to compensate for absent data governance or absent operational control. A process is not ready merely because there is a large archive of tickets, emails, PDFs, or recordings.
Data readiness
For each input and output, answer:
- Is the source authoritative for this decision, current enough, complete enough, and accessible with a legitimate purpose?
- Are important concepts represented consistently, or do statuses and categories mean different things across teams?
- Can the organization identify the case, data version, source, transformation, and output used for a decision?
- Are there privacy, confidentiality, intellectual-property, retention, localization, or contract constraints on using the data with the selected tools?
- Is missing, contradictory, or stale information detected and routed rather than silently filled with a plausible answer?
Data lineage is not bureaucracy. It allows a reviewer to investigate an outcome, correct a bad source, determine which cases need reprocessing, and explain the limitation to an affected person. NIST’s AI RMF Playbook includes documenting data selection, curation, preparation, analysis, lineage, limitations, and the relevant governance policies. NIST AI RMF Playbook, Map guidance
Control readiness
Define what the agent may read, recommend, draft, communicate, change, approve, or execute. Then name the accountable person for each material outcome. Controls commonly include least-privilege access, segmentation of duties, an immutable audit trail, approval thresholds, reference data, versioned prompts and policies, monitoring, a kill switch, and a tested fallback path.
The exact controls should be proportional to the potential impact. Drafting an internal summary has a different risk profile from changing a customer’s entitlement, sending a legal representation, approving a loan, modifying a production setting, or releasing a payment. In regulated or safety-sensitive work, the organization must follow applicable requirements and engage its domain, compliance, legal, and security owners. An AI pilot is not an exemption from those obligations.
Pilot narrowly and compare honestly
A useful pilot is a learning instrument, not a public demonstration. It has a limited population, explicit authority boundary, short enough feedback loop, and a pre-agreed comparison. It should be easy to pause or reverse without trapping users in a half-finished new process.
A practical pilot charter
| Element | Example |
|---|---|
| Problem and outcome | Reduce time for suppliers to receive a correct invoice-status answer without exposing payment information |
| Scope | English-language, low-risk status questions for two vendor groups; no payment changes or account updates |
| Comparator | Redesigned self-service status page and triage rule, with and without retrieval-supported drafting |
| Authority | Agent may retrieve approved status and draft; it cannot disclose restricted fields or change records |
| Evaluation set | Recent routine, difficult, and adversarial examples, plus live sampled cases with user feedback |
| Measures | Correct useful answer rate, escalation rate, time to resolution, complaint rate, review time, cost per resolved case |
| Controls | Approved data source, redaction, access checks, logging, human escalation, rollback to the existing channel |
| Duration and decision | Four weeks or a pre-set number of eligible cases, followed by a documented continue, redesign, or stop decision |
Test in conditions that resemble use, including incomplete data, unusual phrasing, changed policies, slow dependencies, and attempts to make the agent ignore instructions. Include users and workers who are likely to meet the actual system, not only sponsors and model builders. The NIST AI RMF identifies domain experts, end users, and other relevant actors as important perspectives in mapping and measuring AI risks and impacts. NIST AI RMF Core
Avoid defining the pilot’s success as “the model produces plausible answers” or “the team used AI.” A successful pilot may show that AI is unsuitable, that a rule must be clarified, that the input data needs repair, or that a redesigned workflow makes the agent unnecessary. Those are high-value findings if they prevent an expensive rollout.
Involve workers as co-designers and reviewers
Involve the people who handle routine work and exceptions. Workers reveal the practical distinctions between a routine case and an exception, which sources are trusted, which fields are routinely unreliable, and where customers get stuck. They can identify when a proposed automation transfers invisible work to review, correction, or apology.
Involve workers early in the current-state walk, future-state design, evaluation rubric, safety review, and post-pilot review. Give them a way to override or correct the agent, capture the reason in a structured form, and see what happens to recurring feedback. Do not measure worker performance solely by how often they accept the tool. An increased override rate may be a useful warning that the process, data, or authority model is wrong.
Leadership should also be clear about the intended workforce effect. Is the aim to remove a clerical task, improve response capacity, reduce errors, support training, or reduce staffing? The answer affects trust, training, workload, and the quality of feedback. ISO’s quality-management guidance emphasizes customer focus, process approach, defined responsibilities, performance evaluation, and continual improvement based on evidence. ISO on the ISO 9001 process approach Those principles apply whether the improvement includes AI or not.
Define stop criteria before the pilot begins
Pre-commitment prevents a team from explaining away evidence after it has invested in a vendor, a model, or a launch announcement. Stop criteria should be measurable and tied to the harm that matters, not only to aggregate accuracy.
Examples of stop or rollback criteria include:
- A material control, privacy, security, or unauthorized-action incident.
- Error, complaint, or escalation rates above a stated threshold for a critical work class.
- Evidence that a protected or vulnerable group receives materially worse outcomes, pending investigation and remedy.
- Human review, correction, or support time that eliminates the expected capacity benefit.
- Cost per successful outcome above the agreed limit, including vendor, infrastructure, review, repair, and monitoring costs.
- A change in policy, data quality, model behavior, dependency reliability, or context that invalidates the evaluation assumptions.
- Inability to produce the evidence, audit record, or explanation required for an important decision.
Also state what happens on a stop: suspend automated action, return to the prior safe workflow, preserve logs and affected-case lists, notify the accountable owner, remediate user harm where necessary, investigate the cause, and decide whether to redesign, re-test, or retire the use case. NIST’s management guidance includes monitoring, appeal and override, incident response, recovery, change management, and decommissioning as elements of post-deployment AI risk management. NIST AI RMF Playbook, Manage guidance
Example where redesign beats AI automation
Imagine a company wants an AI agent to read vendor invoice emails, classify them, request missing information, and route them to accounts payable. The stated problem is that invoices are paid late and staff spend many hours reading email attachments.
Process discovery finds that the volume is not the main cause. Most delayed invoices are already valid when received. They wait because suppliers use several email addresses, purchase-order numbers are often missing from the request process, the same invoice enters both a shared mailbox and an upload portal, and a manager approval is required even for recurring low-value purchases. Staff repeatedly contact suppliers to reconstruct data that should have been captured at purchase time. The email-reading task is real, but it is an expensive symptom of an upstream design problem.
The redesigned process makes a supplier portal or structured submission channel the default, validates purchase-order and supplier identifiers at submission, shows real-time status, deduplicates email and portal submissions, routes recurring low-value purchases through a policy-approved path, and creates a small exception queue for genuine ambiguities. A deterministic workflow can reject incomplete submissions with a precise explanation and create the required audit record. Baseline measures show fewer duplicate cases, lower queue age, and fewer supplier follow-ups before any AI is introduced.
AI may still be useful at the boundary: extracting fields from genuinely unstructured legacy attachments or drafting a helpful response for the exception queue. It is not the core fix. The redesign resolved the main problem. AI remains useful for the remaining attachments and exceptions that require language interpretation.
A leader’s decision gate
Use this gate before funding or expanding an AI automation initiative:
- Can we state the user outcome and harm boundary without naming a technology?
- Have we observed representative current-state cases and collected data on time, quality, rework, exceptions, and cost?
- Have we removed unnecessary steps and considered policy, form, ownership, upstream-data, and deterministic-workflow fixes?
- Is there a defined remaining task where AI’s ability to handle unstructured or variable input creates a material benefit?
- Are data sources authoritative, permitted, traceable, and sufficiently current for the proposed decision?
- Are authority, escalation, audit, privacy, security, monitoring, and fallback controls clear and staffed?
- Can we run a bounded pilot that compares AI with the best redesigned non-AI alternative?
- Have workers, domain experts, affected users, and control owners participated in the design and evaluation?
- Are success measures and stop criteria agreed before the pilot, including user outcome, quality, cost, and workforce effects?
If the answer to any of the first six questions is no, pause the automation work and improve the process design or operating context. If the last three are no, do not expand beyond a controlled experiment. This is not delay for its own sake. It reduces the chance of scaling an opaque, costly substitute for analysis.
Evidence
Sources used for this answer.
Question signals show what people need. Primary documentation supports the answer. Both remain visible.
- 01Ask HN: Are we automating with AI instead of analysing then designing?Hacker News · question signal · checked 4 Sept 2026
- 02Lean Enterprise Institute on value-stream mappinglean.org · primary evidence · checked 4 Sept 2026
- 03NIST AI RMF Core, Map functionairc.nist.gov · primary evidence · checked 4 Sept 2026
- 04GOV.UK guidance on user needsgov.uk · primary evidence · checked 4 Sept 2026
- 05Process Mining event-data guidanceprocessmining.org · primary evidence · checked 4 Sept 2026
- 06Process Mining on discovery limits and observed behaviorprocessmining.org · primary evidence · checked 4 Sept 2026
- 07NIST AI RMF Playbook, Map guidanceairc.nist.gov · primary evidence · checked 4 Sept 2026
- 08ISO on the ISO 9001 process approachiso.org · primary evidence · checked 4 Sept 2026
- 09NIST AI RMF Playbook, Manage guidanceairc.nist.gov · primary evidence · checked 4 Sept 2026