Skip to main content
AI ProductivityAutomationDecision Systems

How to measure approval fatigue in AI workflows

Measure approval fatigue in AI workflows so human oversight stays focused on costly actions instead of becoming a stream of automatic clicks.

Niels KaspersNiels Kaspers
September 19, 2026
6 min read
How to measure approval fatigue in AI workflows

TL;DR

Approval fatigue is measurable: watch approval volume, timing, risky-action coverage, reversals, and review quality so people intervene where their judgment changes the outcome.

Approval fatigue in an AI workflow is the point where a person is asked to review so often that the review becomes a habit instead of a judgment. The system may look cautious on a diagram. In practice, it is training someone to click through the same low-value interruption until the risky action is hard to see.

That is not just a speed problem. It is an oversight problem. A human gate is valuable when it changes a consequential decision: publishing, spending, changing production data, deleting material, or contacting someone outside the workflow. It is weak when it interrupts routine retrieval or repeats an automatic check the system could have performed itself.

In my own operator workflows, the question is not “where can we put a human in the loop?” It is “where does human judgment actually alter the blast radius?” The answer creates fewer, clearer review moments and makes their evidence inspectable.

What approval fatigue looks like

Approval fatigue normally appears before anyone calls it that. The first signal is volume: a workflow asks for more confirmation without doing more high-impact work. The second is timing: prompts appear in the middle of routine steps, when the reviewer has no context and no reason to pause. The third is behavior: approvals become nearly instantaneous, denials disappear, and the same person carries the burden for decisions they cannot realistically inspect.

The final signal is the most serious. The prompt names an action but not its consequence. “Allow tool call?” is not meaningful oversight when the user cannot see what data will change, who will receive it, or how to undo it. That is blame transfer, not a review mechanism.

Anthropic's research on agent autonomy is useful here: experienced users often move from action-by-action review toward monitoring and intervention. That does not eliminate human accountability. It shows why constant prompting is a poor substitute for a well-designed boundary. Anthropic's containment work makes the adjacent point: a permission prompt and a containment boundary are different controls.

Measure the review system, not only the model

You do not need a governance dashboard before you can see whether approvals are doing their job. Start with five measures.

1. Approval prompts per completed task

Count prompts against completed tasks, not against total model calls. If that ratio rises while the task's risk stays flat, the gate is probably broadening into noise. Segment it by workflow and action type; a production publishing flow should not be compared with a research-only workflow.

2. Risk-weighted approval coverage

Tag actions by consequence. Public, expensive, irreversible, authority-changing, or hard-to-undo actions belong in the high-risk group. The useful question is not whether every action has an approval. It is whether the actions that can do real damage receive the clearest attention.

3. Time to decision, with context

Fast approval is not automatically bad. A reviewer can legitimately approve a well-formed, low-risk request quickly. It is a warning when rapid approval is common for high-risk actions or when the approval packet lacks a clear action, evidence, and rollback path. Track timing alongside action class and prompt contents.

4. Denial, edit, and escalation rate

If nobody ever declines, edits, or escalates a request, the workflow may be exceptionally well designed. More often, it is asking for ceremonial clicks. Look at whether people change tool arguments, request more evidence, defer an action, or stop it. Those are signs that the review point can still carry judgment.

5. Prompt clustering by stage

Map where prompts occur. Early clustering often means the workflow is asking a person to validate facts that an automatic check could verify. Late clustering around the public or irreversible step is usually healthier. This is also where an AI workflow audit helps: trace the action, the state change, the evidence, and the recovery route before moving the gate.

Use four action modes

A simple operating model prevents every uncertainty from becoming a prompt:

  • Act automatically for low-risk, reversible work with clear constraints.
  • Verify automatically when a rule, test, schema check, or preview can catch the obvious failure.
  • Escalate for approval when an action has a meaningful external, financial, production, or irreversible consequence.
  • Stop and ask for judgment when the uncertainty is strategic, ambiguous, or cannot be safely reduced to a rule.

This is more useful than calling everything “human in the loop.” The human should not be a notification sink. They should be the owner of the decision that needs context, tradeoff, or accountability. Approval gates that do not kill automation speed explains the design side; the metrics here show whether that design still works in production.

Make every high-value approval legible

A strong approval request should answer five questions before it asks for a click: what will happen, why now, what evidence supports it, what scope is affected, and what happens if it is wrong. For a publish action, show the destination, the exact artifact, validation results, and rollback path. For a spending action, show the amount, recipient, policy boundary, and reason.

This makes review faster without making it shallow. It also produces an audit trail someone can understand after the event. That matters more than collecting a large number of approval timestamps.

Interactive

Approval-fatigue audit

Use this when an AI workflow feels careful on paper but slow and noisy in practice.

Completion

0%0/5 done

This is the gap between understanding the article and actually using it.

  • Use this block as the practical summary, not just the article ending.
  • If one item feels vague, the article probably needs sharper guidance.
  • A short checklist beats a long recap when the reader needs to act.

What a healthier pattern looks like

The goal is not fewer approvals at any cost. It is fewer low-value approvals and stronger high-value review moments. Routine work stays quiet. Automatic verification handles known failure modes. A human sees the action when their judgment can change the outcome, and they have enough context to exercise it.

That is also the practical answer to what should stay human in an AI workflow. Keep people at the boundary where the choice is consequential, contested, or not reducible to a deterministic rule. Do not use their attention to compensate for missing workflow design.

FAQ

What is approval fatigue in an AI workflow?

It is the loss of meaningful human judgment after repeated, low-context approval prompts train people to approve by habit.

How do you measure approval fatigue?

Track approval prompts per completed task, risk-weighted approval coverage, time to decision, denial or edit rate, and where prompts cluster in the workflow.

Does more approval make AI automation safer?

No. More prompts can create more interruptions without improving control. Safer systems combine clear boundaries, automatic verification, and well-contextualized review for consequential actions.

Which AI actions should require approval?

Prioritize actions that are public, expensive, irreversible, authority-changing, or difficult to undo. Low-risk, reversible work is usually better handled by constraints and automatic checks.

Niels Kaspers

Written by Niels Kaspers

Principal PM, Growth at Picsart

More insight pages

Get in touch

Have questions or want to discuss this topic? Let me know.