How PMs can define an AI prototype's one real user job
Define an AI prototype's one real user job before adding features: set the outcome, boundary, evidence, and review trigger that make a demo useful.

TL;DR
An AI prototype earns its place when it proves one user job can be completed better, within a clear boundary, with evidence a product team can review.
An AI prototype should begin with one real user job, not a list of things a model can do. If a person cannot tell what outcome the prototype helps them reach, the team is not testing a product idea. It is staging a capability demo.
A useful job is concrete enough that a person can recognize completion: turn this customer call into confirmed decisions and owners; identify which support tickets need a human; prepare a first draft of a launch page from approved inputs. "Use AI for research" is not a job. "Help a PM compare three customer segments without inventing evidence" is.
This distinction gets more important as agent infrastructure becomes easier to access. OpenAI's Agents API announcement describes the harnesses, tools, environments, and intermediate state that make long-running work possible. Those capabilities are useful, but they do not decide what a product should do first. That is still product work.
My first-party rule from building product and editorial workflows is simple: before a prototype gets a second capability, it should prove that one person can complete one meaningful job with less uncertainty, less rework, or a better decision. The proof needs to be visible enough that the team can challenge it.
Start with the moment that is currently hard
A user job is not a feature request translated into friendlier language. It is the moment where someone is trying to make progress and currently has to do too much interpretation, coordination, or repetitive work. Start there.
For a product manager, that might be the point after ten research calls when the team needs a decision-ready view of recurring problems. For a growth operator, it might be deciding which existing pages deserve a refresh instead of publishing another generic article. For a support lead, it might be finding the small number of tickets where a wrong answer could create a serious customer problem.
Write the job in this form:
When this person reaches this moment, help them produce this outcome without this costly failure.
That last clause matters. It identifies the thing the prototype must not trade away in pursuit of speed. A research assistant that produces a fast summary but silently drops contrary evidence has not completed the PM's job. A content workflow that drafts quickly but cannot show its sources has only moved the review burden downstream.
This is why an AI product spec needs more than a PRD. The product contract has to name the desired outcome and the quality boundary, not only the prompt or integration.
Make the first version deliberately narrow
A narrow prototype is not a timid prototype. It is a testable one. It gives the team a chance to learn whether the proposed workflow is valuable before it adds multiple tools, modes, user types, and permissions.
Pick one of these boundaries explicitly:
- One user: a PM preparing a weekly decision review, not every employee who has notes.
- One input shape: approved interview transcripts, not any document someone might upload.
- One output: decisions, owners, and open questions in a known format.
- One consequence level: a draft that a person reviews, not an automatic update to a customer record.
- One success measure: fewer missing decisions in the review, not vague "AI adoption."
The boundary is part of the product promise. It tells users what the prototype is reliably for, and it tells the team which attractive requests should wait. This is closely related to a product decision system for AI teams: a clear decision beats a broad list of possibilities because it makes tradeoffs discussable.
Define the evidence before you build the flow
The fastest way to mistake a polished demo for a product is to decide what success means after people have seen it. Set the evidence plan before implementation.
Use three kinds of evidence:
- Outcome evidence: Did the person finish the target job? A PM can identify the decisions and owners; a support lead routes the risky tickets correctly.
- Quality evidence: Did the result preserve the important constraints? It should not invent commitments, hide uncertainty, or take an unauthorized action.
- Workflow evidence: Did the prototype reduce work, or did it merely move it into checking and repair?
OpenAI's agent evaluation guidance makes a useful distinction here: traces help inspect what happened, while datasets and eval runs make a known standard repeatable. For an early product prototype, you do not need a giant benchmark. You need a small set of representative cases and a reviewable definition of good.
That approach also prevents prompt theater. When the result fails, the team can ask a better question than "which prompt wording should we try?" Was the job ambiguous? Was a critical input unavailable? Is the output format wrong? Does this action need a human gate? Or is the model actually failing a well-defined task?
Give the prototype a stop condition
Every prototype needs a reason to stop expanding. Otherwise, capability becomes the roadmap. Set a review trigger before launch: after ten representative uses, after two weeks, after a defined error rate, or when the workflow reaches a higher-risk action.
At the review, decide one of four things:
- Keep: the target user completes the job with credible quality.
- Narrow: the job is real but the current boundary is too broad.
- Rebuild the workflow: the issue is context, interface, ownership, or review design—not raw model capability.
- Stop: the evidence does not show a valuable outcome.
This makes a prototype an evidence-gathering instrument rather than a permanent demo. It also creates a clean connection to an AI eval backlog: recurring misses should become cases that protect the real job, not an endless collection of screenshots.
Use a one-job prototype brief
Before development, I would put the following on one page:
Interactive
One-job AI prototype brief
Use this before a promising capability turns into a feature pile.
Completion
This is the gap between understanding the article and actually using it.
- Use this block as the practical summary, not just the article ending.
- If one item feels vague, the article probably needs sharper guidance.
- A short checklist beats a long recap when the reader needs to act.
The brief is intentionally smaller than a full product requirements document. Its purpose is not to pre-solve every implementation detail. It makes the first bet legible enough that design, engineering, and product can disagree productively.
Do not confuse a capability with a user job
A capability is something the system can do: browse, summarize, call a tool, write code, classify a document, remember a preference. A user job is the outcome someone is hiring that capability to achieve under real constraints.
The difference sounds semantic until a team tries to prioritize. Capability-led roadmaps accumulate buttons because each new demonstration looks impressive. Job-led roadmaps make a harder but healthier choice: which outcome is important enough to earn a reliable workflow, evidence, and maintenance?
The product surface should reinforce that choice. In how to use internal links as product navigation, the same principle appears on a website: a useful route helps a person continue a task; a flat catalogue asks them to choose again. AI products have the same risk. A menu of magical actions is not a workflow.
The practical test
Ask a potential user to finish the job with the prototype. Then ask four questions: What were you trying to get done? What did you trust or distrust? What did you still have to do manually? Would you choose this route again for the same moment?
If the answers are fuzzy, make the job narrower before making the prototype smarter. If they are clear, you have something more valuable than a demo: a bounded product hypothesis that can earn the next investment.
FAQ
What is an AI prototype user job?
It is the specific outcome a person needs at a particular moment, expressed with the constraint that must not be compromised. It is more concrete than a capability such as summarization or tool calling.
How narrow should an AI prototype be?
Narrow enough to name one user, input shape, output, authority level, and success measure. Expand only after the first workflow shows a useful outcome with acceptable quality.
Should product teams build evals before launching an AI prototype?
Build a small set of representative cases and a review rubric before launch. Early evals do not need to be exhaustive; they need to protect the user job the prototype claims to solve.
What should make a team stop an AI prototype?
Stop or narrow it when users do not complete the intended job better, when the review burden cancels the benefit, or when the quality boundary cannot be made credible with the available workflow and evidence.