Semitora.

29 June 2026 · Updated: 2 September 2026

Production-grade AI agents — how they differ from a chatbot, and how to keep control

A chatbot answers. An AI agent acts: it reads documents, uses tools and carries out steps in a business process. That difference matters, but it is not enough. An agent becomes a production system only when the team agrees its specification, data sources, operating boundaries, acceptance criteria and stop condition. Without them, even an impressive demo remains a demonstration.

Three jobs have recurred in recent conversations about new agents: preparing a reply to a request for quote from an email and its attachments, handling an appointment over the phone, and producing content under durable brand rules. The sectors and channels differ. The design problem does not: define what the agent must do, what it may use, who approves the result and how anyone can prove that it completed the job correctly.

A production-grade AI agent needs boundaries, tests and decisions

Table of contents

A chatbot answers; an agent takes action

A chatbot receives a question and generates an answer. If it uses RAG, it should ground that answer in identified sources. An agent adds another layer: it chooses steps, calls tools and may change the state of a process. It may create a draft, complete a form, initiate a call or record a test result.

A useful boundary is:

An agent does not need permission to send or write autonomously to be useful. In a quoting process it can collect requirements and prepare a message while a sales owner approves it. In a content process it can assemble a complete publishing package while the brand owner retains the publishing decision.

What real enquiries reveal about AI agents

The analysis below combines patterns from independent sales and pilot conversations in 2026. It is a qualitative demand signal, not a representative survey. We do not publish names, domains, direct quotations or client-specific configurations. We also separate capabilities already exercised in a pilot from capabilities that still require testing.

Three demand patterns for AI agents: quoting, appointments and content

Requests for quote with documents

A specification rarely fits entirely in the email body. Parameters may sit in a PDF, spreadsheet or technical drawing, while material decisions remain in an older thread with the same customer. The agent should assemble those facts, identify the source of every parameter and prepare a draft for approval.

In pilots of this shape, customer recognition, a durable account card and draft generation are typically the first parts to work. Interpreting complex attachments is a separate capability to evaluate. It is not proven merely because a model handled a few clean PDFs.

Phone appointments and bookings

A voice-AI enquiry often sounds simple: answer the call, provide information and book a slot. In practice, the outcome depends on telephony, the authoritative availability source, service-selection rules, human handoff and evidence that the booking really appeared in the system.

This is currently a demand signal and a design pattern, not public proof of a Semitora voice capability. A defensible acceptance test needs a test number, synthetic caller profiles, recordings, transcripts and a way to inspect the final outcome. Before an end-to-end test, we can publish an evaluation method but not a claim of a completed deployment.

A content agent with durable rules

A brand owner may have documents, previous articles, images, style rules and a long operating protocol. The difficult part is not generating one piece of text. It is retaining consistency across sessions, citing the right sources and adding only approved changes to the rules.

A sensible first stage produces one bounded outcome, such as a complete article with metadata and image descriptions. That is the scope of a planned pilot, not a measured case study. Only work on real material can establish owner time, correction count and protocol adherence.

All three examples lead to the same rule: agree the process and its acceptance criterion before choosing a model. Models can be replaced. An unclear outcome and missing evidence remain defects regardless of the vendor.

Five gates before model selection

Five decisions should close before an agent is configured. Each one produces an artefact that can be reviewed later.

Five AI-agent delivery gates before model selection

  1. Mechanism. What triggers the process and what outcome must exist at the end? “Customer service” is too broad. “Prepare a draft response to a product enquiry” can be specified and tested.
  2. Specification. Which steps, sources, exceptions and business rules apply? The specification should also cover conflicting sources and knowledge updates.
  3. Acceptance criterion. What is a correct result? Define the test case, expected outcome, threshold and blocking error types.
  4. Evidence owner. Who approves sources, thresholds and the result? A consultant can design the method but should not take over the business decision.
  5. No-go condition. When should the project stop, narrow or roll back? Examples include no safe access to the system of record, an unacceptable blocking-error rate or a unit cost above the process value.

This sequence changes a buying conversation. Instead of asking whether a vendor has “the best model”, ask for the specification version, test set, evaluation result, decision owner and rollback path. The AI vendor selection guide provides a broader question set.

Autonomy is a decision for each operation

Do not assign one autonomy level to the product. The same agent may retrieve data without approval, prepare a draft for review and be prohibited from sending it. The boundary depends on the impact of the operation and how easily it can be reversed.

For high-risk systems this maps onto the human-oversight duty — see how to classify AI Act risk.

Operation Initial mode Who approves Evidence
Retrieve approved knowledge automatic source owner approves the scope quotation, source identifier and version
Extract parameters from a document automatic with uncertainty flags person familiar with the document fields, coordinates or source excerpt
Draft a message or article automatic sales or brand owner version diff and approval decision
Send a message after confirmation communication owner content, recipient, time and transport identifier
Create a booking or change a record after confirmation, then based on test results process owner before-and-after state and operation identifier
Handle an exception human handoff nominated operator escalation reason and complete context

Test evidence may justify more autonomy for selected operations. Such a change is a versioned decision, not a reward for the model after a few successful demos.

How to define acceptance criteria

“The agent works” settles nothing. Criteria must describe the process outcome, not the fluency of the model response.

Process Outcome to inspect Example measures Blocking error
Request for quote complete draft grounded in the email and attachments required-field coverage, source correctness, material correction count invented parameter or ignored attachment
Phone appointment correct conversation and verified outcome required-question completion, latency, handoff success, booking result invented slot, wrong price or missing escalation
Brand content package consistent with sources and approved rules owner time, correction count, protocol adherence, metadata completeness unsupported claim or obsolete rule

The process owner sets thresholds with the people accountable for data, risk and operations. There is no universal safe percentage for every agent. An error in an internal draft has a different consequence from an incorrect customer booking.

What separates a production system from a demo

A strong model can sit inside either one. The operating environment makes the difference:

A production system does not mean full autonomy. It means the team understands the operating range and can detect, stop and explain a failure.

How to monitor an agent after release

A pre-release test result is not enough because sources, instructions, tools and case mix change. Track three operational signals after deployment:

Do not compare these values without context. The same rate can mean something different for internal drafts and for customer-visible bookings. Record the process type, failure impact, and the model, instruction, knowledge and integration version with every result.

How to test an agent before production

Testing starts with the specification, not an open conversation with a bot. Build normal, boundary and blocking scenarios first. Then run the agent on synthetic data, a controlled real-data slice and the full workflow. Only the result with open risks should reach the go/no-go decision.

AI-agent test loop from specification to monitoring

A practical sequence is:

  1. describe the input, expected action, outcome and evidence;
  2. build scenarios for missing data, conflicts and attempts to leave scope;
  3. run synthetic tests with no production consequence;
  4. use a representative and approved slice of real data;
  5. perform an end-to-end test in a controlled environment;
  6. record failures, fix the system and rerun failed tests;
  7. make a go/no-go decision with the process owner;
  8. after release, monitor the trend and run regression after every material change.

Our safe AI PoC in 30 days explains the bounded experiment in more detail. An agent that works with correspondence also benefits from the separate guide to AI email query handling.

Common mistakes

The instruction form appears at the end of delivery

A common delivery pattern is that the full instruction template only surfaces weeks into the engagement. Approved rules are necessary; the form itself is not the defect. The defect is revealing this work late and shifting the whole analysis onto the process owner.

A better approach is for the agent or delivery team to draft the specification from sources. The owner answers only unresolved questions and approves the rules.

A demo replaces an outcome test

A fluent call does not prove that an appointment was booked. A polished email does not prove that drawing parameters are correct. On-brand copy may still contain an unsupported claim. Test the result in the target system and inspect its sources.

The agent receives permissions that are too broad

Access “just in case” speeds up a prototype but increases failure impact. Start with read-only access, then drafts, then confirmed operations. Consider additional autonomy after evidence, not before it.

There is no no-go path

If every PoC must become a deployment, evaluation is decorative. No-go can be the correct result: the source is too unstable, the business system exposes no safe integration, or exception handling consumes the expected benefit.

When you do not need an agent

If the task ends with an answer grounded in sources, RAG may be enough. If the rules are stable and deterministic, an API or ordinary automation may be better. An agent makes sense when the workflow requires interpreting heterogeneous inputs, selecting subsequent steps and using several tools under control.

Our API, RPA, RAG or AI agent decision matrix compares these options. The cheapest sound architecture is not always an agent. Sometimes the best intervention is removing an unnecessary step from the process.

FAQ

How does an AI agent differ from a chatbot?

A chatbot generates an answer. An agent performs a sequence of steps and uses tools, for example reading a document, preparing a draft and recording a test result. If it can change data or send messages, it needs separate boundaries, approvals and an action trail.

Can an AI agent operate without human approval?

It can perform low-impact operations automatically inside an approved scope, such as retrieving knowledge or producing a draft. Sending, publishing and writing to a business system usually begin behind an approval gate. Further autonomy depends on test results and reversibility.

What data does an agent pilot need?

The minimum is one process description, approved sources, representative inputs, expected outcomes and exceptions. There is no need to copy an entire inbox or database at the start. The tests determine the required data scope.

How should agent quality be measured?

Measure the process outcome: completeness, source correctness, correct action, escalation, correction count, unit cost and failure handling. Link every result to the model, instruction, knowledge and integration version.

When should an AI-agent PoC stop?

Stop when no measurable outcome can be defined, there is no lawful or safe data access, blocking errors remain above the agreed threshold, or post-automation process cost does not justify further investment. Such a no-go is a result, not a failed evaluation.

What next

If you have a process built around email, documents, spreadsheets or calls, send us a short description. Include the triggering event, expected outcome, sources and the operations the agent must not perform alone. We will tell you whether the process needs an agent or a simpler architecture, and which acceptance criteria must be agreed before model selection.

How we build agents in production is on our AI Agents page.