An AI agent is a language model that works toward a goal in a loop: it decides on a step, calls a tool (search a database, create a ticket, send a draft), looks at the result and decides what to do next. That loop is powerful and, without limits, can also take wrong actions quickly.
Good first projects
- Triage incoming support emails or tickets: classify, summarize and route, with a person approving anything customer-facing.
- Drafting replies, quotes or reports from your own data for a human to review.
- Reconciling records across systems and flagging mismatches for staff.
- Answering internal questions from your documents, with citations so answers can be checked.
- After-hours voice assistants that take messages, book appointments and hand off to a person.
Poor first projects
- Anything irreversible, such as payments, deletions or contract changes, that runs without approval.
- Decisions about people (hiring, credit, health) where errors are costly and regulated.
- Vague goals with no way to tell whether the answer was right.
Guardrails to build in from the start
- Least privilege: give each tool the narrowest access it needs, read-only wherever possible.
- Human approval for any action that writes, sends or spends.
- Treat retrieved text (emails, web pages, documents) as untrusted. It can contain instructions meant to hijack the agent, so never let it widen the agent's permissions.
- Log every tool call and decision so you can audit and replay what happened.
- Set spend and step limits, and stop the loop when they are reached.
- Keep a test set of real examples with known good answers and run it on every change.
Measure what matters
| Measure | Why it matters |
|---|---|
| Task success rate on your test set | Shows whether a change helped or hurt |
| Human override rate | High overrides mean the agent is not ready for more autonomy |
| Cost and time per task | Compare against the manual process, including review time |
| Incidents and near misses | Tells you whether the guardrails are real |
