Why an agent needs limits
A chat assistant that answers questions can be wrong, which is a problem. An agent that acts can be wrong and do something about it: refund the wrong order, email the wrong customer, delete the wrong file. Language models are also easy to mislead, either by the person using them or by text they read along the way.
OWASP's Top 10 for LLM Applications names the risks that matter most here, including prompt injection, excessive agency and unbounded consumption [1]. Each one has a practical answer, and none of the answers depend on the model behaving perfectly. They are the same kinds of control you would put around a new employee with access to real systems: limited access, supervision on important decisions, a spending limit and a record of what they did.
Give it the fewest tools, with the least access
An agent's tools are the functions it can call: search the knowledge base, look up an order, create a ticket, issue a refund. OWASP describes excessive agency as a system given more functionality, more permissions or more autonomy than its task needs, and recommends limiting both the tools an agent may call and what each tool can do to the minimum necessary [2].
- Narrow tools beat general ones. A tool that looks up one order by number is safer than one that runs any database query.
- Read before write. Start an agent with read-only tools; add the ability to change things one action at a time, as it earns trust.
- Act as the user, not as an administrator. OWASP recommends running tools in the user's own context rather than through a privileged account, so the agent can never do more than the person it is helping could [2].
- Check permissions in the system being called, not only in the prompt. The downstream system should refuse an action the user is not allowed, whatever the model asked for [2].
Treat everything it reads as untrusted
Prompt injection is text that tries to change what the model does. It can come directly from the user, or indirectly from content the model reads, such as a web page, an email or a file [3]. An agent that reads a customer's email and also has a refund tool can be told, by that email, to issue a refund.
- Keep instructions and data apart, and mark outside content clearly as untrusted, as OWASP recommends [3].
- Define the output format and check it with ordinary code, not with another model, before acting on it [3].
- Never let content the agent reads raise its own permissions. A document cannot grant access; only your systems can.
- Test with adversarial inputs before launch and regularly after, the same way you would test any other security control [3].
Put people at the right checkpoints
Not every action needs a person, and asking for approval on everything teaches people to click yes without reading. Sort actions by what a mistake would cost, and require approval only where it matters. OWASP recommends human approval for high-impact actions [2].
| Action | Risk if wrong | Control |
|---|---|---|
| Search documents, look up an order | Low: a wrong answer, visible to the user | None beyond permissions and logging |
| Prepare a reply or fill in a form for review | Low: a person sees it before it goes anywhere | Shown to the user to edit and send |
| Create a ticket or update a non-financial record | Medium: a messy record to correct | Allowed, logged, reversible |
| Send an email to a customer | Medium to high: cannot be unsent | Approval by the user, or limited to templates |
| Refund, pay, delete, change permissions | High: money or data lost | Approval by an authorised person every time |
Budgets: steps, time and spend
An agent works in a loop: think, call a tool, read the result, think again. Without limits, a confused agent can loop for a long time, and every turn uses tokens or GPU time. OWASP lists unbounded consumption as a risk in its own right, including attackers running up a pay-per-use bill on purpose, and recommends rate limits, quotas, timeouts and monitoring [4].
- A maximum number of steps per task. When it is reached, the agent stops and hands over to a person with what it has so far.
- A time limit per task and per tool call, so one slow system does not hold everything up.
- A spend limit per task, per user and per day, enforced by your code before each model call.
- Limits on input size, so nobody can paste a library into the chat box.
- Alerts when usage departs from normal, sent to someone who can act on them.
Budgets also make the running cost predictable. The examples below show the monthly model cost at a given level of use; limits are what keep it there.
Evaluation before and after launch
You cannot tell whether an agent is safe and useful by trying it a few times. You need a set of realistic tasks with known good outcomes, run automatically, before launch and before every change. NIST's AI Risk Management Framework, a voluntary framework for building trustworthiness into the design, development, use and evaluation of AI systems, is a useful structure for deciding what to measure and who is accountable [5].
- 01
Collect real tasks
Questions and requests from the people who will use the agent, with the right outcome for each, including tasks it should refuse or hand over.
- 02
Add adversarial cases
Injected instructions in documents, requests for things the user is not allowed, and attempts to make it loop.
- 03
Score more than the answer
Did it use the right tools, in the right order, within budget, and stop when it should have?
- 04
Set a bar and hold it
Agree the score it must reach to go live. A change that drops below it does not ship.
- 05
Keep evaluating in production
Sample real conversations, review them, and add the failures to the test set.
Logs, monitoring and an off switch
- Log every step: what the agent was asked, which tools it called with which inputs, what came back and what it did. Mask personal information in the logs as you would anywhere else.
- Make every action traceable to a user and a task, so an odd change in a record can be explained.
- Watch error rates, handovers, refusals and spend, and review a sample of conversations each week.
- Have an off switch that does not need a deployment: one setting that disables the agent, or a single tool, at once.
- Write down who is accountable for the agent's behaviour, and who decides when to switch it off.
Start narrow, widen with evidence
The safest agents start small. Launch with one well-defined job, read-only tools and a handful of trusted users. Watch the logs and the evaluation scores for a few weeks. Then widen one thing at a time: more users, then a tool that writes, then fewer approvals on actions that have proved safe. Each step is a decision made on evidence, and each can be rolled back.
This is slower than switching everything on at once, and it is how agents end up trusted rather than switched off after the first incident. It also gives the people who work alongside the agent time to learn what it is good at, and to tell you where it is not.
Two agents, priced from our rate card
The worked examples below show the range from our rate card today, with the build and the monthly running cost. The build includes the tools, the permissions, the approval steps, the budgets and the evaluation set; they are part of the agent, not extras.
Worked example, priced now
An operations agent on a frontier model
An internal agent that looks up orders and customers, creates tickets and prepares refunds for approval, connected to existing systems, on a frontier model through its API.
- Build
- ≈ US$70,200 to US$107,000, delivered within 19 weeksCAD 99,900 to 152,800
Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.
Running it
- Hosting
- ≈ US$558 a monthCAD 795 a month
- Support
- ≈ US$2,500 a monthCAD 3,565 a month
- Model running cost
- ≈ US$383 a monthCAD 545 a month
Worked example, priced now
An agent for regulated data on an open-source model
An agent that searches regulated documents and acts in internal systems, on an open-weight model served on rented GPUs in your own cloud account, with priority support.
- Build
- ≈ US$115,000 to US$175,000, delivered within 23 weeksCAD 163,200 to 249,400
Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.
Running it
- Hosting
- ≈ US$7,020 a monthCAD 10,000 a month
- Support
- ≈ US$5,010 a monthCAD 7,135 a month
- Model running cost
- ≈ US$6,530 a monthCAD 9,300 a month
A checklist before an agent goes live
- 01
Tools listed and narrowed
Every tool named, with what it can read and change, and nothing more.
- 02
Permissions enforced downstream
The systems the agent calls check the user's rights themselves.
- 03
Approvals for high-impact actions
Money, deletions, permissions and outbound messages need a person.
- 04
Budgets set and enforced in code
Steps, time, spend per task and per day, and input size.
- 05
Evaluation passing
Real and adversarial tasks, scored, at or above the agreed bar.
- 06
Logs, alerts and an off switch tested
Someone has switched it off and on again, on purpose, before launch.