Skip to content
AtheronLABS

You're visiting from the United States. Prices are shown in US dollars. Not right?

labs@atheron:~/insights/agent-guardrails$ agent run --tools read-only --max-steps 12 --approve writes

Limits for an AI agent

Tools, budgets and evaluation.

An agent is a language model that can take actions: look things up, fill in forms, send messages, change records. That is what makes it useful, and what makes it risky. You do not fix that by trusting the model more. You design the limits around it, and this guide shows you how.

AI · Published October 2, 2026 · 7 min read

Why an agent needs limits

A chat assistant that answers questions can be wrong, which is a problem. An agent that acts can be wrong and do something about it: refund the wrong order, email the wrong customer, delete the wrong file. Language models are also easy to mislead, either by the person using them or by text they read along the way.

OWASP's Top 10 for LLM Applications names the risks that matter most here, including prompt injection, excessive agency and unbounded consumption [1]. Each one has a practical answer, and none of the answers depend on the model behaving perfectly. They are the same kinds of control you would put around a new employee with access to real systems: limited access, supervision on important decisions, a spending limit and a record of what they did.

Give it the fewest tools, with the least access

An agent's tools are the functions it can call: search the knowledge base, look up an order, create a ticket, issue a refund. OWASP describes excessive agency as a system given more functionality, more permissions or more autonomy than its task needs, and recommends limiting both the tools an agent may call and what each tool can do to the minimum necessary [2].

  • Narrow tools beat general ones. A tool that looks up one order by number is safer than one that runs any database query.
  • Read before write. Start an agent with read-only tools; add the ability to change things one action at a time, as it earns trust.
  • Act as the user, not as an administrator. OWASP recommends running tools in the user's own context rather than through a privileged account, so the agent can never do more than the person it is helping could [2].
  • Check permissions in the system being called, not only in the prompt. The downstream system should refuse an action the user is not allowed, whatever the model asked for [2].

Treat everything it reads as untrusted

Prompt injection is text that tries to change what the model does. It can come directly from the user, or indirectly from content the model reads, such as a web page, an email or a file [3]. An agent that reads a customer's email and also has a refund tool can be told, by that email, to issue a refund.

  • Keep instructions and data apart, and mark outside content clearly as untrusted, as OWASP recommends [3].
  • Define the output format and check it with ordinary code, not with another model, before acting on it [3].
  • Never let content the agent reads raise its own permissions. A document cannot grant access; only your systems can.
  • Test with adversarial inputs before launch and regularly after, the same way you would test any other security control [3].

Put people at the right checkpoints

Not every action needs a person, and asking for approval on everything teaches people to click yes without reading. Sort actions by what a mistake would cost, and require approval only where it matters. OWASP recommends human approval for high-impact actions [2].

An example of approval by risk
ActionRisk if wrongControl
Search documents, look up an orderLow: a wrong answer, visible to the userNone beyond permissions and logging
Prepare a reply or fill in a form for reviewLow: a person sees it before it goes anywhereShown to the user to edit and send
Create a ticket or update a non-financial recordMedium: a messy record to correctAllowed, logged, reversible
Send an email to a customerMedium to high: cannot be unsentApproval by the user, or limited to templates
Refund, pay, delete, change permissionsHigh: money or data lostApproval by an authorised person every time

Budgets: steps, time and spend

An agent works in a loop: think, call a tool, read the result, think again. Without limits, a confused agent can loop for a long time, and every turn uses tokens or GPU time. OWASP lists unbounded consumption as a risk in its own right, including attackers running up a pay-per-use bill on purpose, and recommends rate limits, quotas, timeouts and monitoring [4].

  • A maximum number of steps per task. When it is reached, the agent stops and hands over to a person with what it has so far.
  • A time limit per task and per tool call, so one slow system does not hold everything up.
  • A spend limit per task, per user and per day, enforced by your code before each model call.
  • Limits on input size, so nobody can paste a library into the chat box.
  • Alerts when usage departs from normal, sent to someone who can act on them.

Budgets also make the running cost predictable. The examples below show the monthly model cost at a given level of use; limits are what keep it there.

Evaluation before and after launch

You cannot tell whether an agent is safe and useful by trying it a few times. You need a set of realistic tasks with known good outcomes, run automatically, before launch and before every change. NIST's AI Risk Management Framework, a voluntary framework for building trustworthiness into the design, development, use and evaluation of AI systems, is a useful structure for deciding what to measure and who is accountable [5].

  1. 01

    Collect real tasks

    Questions and requests from the people who will use the agent, with the right outcome for each, including tasks it should refuse or hand over.

  2. 02

    Add adversarial cases

    Injected instructions in documents, requests for things the user is not allowed, and attempts to make it loop.

  3. 03

    Score more than the answer

    Did it use the right tools, in the right order, within budget, and stop when it should have?

  4. 04

    Set a bar and hold it

    Agree the score it must reach to go live. A change that drops below it does not ship.

  5. 05

    Keep evaluating in production

    Sample real conversations, review them, and add the failures to the test set.

Logs, monitoring and an off switch

  • Log every step: what the agent was asked, which tools it called with which inputs, what came back and what it did. Mask personal information in the logs as you would anywhere else.
  • Make every action traceable to a user and a task, so an odd change in a record can be explained.
  • Watch error rates, handovers, refusals and spend, and review a sample of conversations each week.
  • Have an off switch that does not need a deployment: one setting that disables the agent, or a single tool, at once.
  • Write down who is accountable for the agent's behaviour, and who decides when to switch it off.

Start narrow, widen with evidence

The safest agents start small. Launch with one well-defined job, read-only tools and a handful of trusted users. Watch the logs and the evaluation scores for a few weeks. Then widen one thing at a time: more users, then a tool that writes, then fewer approvals on actions that have proved safe. Each step is a decision made on evidence, and each can be rolled back.

This is slower than switching everything on at once, and it is how agents end up trusted rather than switched off after the first incident. It also gives the people who work alongside the agent time to learn what it is good at, and to tell you where it is not.

Two agents, priced from our rate card

The worked examples below show the range from our rate card today, with the build and the monthly running cost. The build includes the tools, the permissions, the approval steps, the budgets and the evaluation set; they are part of the agent, not extras.

Worked example, priced now

An operations agent on a frontier model

An internal agent that looks up orders and customers, creates tickets and prepares refunds for approval, connected to existing systems, on a frontier model through its API.

Build
≈ US$70,200 to US$107,000, delivered within 19 weeksCAD 99,900 to 152,800

Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.

Running it

Hosting
≈ US$558 a monthCAD 795 a month
Support
≈ US$2,500 a monthCAD 3,565 a month
Model running cost
≈ US$383 a monthCAD 545 a month

Worked example, priced now

An agent for regulated data on an open-source model

An agent that searches regulated documents and acts in internal systems, on an open-weight model served on rented GPUs in your own cloud account, with priority support.

Build
≈ US$115,000 to US$175,000, delivered within 23 weeksCAD 163,200 to 249,400

Prices in your currency are estimates from today's Bank of Canada rate. All invoicing is in CAD or USD.

Running it

Hosting
≈ US$7,020 a monthCAD 10,000 a month
Support
≈ US$5,010 a monthCAD 7,135 a month
Model running cost
≈ US$6,530 a monthCAD 9,300 a month

A checklist before an agent goes live

  1. 01

    Tools listed and narrowed

    Every tool named, with what it can read and change, and nothing more.

  2. 02

    Permissions enforced downstream

    The systems the agent calls check the user's rights themselves.

  3. 03

    Approvals for high-impact actions

    Money, deletions, permissions and outbound messages need a person.

  4. 04

    Budgets set and enforced in code

    Steps, time, spend per task and per day, and input size.

  5. 05

    Evaluation passing

    Real and adversarial tasks, scored, at or above the agreed bar.

  6. 06

    Logs, alerts and an off switch tested

    Someone has switched it off and on again, on purpose, before launch.

// sources

Where the figures come from.

Every statistic in this guide links here. Prices come from our estimator, not from a source.

  1. [1]OWASP GenAI Security Project, Top 10 for LLM Applications, 2025. genai.owasp.org/llm-top-10/
  2. [2]OWASP GenAI Security Project, LLM06:2025 Excessive Agency. genai.owasp.org/llmrisk/llm062025-excessive-agency/
  3. [3]OWASP GenAI Security Project, LLM01:2025 Prompt Injection. genai.owasp.org/llmrisk/llm01-prompt-injection/
  4. [4]OWASP GenAI Security Project, LLM10:2025 Unbounded Consumption. genai.owasp.org/llmrisk/llm102025-unbounded-consumption/
  5. [5]NIST, AI Risk Management Framework, 2023. www.nist.gov/itl/ai-risk-management-framework

// questions

Short answers.

Can an agent be fully autonomous?

For low-risk, reversible actions inside tight limits, yes. For anything involving money, deletions or messages to customers, we recommend a person approves, at least until months of logs show it is safe to relax.

Does a better model remove the need for limits?

No. Better models make fewer mistakes, but they can still be misled by what they read. Limits are what make the remaining mistakes harmless.

Who is responsible when an agent makes a mistake?

The organisation that deployed it. That is why every action should be logged, traceable to a user and a task, and reversible where possible. Your advisers can tell you what applies in your sector.

// next

Have a project in mind?

Put it through the estimator and get a range in a few minutes. Or tell us about it, and we will come back to you with questions.

AI agent guardrails: tools, permissions, budgets, evaluation | Atheron Network Labs