A prototype agent runs as its developer. It inherits that person’s access, which is usually broad, and it works. The demo is convincing, the business case is written, and then security asks a question that has no good answer: in production, who is this agent, and what can it reach?
This is where most agent programmes lose two quarters. The problem is not that the question is hard to answer. It is that answering it properly requires design work that nobody scoped, because the prototype made the hard part look solved.
An agent is a new kind of principal
Existing access models assume two kinds of actor. Humans are slow, accountable, and subject to review. Services are fast, narrow, and deterministic. An agent is neither. It is fast, broad, and non-deterministic, and it takes actions a human did not individually approve.
Granting it a human’s credentials makes attribution impossible and privilege excessive. Granting it a single service account across many workflows makes the blast radius the union of everything it has ever needed. Both are common, and both fail the first audit.
What the model actually needs
- An identity per agent and per workflow, not per team, so the audit trail attributes an action to a specific purpose
- Scopes derived from the workflow decomposition, so every permission traces back to a step that needs it
- Short-lived credentials with automatic rotation, because an agent credential leaked into a log or a prompt is a credential in untrusted hands
- Separate read and write paths, with write paths gated differently from read paths by default
- A budget and a rate limit attached to the identity, so cost is bounded by design rather than by alerting
- Revocation that takes effect within a run, not at the next token refresh
Untrusted content is the hard case
The failure mode that makes permission design non-negotiable is prompt injection. An agent that reads a web page, a support ticket, or an email is reading text an attacker may control. That text can attempt to redirect the agent toward its tools.
No prompt-level defence removes this risk. The durable mitigation is architectural: assume instructions in retrieved content may be hostile, and ensure the agent’s permissions make the worst plausible instruction survivable. If an injected instruction could cause a customer refund, an outbound email, or a data export, the permission model is doing the security work, and it has to be designed accordingly.
A useful test: write down the single most damaging action your agent could take if an attacker fully controlled the content it reads. If that action is possible, the guardrail is missing, not weak.
Where the human belongs
Human-in-the-loop is often added uniformly, which trains reviewers to approve everything. Approval is a scarce resource and should be spent where reversal is expensive.
- 01Cheap and reversible actions run unattended, with logging.
- 02Expensive but reversible actions run unattended, with a notification and an easy undo.
- 03Irreversible actions, or actions visible to a customer, require approval before execution.
- 04Actions outside the agent’s declared scope fail closed and escalate. They are never silently retried.
Design it before the framework
Teams usually choose an agent framework first and discover the permission model second. Reversing that order costs a week and saves the rework. Decompose the workflow into tools, state, and decision points. Derive the scopes. Decide where approval sits. Then choose the framework that supports what you designed.
The same decomposition produces your evaluation suite, because the decision points you identified are precisely what needs testing.