# Agentic authority boundaries

> Agentic authority boundaries are the controls that let an AI system that can act do so safely: tool grants are bounded, actions are staged propose-verify-execute, and the authority to take a consequential action is separate from the ability to recommend it.

Category: AI governance
Also searched as: AI agents, excessive agency, prompt injection, tool use
Source: Mediator Solutions — https://mediatorsolutions.io/learn/#agentic-authority-boundaries
License: free to read, learn, cite, and apply, with attribution to Mediator Solutions.

## What it is

A system that generates text can be wrong on a page; a system that can act can be wrong in the world. Agentic authority boundaries treat an acting model like an actuator: it proposes, the proposal is verified, and only then is it executed, in that order and not the reverse. Tool grants are scoped, untrusted input is treated as data rather than instruction, and excessive agency is designed out rather than hoped away.

## Why it matters

A model that writes text can be wrong on a page; a model that can act can be wrong in the world, irreversibly and at machine speed. The boundary treats an acting model like an actuator in any safety-critical system: it proposes, the proposal is checked, and only then does it execute — never the reverse, because execute-then-check has no undo. The two failure modes to design against are prompt injection, where untrusted input is read as instruction, and excessive agency, where broad goals plus tool access plus weak supervision plus no clear stop condition compound into a system that can do far more than anyone intended.

## When to use it

- Designing or reviewing any system where a model can take actions, not just produce text.
- A model is given tools, broad goals, and weak supervision at the same time.
- Untrusted input (a document, a web page, a user message) could be read as an instruction.

## Principles

- Stage actions propose, verify, then execute — not execute then check.
- Bound every tool grant to a defined scope; broad agency is a liability, not a feature.
- Treat external content as untrusted data, not as instructions (guard against prompt injection).
- Language outputs are proposals, not truth; a model advises and a governed path decides.

## Practice

1. List what the agent can actually do, and scope each capability explicitly.
2. Insert a verification step between an agent’s proposal and any consequential action.
3. Isolate untrusted input so it cannot redirect the agent’s authority.
4. Route binding actions through the measurement-is-not-authority boundary.

## Where it fails

- **Prompt injection** — Untrusted content is interpreted as instruction, so an attacker steers the agent by planting text in data it reads.
- **Excessive agency** — Broad goals plus tool access plus weak oversight plus no clear stop condition compound into a system that can do far more than intended.
- **Execute-then-check** — The action runs before it is validated, so there is no undo for the irreversible mistake that validation would have caught.

## In practice

An agent that can send payments is asked to settle invoices from an inbox. The unsafe design lets it read, decide, and pay in one motion — an injected invoice becomes a real transfer. The bounded design makes the agent propose a payment, routes the proposal through a validation that checks the invoice against known vendors and limits, and only then executes, with a human approving anything outside the envelope. The model proposes; the boundary disposes; nothing irreversible happens before the check.

## Verification

For each agent capability, the record shows a scoped grant, a verification step before execution, and a separate authorization for any consequential act.

## Reference

### Agent risk classes

- Class 0 — text only: draft or summarize, no tools, no external action
- Class 1 — retrieval: reads approved sources; risk is leakage, misquotation, stale information
- Class 2 — tool assisted: calls tools but changes no systems; risk is tool misuse, data exposure, false confidence
- Class 3 — execution capable: writes files, runs commands, deploys, transacts; requires approval, rollback, receipt, allowlist, sandbox, review
- Class 4 — autonomous external action: acts externally without immediate approval; default status prohibited unless explicitly authorized

### Deterministic controls for probabilistic outputs

- Scope the allowed task
- Define allowed tools
- Restrict access by least privilege
- Log prompts, sources, tool calls, and outputs
- Validate outputs before operational use
- Require human approval for external action
- Store receipts
- Monitor drift
- Set cost and rate limits
- Maintain a stop condition

---

Previous: https://mediatorsolutions.io/learn/#drift-register
Next: https://mediatorsolutions.io/learn/#secure-by-design
