Spurlock Studios
Contact
Share LinkedIn X
A cracked amber fuse. Thesis: SET POLICIES AGENTS CANNOT.

You set policies for what agents can and cannot do by writing a versioned catalog with a named owner: every tool proposal is allow, deny, or pending-approval before the tool runs, and the catalog fails closed if it cannot load. Prompts explain tone. IAM holds keys. Counsel writes contracts. None of those three is the agent policy. This post is production ops for side effects — not legal advice.

This spoke sits under the Agentic Systems Operating Manual. It owns the rules humans author and the person who can change them. The on-ramp for a control-plane pilot is /agentic.

The short answer

  • Interview the business for side effects, then encode each as a predicate → allow / deny / pending-approval.
  • Default is deny. Unknown tool, failed schema, missing owner, and catalog outage are deny for writes.
  • Put a named human on the catalog: change authority, backup, review date. A team name is not an owner.
  • Bind pending-approval to a payload hash. A later arg change invalidates the yes.
  • Do not treat this catalog as a privacy program, an employment policy, or a substitute for an attorney.

What is an agent policy if it is not a prompt?

An agent policy is a machine-checked table that answers one question per proposed tool call: may this principal run this tool with these args, right now? If the table has no row, the answer is no.

OWASP LLM07:2025 System Prompt Leakage is blunt: do not put authorization in the system prompt. OWASP LLM01:2025 Prompt Injection exists because instructions and untrusted data share one channel. A sentence that says “never refund over $50” is a hope. A row that says amount_usd > 50 → pending-approval is a policy.

ArtifactWhat it can doWhat it cannot do
System promptTone, format, “please ask”Stop a tool the runtime still executes
Legal / HR PDFContracts, employment, privacy noticesInspect {"amount": 4900} at call time
IAM / OAuth scopesWhich keys existExpress “over $50 needs Maya”
Vendor topic filterBlock classes of textRead your refund cap or recipient allowlist
Policy catalogallow / deny / pending-approval on argsReplace counsel, IAM, or a sandbox

OpenAI’s function calling guide is explicit: the API returns a proposed call; your application executes it. Anthropic’s tool-use loop says the same for client tools. The catalog lives in that gap. If you skipped the gap and “set policy” by editing a prompt, you did not set policy.

A catalog row that cannot name its owner is a rumor with YAML syntax.

How do you write allow, deny, and pending-approval?

You do not start from a philosophy deck. You start from side effects the business already fears, then you force every fear into one of three decisions.

allow

Use when the tool is on the principal’s allowlist, the args pass a strict schema, the side-effect class is permitted at this autonomy level, and no freeze is on. Allow is a positive grant, not “we forgot to list it.”

deny

Use for unknown tools, schema failures, disallowed destinations, amounts over a hard cap, tenant mismatches, and any time the catalog cannot be evaluated. Deny is the default. Write it down so on-call does not treat it as a bug.

pending-approval

Use when the args are otherwise valid but the action is irreversible, over a soft cap, novel (new domain, new payee), or the autonomy level is “draft + human.” Approval binds to a hash of normalized args. If the model edits the payload, prior approval is void.

Encoding procedure:

  1. List every tool the agent can see. Delete any the job does not need (OWASP LLM06:2025 Excessive Agency — too much functionality, too many permissions, too much autonomy).
  2. Tag each tool read / write_reversible / write_irreversible / exfil_risk at registration, not by guessing from the name.
  3. Sit with the operator who already approves the human version of this action. Ask four questions: who may, who may not, who must approve, what is the number.
  4. Turn each answer into a predicate on args (amount_usd, recipient_host, table, op).
  5. Map predicate → allow | deny | pending-approval. No fourth outcome named “model’s best judgment.”
  6. Set default deny. Publish policy_id, version, owner, backup, review_by.
Business sentencePredicateDecision
“Support can refund small tickets”orders.refund AND amount_usd <= 50 AND same tenantallow
“Bigger refunds need a human”amount_usd > 50 AND amount_usd <= 500pending-approval
“We never refund over $500 from the bot”amount_usd > 500deny
“Only template email to our domain”email.send AND host allowlisted AND template_id setallow
“Free-text to a new domain”body free-text OR host unknownpending or deny
“No mass CRM rewrite”op in delete/truncate OR missing object iddeny

If two people disagree on the number, you do not have a policy yet. You have an argument. Freeze the tool until they pick a number.

Pack header you can actually load (illustrative — not a vendor schema):

FieldExampleLoad rule
policy_pack_idsupport.refund.v7Required
version7Integer; cache key
owneremail of a humanRequired; not a group alias
backupsecond humanRequired
review_by2026-09-19Invalid pack if missing or past
defaultdenyWrites; do not omit
outage_writesdenyFail closed
outage_readsnamed read-tool list or denyNo “all GET”
job_typessupport.ticketScope
envprodStaging pack is a different file

Interview script for the operator who already does this job by hand:

  1. What is the last action you personally approved this month? Name the tool, not the vibe.
  2. What number would have made you say no without a meeting?
  3. Who is not allowed to do this even if they ask nicely?
  4. What destination, table, or payee is always wrong?
  5. If the policy service is down at 6pm, do we stop or do we improvise? Write the answer they actually give.

If they say “use your judgment,” you still do not have a predicate. Stay on deny until they pick a number or a named approver.

Where must those decisions fire?

Before the side effect leaves the process. After Stripe, after SMTP, after the row update is an incident report, not a policy.

OWASP LLM06 calls this complete mediation: every downstream request checked against policy. You set that requirement in the catalog (“no tool adapter may execute before a decision”) even though the interceptor is code. If any path reaches a client without a decision, that path is a bypass — treat it like a missing rule, not a clever shortcut.

Side effectCatalog must constrainTypical pending trigger
Refund / payoutAmount, currency, order id, tenantOver soft cap
Outbound email / SlackRecipients, template id, attachment refsNew host; fan-out over N
CRM / DB writeTable, object id, field allowlist, opProduction table; delete
Calendar createCalendar id, attendee domainExternal attendees
HTTP fetchURL allowlist, method, sizeNon-allowlisted host; POST as GET
Deploy / schemaTarget, diff sizeAny production mutate

Checklist before you call the catalog “set”:

  • Every registered tool has a side-effect class
  • Every write tool has at least one predicate row
  • Extra arg fields are rejected (additionalProperties: false)
  • Money fields require currency + max
  • Recipients resolve against an allowlist you maintain, not a regex the model wrote
  • “Update all” / missing object id is deny
  • Kill-switch and budget freeze are deny predicates the catalog names

Reads can be allowlisted. Writes earn the three-way decision. Irreversible writes default to pending or off until the owner promotes them with evidence, not with a pep talk in the prompt.

Parallel proposals are still one catalog. OpenAI will return multiple tool_calls in one turn. Your rule is: evaluate each call, in order, and stop executing the rest after a deny unless the owner wrote an explicit “continue siblings” row (almost nobody should). A deny that “rides along” with an allow is a bypass you authored by accident.

EnvironmentPackTypical loosen
Local / CIfixture pack, default denyNone — tests assert decisions
Stagingcopy of prod predicatesPending may auto-approve to a sink, never to prod credentials
Prodsigned pack, named ownerOnly via the slow loosen path

Staging credentials plus a prod pack is how “it worked in staging” becomes a refund. The catalog is not portable across secret sets. Name the env in the header.

Who owns the catalog, and what does ownership mean?

Someone with a name, a calendar, and the authority to change a row. Not “engineering.” Not “legal signed off once.” Not #agent-approvals with no on-call rotation.

NIST AI RMF 1.0 and the July 2024 Generative AI Profile (NIST AI 600-1) ask you to define roles for human–AI configurations and to refuse work the system should not run. A named catalog owner is that role in ops language. I am not mapping your company onto a NIST control ID. I am saying: if you cannot name the human, you have not implemented the role.

RoleDoesDoes not
Catalog ownerApproves row changes; sets review_by; is pageable for freezeWrite every predicate on day one
Backup ownerSame rights when owner is outA “FYI” Slack mention
Job operatorSupplies the four questions (who may / may not / approves / number)Merge YAML without owner ack
Runtime / platformEnforces the packed catalog; fail closed on load failureInvent business caps
CounselContracts, privacy notices, regulated claimsAuthor amount_usd > 50 as a substitute for the catalog
ModelProposes tool + argsDecide

RACI you can actually run:

  • Responsible: one owner per job_type catalog (or per tool class if you split packs)
  • Accountable: the person whose budget eats a bad refund / send
  • Consulted: operator who does the human version of the task today
  • Informed: on-call, who must know that a red policy dependency means writes stop

Owner checklist:

  • Name and backup in the pack header, not in a wiki nobody deploys
  • Change path: PR or config review that the owner must ack
  • Freeze path: owner or incident commander can deny-all writes without a prompt edit
  • Review date in the future, not “we’ll revisit”
  • Handoff note when the owner leaves the company — catalogs do not inherit by osmosis

If the owner is “the founder,” write that down and put a backup anyway. Founders take flights.

Solo-operator shape (this studio’s default): you are owner. Name a backup who can freeze writes even if they cannot author predicates — incident commander, contractor, whoever can reach the flag. A one-person company still needs a freeze path that does not require editing a prompt from a phone.

Team shape: one owner per job_type pack. Payments owns refunds. Support owns ticket sends. Platform owns the loader and the fail-closed contract. If both payments and support can change the same row, you do not have an owner. You have a wiki.

Handoff when the owner leaves:

  1. New owner acks the current version in the pack header.
  2. Backup confirms they still want the pager.
  3. Freeze path is tested with the new people, not assumed.
  4. Old owner’s credentials are removed from the approval queue.

An owner who cannot be paged is a caption.

How do you fail closed without stalling the business?

You write the outage mode as a row, you sign it, and you drill it. Fail closed means: if the catalog cannot be evaluated, writes do not run. It does not mean the company halts. Reads may continue only if you authored an explicit outage allowlist for pre-classified read tools.

CISA’s joint work on AI in operational technology tells operators to keep a human in the loop on critical decisions and to implement fail-safe limits on worst-case harm. The April 2024 Deploying AI Systems Securely CSA (CISA, NSA, FBI, and partners) uses the same shape: human-in-the-loop as a failsafe, rollbacks ready. Your support bot is not a turbine. The outage posture still applies: empty safety channel, no writes.

FailureCatalog saysNot allowed
Pack timeout / 5xxdeny writes; optional pending queue with expiry → deny“Let it through, we’ll audit”
Rule pack failed to loaddenyLast-known-good from a laptop cache you cannot hash
Schema fail / unknown tooldenyAlias guess (refund_v2)
Approval service downdo not allow irreversible; abort or wait then denyConvert pending to allow
Owner on PTO, no backupfreeze writes for that job_typeIntern merges a cap change

Drill, because unsigned fail-closed is a slide:

  1. Break the pack on purpose in staging (empty file, bad signature, 30s timeout).
  2. Confirm every write tool returns policy_unavailable and does not execute.
  3. Confirm the read exception, if any, is the one you wrote — not “all GETs.”
  4. Time how long on-call takes to name the owner and flip the freeze.
  5. Write the result in the runbook next to the catalog version.

Fail open during an outage is how a refund agent empties the till while the policy repo is “having a moment.” If that sentence makes finance nervous, good — they just became consulted.

Read exceptions are authored, never implied:

  • List the exact read tool names that may run when the pack is down
  • Ban any read that can concatenate into a send (mailbox-read plus a still-live send path)
  • Cap bytes and fan-out on those reads
  • Expire the exception (hours), not “until the repo is healthy”
  • Emit policy_unavailable even on allowed reads so the drill is visible

If you cannot list the read tools, the outage mode is deny-all. That is a valid business decision. Write it. Customers would rather wait than receive a public send composed from an ungoverned inbox. OWASP LLM06’s mailbox scenario is this failure with a bow on it: summarize, then a planted message talks the agent into forwarding secrets. An outage is not permission to grow a send path.

How do you handle exceptions and promotions?

Exceptions are how catalogs rot. Someone asks for “just this customer,” you paste an allow, and six months later the intern thinks that row is the product. Set the exception shape before the first plea.

RequestCatalog responseForbidden response
One-off over-cap refundHuman runs the native tool; agent stays pending/denyPermanent cap raise in the agent pack
New destination domainTime-boxed allowlist add with review_byRegex the model supplied
Demo tomorrowStaging pack + sink credentialsProd loosen for a recording
“VIP always allow”Separate principal with a documented pack, still cappedHidden prompt: “VIP means skip policy”
Model keeps getting deniedFix args or the jobLower the cap to make the graph green

Promotion from pending-approval to allow is a loosen change. Evidence, not vibes:

  1. Frozen golden set of payloads (including the injection cases you already lost to).
  2. N consecutive human approvals with zero arg edits on that tool + cap.
  3. Deny rate explained — falling because args got cleaner, not because someone fail-opened.
  4. Trace coverage still 100% of tool spans.
  5. Owner and budget holder both ack a version bump.
EvidenceNot evidence
Approvals without editsA longer system prompt
Payload hashes matching at executeA Slack thumbs-up on a screenshot
Drill of pack-down → deny“We’ll watch it this week”
Kill-switch tested this monthA channel named #approvals with no hash

If promotion evidence is missing, the tool stays pending. Unstaffed pending stays deny. You are not cruel. You are refusing to launder a missing owner through a confidence score.

What belongs in the catalog versus IAM, spend, and traces?

Split the layers or you will jam every fear into one YAML file and then stop updating it.

ConcernHomeCatalog’s job
Which keys existIAM / secretsDo not store secrets in rules
Tenant isolation at the providerIAMMirror tenant checks on args anyway
“Refunds over $50 need a human”CatalogPredicate + pending
Token / USD spirals, revision capsCost controls for agent fleetsTreat budget freeze as a deny predicate the pack names
Did the decision fire?Observability for agentsEmit policy_id, version, decision, reason, payload hash
Network blast radiusSandbox / egress allowlistMention exfil_risk; do not pretend YAML is a jail
Contracts, DPIAs, employmentCounsel / privacyOut of scope for this catalog

Anthropic’s managed-agent permission policies use always_allow / always_ask for their managed tools. Custom tools are still your catalog. Vendor vocabulary is not a substitute for your rows.

A spend kill switch without a catalog deny is a finance backstop. A catalog deny without a budget counter still lets a loop burn tokens on allowed reads. You want both. After 500+ automations and 20,000+ hours on agentic systems, the expensive miss is usually the layer nobody named — not the lack of another adjective in the prompt.

Keep the catalog small enough to review. Caps and traces have their own spokes. Do not copy those tables here and call it governance.

How do you version, review, and sunset a rule?

Treat the pack like production config. A hot cache that serves last week’s cap is not a policy. Cost-control caches already need policy_version in the key; your catalog is that version.

Change typeHow it shipsWho acks
Tighten (lower cap, extra deny)Fast path; still version bumpOwner or incident commander
Loosen (raise cap, new allow)Slow path; golden-set + deny-rate checkOwner and accountable budget holder
New toolDeny by default until a row existsOwner
Sunset a toolRemove from allowlist; keep deny traces with the rest of the incident windowOwner
Owner transferHeader change + backup confirmOld owner, new owner, backup

Review procedure (cadence you can keep — pick one and calendar it):

  1. Export decisions grouped by policy_id for the window (allows, denies, pending, policy_unavailable).
  2. Open the ten loudest pending items. If humans always click yes with zero edits, the cap is theater — tighten or allow under a lower blast radius.
  3. Open the ten loudest denies. If operators keep asking for a bypass, you have a missing row or a job that should not be an agent.
  4. Confirm every tool in the registry still has a row. Friday’s new MCP wrapper is how allow-by-default returns through the side door.
  5. Move review_by forward only if the owner acks. Expired review → freeze loosens, not “hope.”

Sunset rules you should write on day one:

  • Caps without a review_by are invalid at load time (fail closed)
  • Tools unused for N days stay deny until re-justified
  • Approvals expire (minutes, not “this sprint”)
  • A loosen change cannot ride along in a “docs only” PR

Version is an integer in the pack, in the trace, and in the cache key. “Latest” is not a version.

What breaks when the policy has no named owner?

What breaks: a Notion page titled “Agent Guardrails,” a prompt paragraph, and a legal acceptable-use PDF. None of them sit in front of orders.refund. An intern adds slack.post_public on Friday. Nobody is pageable. Monday a send goes to a lookalike domain. The incident channel asks “who approved this tool?” and the answer is a shrug.

What it costs: I will not invent a dollar figure you can take to finance. The loss is every side effect the tool actually posted, plus the hours to reconstruct which ones were real, plus the freeze you now apply in panic. Prompt-only agents fail this way on a schedule. After 500+ automations, the report that keeps repeating is “we thought the prompt covered it,” not “the catalog owner denied it.”

What you do instead:

  1. Stop writing new tools until a pack header has owner + backup + default deny.
  2. Register the three scary classes first: refund/payout, outbound send, production write.
  3. Encode the operator’s existing approval number as pending/deny. Do not invent a tighter cap than the humans already use — you will get shadow bypasses.
  4. Fail closed on load failure. Drill it once.
  5. Put policy_decision on every tool span before you argue about dashboards.

Secondary failure modes that look like ownership and are not:

CounterfeitWhy it fails
Legal PDF as the catalogNo arg predicates; no fail-closed loader
Slack channel as ownerNo ack, no freeze, no version
“The model will ask”Injection and non-determinism
Allow-by-default plus a deny listYou will miss the new alias
Approvals without payload hashHumans approve vibes; agents send payload B
Owner = rotating internCaps loosen under pressure; nobody is accountable

Bravery is not a policy owner.

How do you measure whether the catalog is working?

You measure whether every tool span has a decision, whether humans still agree with pending, and whether bypasses exist. You do not measure “policy quality” with a thumbs-up on the prompt.

Wire the fields into the same timeline ops already reads — that screen is observability for agents, not a second graveyard CSV.

SignalHealthySick
Tool spans with policy_decision100% of executionsAny execute-without-decision
policy_unavailable in prod0 except scheduled drillsNonzero on ordinary hours
Pending wait (p50 / p95)Inside the SLA the owner publishedQueue nobody staffs
Approval-then-mutation rejectsHappen in tests; rare in prodZero because you do not hash
Deny rate by policy_idExplained; trending with job mixSudden drop with no pack change (fail open or bypass)
Loosen changesSlow, dual-ackedFriday night cap raises
Owner response on freezeMinutes“Who owns this?” in the incident doc

Evaluation order you should publish next to the metrics (so on-call knows what “working” means):

  1. Kill switch / budget freeze → deny
  2. Tool on principal allowlist → else deny
  3. Args match schema → else deny
  4. Pending predicates → pending-approval
  5. Deny predicates → deny
  6. Default allow only if the owner wrote a default-allow for that class (most packs should never)

Shadow mode is allowed as a measurement step: log what the catalog would have decided while the old path still runs. Shadow mode is not production policy. If you ship shadow and skip the interceptor, you measured a slideshow.

A catalog that never denies is not “high trust.” It is unused or fail-open. Check which.

When is a workflow enough instead of an agent?

When the can/cannot table is a flowchart with no mid-run tool choice. If every refund follows the same schema-checked steps, you want a workflow with a gate on the write node — not an agent that “decides” which tools exist.

You still need a catalog if…A workflow is enough if…
The model picks among tools mid-runTool sequence is fixed in the graph
Args are free-form and must be cappedArgs are form fields with the same caps
Novel recipients / payees appearDestinations are a static allowlist in the node
Autonomy might rise laterAutonomy is permanently “run this graph”

Decision list:

  1. Can you draw the happy path without a tool-choice step? If yes, start as a workflow.
  2. Does a human already follow a script with one number (cap, template, queue)? Encode that number in the write node. Do not hire a model to remember it.
  3. Are you adding an agent because the ticket text is messy, not because the action is messy? Classify with a model; execute with a workflow.
  4. If you still want an agent, the catalog is the promotion exam — not a consolation prize after the demo.

An agent whose entire policy is “always pending” is a slow workflow with extra tokens. Be honest and draw the graph.

Setting can/cannot policy does not obligate you to keep the agent. Sometimes the policy work reveals the job should never have been autonomous. That is a successful policy project.

What should you ship if you only have a week?

A v1 catalog that covers the irreversible tools you already have, with a named owner and a fail-closed load path. Not a platform. Not a 40-page acceptable-use rewrite. Spurlock Studios does not pretend a week is a governance program. It is enough to stop prompt-only side effects.

DayCatalog outputDone when
1Inventory tools; tag side effects; name owner + backupRegistry has classes; header has humans
2Default deny; allowlist per principalUnknown tool cannot execute in staging
3Predicates for refund / send / write; pending + payload hashOperator’s existing number is in a row
4Outage mode written; fail-closed drillpolicy_unavailable blocks writes
5Version 1 frozen; review_by on the calendar; traces onFreeze path works without a prompt edit

Skip if the week is small:

  • Multi-agent policy inheritance
  • Per-employee exception matrices
  • Mapping the pack onto ISO / SOC control IDs
  • Letting the model draft the predicates (it will allow itself)
  • A public “ethics” page as a substitute for rows
  • Loosening caps to make the demo prettier

Do not skip:

  • Named owner and backup
  • Three-way decision before tools
  • Fail closed on pack failure
  • Payload-bound pending
  • Decision fields on traces

If you cannot staff pending-approval, those tools stay deny. Unstaffed pending is fail-open with extra latency.

This catalog is how you stop an agent from running a tool you did not mean to allow. It is not legal advice. It does not make you compliant with a privacy law, a healthcare rule, an advertising standard, or a contract you have not read. If your agents touch regulated data or make statements that bind the company, talk to counsel. Do not paste this post into a DPIA and call it done.

People ask forWhat to give themWhat not to give them
“Set agent policy”Catalog + owner + fail closedA longer system prompt
“Are we allowed to do this?”Counsel + the actual contract / statuteA policy_id
“Make it ethical”Product decisions, disclosed to usersA deny list of adjectives
“Pass the audit”Traces that prove the gate firedA slide titled Guardrails
“Ship faster”Deny the scary tools until rows existShadow mode in production

Hard boundaries for this document:

  • No claim that a catalog replaces an attorney
  • No claim that OWASP, NIST, or CISA citations certify your deployment
  • No invented pass rates, dollar saves, or client outcomes
  • No prompt pack presented as enforcement
  • No “the model agreed to the policy” as evidence

Counsel can require human review for a class of actions. Your job is to turn that requirement into pending-approval on named tools with a hash, an expiry, and an owner. If counsel’s memo cannot be mapped onto a tool name and a predicate, it is not yet an agent policy — it is homework.

A prompt is not a policy. A policy without an owner is a rumor. A rumor that fails open is an incident.

FAQ

How do I set policies for what agents can and cannot do?

Write a versioned catalog with a named owner that returns allow, deny, or pending-approval before every tool runs, and fail closed if that catalog cannot load. Interview the operator who already approves the human version of the action, encode their numbers as predicates, and default-deny unknown tools. Prompts, PDFs, and IAM scopes are not a substitute for those rows.

How do I measure whether agent policies are working?

Require 100% of tool executions to carry a policy_decision, treat policy_unavailable in ordinary production as an incident, and watch pending wait times against the SLA the owner published. A sudden deny-rate drop with no pack change usually means a bypass or fail-open. Trace fields and the weekly screen live in observability; the catalog owner still owns the caps.

What usually fails first when teams try this?

Ownership. The pack is a wiki page, a prompt, or “legal signed something,” and Friday’s new tool ships with no row. Second is pending-approval with no one on the queue, which becomes allow under demo pressure. Third is fail-open on pack timeout. Fix the header (owner, backup, default deny) before you write clever predicates.

How long does this take to show results?

A week is enough to freeze irreversible tools behind a v1 catalog and prove fail-closed in staging. You will see results the first time a proposed refund or send dies at a row instead of in a chat log. Broader coverage — every job_type, reviewed caps, staffed pending SLAs — is a calendar, not a sprint slogan. If you cannot name an owner on day one, you are not starting the week.

What should I skip if I only have a week?

Skip control-framework mapping, per-user exception matrices, multi-agent inheritance, and letting the model draft its own allowlist. Do not skip owner + backup, default deny, the three scary tool classes, payload-bound pending, and a fail-closed drill. Unstaffed pending stays deny. A prettier demo is not a loosen change.

When is this not worth doing yet?

If you have no tools with side effects — no send, no write, no refund, no deploy — you do not have an agent policy problem yet; you have a chatbot. If the job is a fixed flowchart, ship a workflow with a gate on the write node instead of an agent. If nobody will own the catalog or staff pending, do not turn the model loose on production writes.

CTA

Put can/cannot in a named-owner catalog that decides before the tool runs, then fail closed — not in a paragraph the model can ignore.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

How do I set policies for what agents can and cannot do?
Write a versioned catalog with a named owner that returns `allow`, `deny`, or `pending-approval` before every tool runs, and fail closed if that catalog cannot load. Interview the operator who already approves the human version of the action, encode their numbers as predicates, and default-deny unknown tools. Prompts, PDFs, and IAM scopes are not a substitute for those rows.
How do I measure whether agent policies are working?
Require 100% of tool executions to carry a `policy_decision`, treat `policy_unavailable` in ordinary production as an incident, and watch pending wait times against the SLA the owner published. A sudden deny-rate drop with no pack change usually means a bypass or fail-open. Trace fields and the weekly screen live in observability; the catalog owner still owns the caps.
What usually fails first when teams try this?
Ownership. The pack is a wiki page, a prompt, or “legal signed something,” and Friday’s new tool ships with no row. Second is pending-approval with no one on the queue, which becomes allow under demo pressure. Third is fail-open on pack timeout. Fix the header (owner, backup, default deny) before you write clever predicates.
How long does this take to show results?
A week is enough to freeze irreversible tools behind a v1 catalog and prove fail-closed in staging. You will see results the first time a proposed refund or send dies at a row instead of in a chat log. Broader coverage — every `job_type`, reviewed caps, staffed pending SLAs — is a calendar, not a sprint slogan. If you cannot name an owner on day one, you are not starting the week.
What should I skip if I only have a week?
Skip control-framework mapping, per-user exception matrices, multi-agent inheritance, and letting the model draft its own allowlist. Do not skip owner + backup, default deny, the three scary tool classes, payload-bound pending, and a fail-closed drill. Unstaffed pending stays deny. A prettier demo is not a loosen change.
When is this not worth doing yet?
If you have no tools with side effects — no send, no write, no refund, no deploy — you do not have an agent policy problem yet; you have a chatbot. If the job is a fixed flowchart, ship a workflow with a gate on the write node instead of an agent. If nobody will own the catalog or staff pending, do not turn the model loose on production writes.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot