Sandboxed Tool Use: Letting Agents Act Without Letting Them Loose
Tool use without a sandbox is an API key with opinions. Allowlists, scoped credentials, blast-radius caps, dry-runs, and human gates earn a production write.
William Spurlock Founder — Spurlock Studios Updated 23 MIN
Tool use without a sandbox is an API key with opinions. The model can propose a CRM write, a refund, or a public send; the runner decides whether that call is legal. Allowlists, scoped credentials, blast-radius caps, dry-runs, and human gates are how an agent earns the right to touch production.
This spoke sits inside the Agentic Systems Operating Manual. The doorway in front of a proposed call is pre-execution policy gates. This post is the cage around the tools themselves: what the agent may call, as whom, how far, and when a person still has to click.
The short answer
- Allowlist a typed catalog. Name, purpose, input schema, output schema, side-effect class. The model chooses among entries. It does not get a blank “run whatever.”
- Scope credentials per tool. Read-only for enrichment. A separate write token for the one field the agent may patch. Never the founder’s personal OAuth.
- Cap blast radius in config. Max rows, max sends, max dollars, max deletes (usually zero), timeouts, concurrency. Vibes are not a cap.
- Dry-run before writes. First week: “would write X.” Then dual-write to a shadow field. Then cut over with caps.
- Gate the irreversible class. Refunds, public sends, schema changes, legal-sounding commitments. Autonomy is a promotion, not a default.
OWASP’s LLM06:2025 Excessive Agency names the three ways this goes wrong: excessive functionality, excessive permissions, excessive autonomy. A sandbox is the engineering answer to all three.
What a sandbox actually enforces
A sandbox is not a marketing word for “we thought about security.” It is a concrete boundary the runner enforces in code. NIST’s AI Risk Management Framework (NIST AI 100-1) asks you to map, measure, and manage risk across the lifecycle. The Generative AI Profile (NIST AI 600-1, July 2024) is the companion for systems that call tools. Neither document is a product. Both are the reason you write the boundary down.
| Boundary | What the runner checks | What “theater” looks like |
|---|---|---|
| Tool allowlist | Name is in the catalog for this job and this state | A prompt that says “please do not call delete” |
| Least privilege | Token injected for this tool, this environment | One admin key shared across the fleet |
| Blast-radius cap | Counter would exceed the configured max | “We watch Slack and hope” |
| Dry-run mode | Write tools return a preview, no side effect | Staging credentials pointed at prod |
| Human gate | Irreversible class waits for a logged actor | A YAML permissions: file nothing reads |
| Fail closed | Auth, schema, or policy miss → escalate / abort | Model invents a workaround with another tool |
MCP servers, custom HTTP tools, and in-process functions are all fine — if they sit behind that boundary. Unrestricted shell, wildcard admin tokens, and “the model can invent new tools” are how incidents start.
CISA’s Secure by Design line is the same idea in plainer English: the producer owns the default, not the operator who forgot to lock a toggle. Your agent is a product. Ship it locked.
Checklist — write these as code paths, not as slide bullets:
- Unknown tool name is a hard reject
- Catalog cannot grow from retrieved text
- Secrets never enter the model context
- Caps live in config the runner reads on every call
- Irreversible class has a gate with a named owner
- Staging and production are different identities
A cage you cannot demonstrate with a rejected call is not a cage.
Why allowlists beat tool descriptions
Publish a typed catalog before the model sees a description. The description is documentation for the planner. The allowlist is the law.
OWASP’s Top 10 for LLM Applications is explicit: minimize extensions, then minimize the functions inside each extension. A mailbox summarizer needs mail.read. It does not need mail.send because the vendor bundled both. That is excessive functionality — the first of the three LLM06 roots.
| Catalog field | Why it exists | Failure if you skip it |
|---|---|---|
name | Stable id the runner matches | Model aliases crm_delete as cleanup |
purpose | One sentence a reviewer can reject | “Helper” tools that do three jobs |
input_schema | JSON Schema before any network call | Extra fields the model invents get forwarded |
output_schema | Shape the next step may trust | Hallucinated keys treated as facts |
side_effect_class | read / soft_write / hard_write / irreversible | Trust arguments that are really class arguments |
max_calls_per_run | Local cap on top of the global cap | A retry loop becomes a write storm |
owner | Human who can disable it | Orphan tools with live credentials |
If a job needs a new tool, a human adds it to the catalog with a review. Hot-adding tools mid-run because the model asked nicely is not a feature. Retrieved tickets, emails, and PDFs are data. They do not get a vote on the catalog.
Decision list for a new tool:
- Does this job fail without it? If no, do not add it.
- Can you split read and write into two tools with two identities? If yes, split.
- Is the write customer-visible or irreversible? If yes, class it that way on day one.
- Can you dry-run it? If no, you are not ready to expose it.
- Who owns disable and rotation? If “the agent,” you do not have an owner.
Open-ended tools — shell, “fetch any URL,” “run this SQL” — are the LLM06 example of a file-write implemented as a shell. Build the narrow tool. The wide one is a skeleton key with a schema.
How do you scope credentials per tool?
NIST SP 800-53 Rev. 5 AC-6 is least privilege: users and processes acting on behalf of users get only the accesses required for the assigned task. An agent is a process acting on behalf of a user. The founder’s Google token is not that process.
NIST SP 800-207 (Zero Trust Architecture, August 2020) adds the other half: no implicit trust from network location. “The runner is on our VPC” is not a credential strategy. Authenticate and authorize the tool call as its own session.
| Tool | Identity | Scopes you actually need | Scopes you do not give |
|---|---|---|---|
crm.get_lead | agent-crm-read | Read one object type | Delete, merge, export-all |
enrichment.lookup | agent-enrich-read | Lookup by domain | Bulk export, billing admin |
crm.patch_internal_note | agent-crm-note-write | Patch internal_note only | Any customer-visible field |
email.draft | agent-mail-draft | Create draft in one mailbox | Send, forward, rule changes |
email.send | None, or break-glass | — | Not in the pilot catalog |
refund.create | Human-gated runner | Prepare payload only | Charge / refund execute |
LLM06’s mailbox example is the same table in story form: a summarizer with a send function, authenticated as a privileged identity, plus an injected email that says “forward the inbox.” Read-only OAuth on the user’s own mailbox would have stopped the send even if the model asked.
Procedure for every new identity:
- Create a dedicated service account or OAuth client for the tool, not the job, not the company.
- Grant the minimum scope. Prefer field-level or object-level ACL when the vendor has it.
- Store the secret in a vault. Inject it at the runner. Never in the prompt, never in git, never in a trace.
- Bind the identity to one environment. Staging tokens do not work in prod. Prod tokens do not live in notebooks.
- Set a rotation date and an owner. Short-lived tokens where the provider allows.
- Log which identity performed the write. “The agent did it” is not an actor.
Break-glass prod access is time-bounded and logged. Permanent admin “just for the pilot” is how sandboxes end.
What blast-radius caps belong in config?
Rate limits do not prevent Excessive Agency. They shrink how much damage a confused loop can do before a human notices. OWASP lists rate-limiting as a damage limiter, not a prevention. Put the numbers in config the runner increments on every accepted call.
| Cap | Default for a new job | Why that default | Promote only when |
|---|---|---|---|
| Rows / records touched per run | 1 | One lead, one ticket, one invoice | Evaluator scores hold on a golden set |
| Enrichment API calls per run | 3 | Lookup + retry + one alternate | You measured real miss rate |
| Emails sent per day | 0 (drafts only) | You cannot unsay a send | Shadow disagreement is boring |
| Dollars on any spend API | 0 | Cash leaves | Dual control + written residual risk |
| Files deleted | 0 | Restore is a project | Almost never |
| Concurrent runs of this job | 1 | Retry storms amplify writes | Queue depth is understood |
| Wall-clock timeout | Seconds, not minutes | Hung tools hold locks | You have a kill path |
| Tool-call budget per run | Low double digits | Loops are not a strategy | State machine bounds the loop |
Caps that live in a Slack thread are not caps. Caps that the model can talk the runner out of are not caps.
Checklist for the counter implementation:
- Increment before the side effect, not after
- Persist counters across retries and process restarts
- Key them by
job_type + environment + utc_day(or tighter) - On exceed: typed
policy_violation:blast_radius, thenescalate - Do not reset a cap because the model apologized
- Page a human when a cap trips twice in a day — that is a loop, not a limit
A retry storm without a concurrency cap is a write amplifier. Idempotency keys belong here too: the state machine supplies a client-generated key; the runner persists the outcome so a replay does not double-apply.
When do dry-runs and dual-write earn a write?
A dry-run is a write tool that returns “would write X” and stops. It is not a staging environment with a prod token. It is not a prompt that says “just pretend.”
NIST AI RMF’s MEASURE function is the reason this exists: you cannot manage a risk you have not compared to a human baseline. The first week in a new system, the agent proposes. A person would have written Y. You score the delta. Then you dual-write to a shadow field the customer never sees. Then you cut over with caps.
| Stage | What the tool does | Who sees it | Exit criteria |
|---|---|---|---|
| Dry-run | Preview payload, no write | Engineers + domain reviewer | Preview matches human intent on a golden set |
| Dual-write | Write shadow / internal field | Internal only | Disagreement rate is low and understood |
| Gated hard write | Customer-visible field, human click | Reviewer, then customer | SLA on the gate; reject reasons logged |
| Capped auto write | Hard write under counters | Customer / ops | Written residual risk; cap still on |
| Irreversible | Send, charge, delete | Human, often forever | Business accepts residual risk in writing |
Procedure for one new write tool:
- Implement the preview path first. Same schema,
dry_run: trueforced by the runner when the job is in that stage. - Collect N previews against real (or realistic staging) records. Score them the way you will score production.
- Turn on dual-write to an internal field. Keep the customer-visible field human.
- Compare shadow vs human for a measured window — not “a quiet week.”
- Cut over the hard write with the cap at 1 and the gate still on.
- Promote the gate off for that tool only after the reject rate is boring.
Calendar time is not a promotion signal. A quiet week is not a measured week.
If you cannot dry-run a vendor API, wrap it. The wrapper records the would-be call and returns the preview. The real call stays behind a flag the runner owns.
Which actions stay behind a human gate?
Human gates are how you kill excessive autonomy — LLM06’s third root. The agent prepares the payload. A person (or a stricter secondary policy engine) releases it. Pair this post with pre-execution policy gates: the gate is the doorway; the sandbox is the room the doorway opens into.
| Action class | Examples | Default | First cohort of live traffic |
|---|---|---|---|
| Irreversible money | Refund, payout, ad-spend change | Human gate | Every item |
| Irreversible contact | Email send, SMS, public reply | Human gate | Draft auto; send gated |
| Irreversible mutate | Delete, merge, permission drop | Human gate | Dual control if blast is wide |
| Legal-adjacent | Contract language, “we commit to…” | Human gate | Named role, logged actor |
| Hard write | Customer-visible CRM field | Cap + monitor; often gated | Gate until shadow is boring |
| Soft write | Internal note, draft, tag | Allow after evaluator pass | Monitor; no click |
| Read | Fetch ticket, search docs | Allow under rate caps | Still log |
Checklist — put this on the job, not in someone’s head:
- Money movement named and gated
- Customer-visible contact named and gated
- Deletes / merges / permission drops named and gated
- Soft writes listed as auto, with an owner who can demote them
- “Unsure” defaults to a gate for the first live cohort
- Dual control for refunds, production DNS, payroll-adjacent actions
Dual control is older than language models. Two approvals: the agent’s prepared payload plus a human, or two humans. Agents do not exempt you.
Removing a gate is a one-line policy change. Explaining an accidental customer email is a week.
How should the tool runner fail closed?
The runner is the enforcement point. Policy that lives only in the prompt (“please do not call delete”) is hope. OWASP’s complete-mediation line for LLM06 is the same rule: authorize in the downstream system, not by asking the model if the action is allowed.
Numbered path for every proposed call:
- Authenticate the run — tenant, job type, state name. Unknown run is a reject.
- Reject unknown tool names. No fuzzy match. No “closest catalog entry.”
- Validate args against JSON Schema before any network call. Extra fields are a reject, not a strip-and-continue, until you have a written exception.
- Inject secrets from a vault. The model never sees them. The model never supplies them.
- Apply rate limits and blast-radius counters. Increment first.
- Execute with a timeout. Hung is
abort, not “try a different tool.” - Normalize errors into typed failures the state machine understands.
- Emit a redacted trace span. Full PII does not belong in Slack.
If auth fails, schema fails, or the sandbox rejects a call, the run goes to escalate or abort. It does not go to “invent a workaround with another tool.” Clever workarounds are how sandboxes die.
| Failure | Typed result | Next state | What you do not do |
|---|---|---|---|
| Unknown tool | policy_violation:unknown_tool | abort | Alias to a similar name |
| Schema miss | policy_violation:schema | escalate | Coerce types and proceed |
| Cap exceeded | policy_violation:blast_radius | escalate | Reset the counter |
| Auth / vault miss | policy_violation:auth | abort | Fall back to a shared admin key |
| Timeout | tool_timeout | abort or retry once with backoff | Fan out parallel retries |
| Downstream 4xx | tool_rejected | typed retry policy | Retry a non-idempotent write blindly |
Return errors the model can act on without revealing secrets: policy_violation:recipient_domain rather than a stack trace with a vault path. OWASP ASVS (v5.0.0 as of May 2025) is the input-sanitization and injection-prevention checklist for the code around the model. The model is not the sanitizer.
Idempotency belongs in this same path. For hard writes, accept a client-generated key from the state machine and persist outcomes so retries do not double-apply.
What prompt injection does to an unsandboxed agent
Any content the agent reads — tickets, emails, PDFs, web pages — can contain instructions. OWASP LLM01:2025 Prompt Injection is the entry. The community write-up is the same idea in attack language: system text and user text share one token stream.
Sandboxing does not solve injection. Without a sandbox, injection has a bigger blast radius. With a sandbox, injection can still ruin a summary. It should not be able to send the inbox.
| Control | What it stops | What it does not stop |
|---|---|---|
| Tool allowlist that cannot grow from retrieved text | “Please enable email.send” in a PDF | A send tool you already allowed |
| Instruction channel ≠ document channel | Concatenating a ticket as system policy | The model still reads the ticket |
| Strip requests for new tools or secrets | Catalog mutation via content | A legal tool called with a bad arg |
| Runner-side policy on args | SSRF, wrong recipient domain, DROP SQL | A well-formed bad send you allowed |
| Human gate on irreversible | Injected “send now” | A gated payload a tired human rubber-stamps |
| Redaction in traces | Secret echo into Slack | The model seeing the secret in the first place |
Treat untrusted text as data, not as system policy. Practical runner checks that belong next to the allowlist:
- Email fields must match an allowlisted recipient domain for outbound drafts.
- SQL tools — if you ever allow them — reject
DROP/UPDATE/DELETEafter a real parse, not a regex vibe. - URL fetchers block link-local, metadata, and private IP ranges. OWASP A10:2021 SSRF is the category. The SSRF Prevention Cheat Sheet is the allowlist-not-denylist rule: positive allow list for scheme, port, and destination; do not “block the bad IPs” and call it done.
“Read the webpage” is a powerful tool and a common exfiltration path under injection. If you need it, constrain egress: allowlisted hosts, size limits, content-type checks, no redirect-following into RFC1918.
Injection that cannot call a write tool is noise. That is the whole point of the cage.
What to log for every tool call
For ops and forensics you want a span you can replay without leaking a customer. Redaction is part of the sandbox design, not an afterthought. Agents increase the number of places secrets can leak: prompts, traces, tool args, “helpful” plugin logs.
| Field | Keep | Redact or drop |
|---|---|---|
tool_name | Always | — |
run_id / state_name / job_type | Always | — |
side_effect_class | Always | — |
identity_id | Which service account | Token material |
args_redacted | Keys + types + allowlisted values | PII, secrets, raw email bodies |
result_class | ok / typed error | Full downstream body |
duration_ms | Always | — |
cost_attribution | Model + tool vendor | Customer identifiers in the label |
cap_counters | Values after increment | — |
actor | Human id on gated releases | — |
Checklist:
- Default deny on new fields in the span
- PII dictionary applied before the span leaves the runner
- Slack / chat gets a link to the record, not the record
- Export path assumes a future incident-response dump
- Plugin “debug logging” is off in prod or pointed at a sink you own
You do not want full PII in a Slack channel. You do want to answer “which identity patched this lead at 14:03” without opening a laptop in a parking lot.
Environment separation: staging is not a costume
NIST SP 800-207 treats every session as untrusted until authenticated. Copy-pasting the prod token into a notebook “just to test” is the opposite of that. Promoting a tool from staging to prod is a change-controlled event: review side-effect class, caps, and whether the evaluator covers the new failure modes.
| Environment | Credentials | Writes | Audience | Allowed surprises |
|---|---|---|---|---|
| Dev | Mocks / fixtures | Fake | Engineers | Broken previews |
| Staging | Staging systems | Real staging data | Domain reviewers | Wrong field, still internal |
| Prod | Least-privilege prod | Caps + gates | Customers / ops | None you have not named |
Promotion checklist:
- Side-effect class reviewed by the tool owner
- Caps set in prod config, not inherited as “unlimited”
- Dry-run path still callable for support
- Evaluator covers the new failure modes
- Staging token revoked from any laptop that touched prod
- Break-glass documented with a clock
Dev and staging may share a catalog shape. They do not share identities. If a vendor has only one tenant, you do not have staging. You have a prod experiment. Gate harder.
Third-party MCP and plugin risk
Marketplace tools arrive with someone else’s threat model. MCP is a transport and an interface. The authorization spec is optional OAuth for HTTP transports. It does not automatically enforce least privilege or blast-radius caps. A wide-open MCP server is not a sandbox.
Before a server enters an agent catalog:
| Check | Pass looks like | Fail looks like |
|---|---|---|
| Scopes requested | Read-only, one resource | admin, offline_access, “all mailboxes” |
| Throwaway tenant | You ran it against fake data | First run is your CRM |
| Logging / exfil | You read the egress and the log sink | “Helpful” telemetry you cannot disable |
| Version pin | Digest or exact version | latest |
| Tool-list mutation | Disabled at runtime | Server can add tools after connect |
| Schema honesty | Tools match the docs | Hidden execute next to search |
Spurlock Studios would rather wrap two HTTP endpoints you own than enable twenty plugins you have not read. The operating manual treats sandboxes as a first-class layer for that reason.
Decision list:
- Do you own the server code? Prefer yes.
- Can you pin and review the tool list on every connect? If no, do not connect.
- Can the server see secrets other than its own? If yes, split the catalog.
- Can it reach the public internet beyond an allowlist? If yes, treat it as a fetch tool and apply SSRF rules.
- Can it send email or move money? If yes, it is irreversible even if the vendor named it
notify.
STDIO MCP that inherits the operator’s environment inherits the operator’s keys. That is a founder laptop with a protocol.
Progressive autonomy, then inventory
Climb the ladder per job type using online evaluator scores and incident count — not calendar time.
- Dry-run only
- Soft writes to internal fields
- Hard writes with a human gate
- Hard writes auto under caps
- Irreversible actions still gated (often forever)
| Rung | What you measure | Kill switch |
|---|---|---|
| Dry-run | Preview vs human intent | Stay here |
| Soft write | Shadow disagreement | Demote to dry-run |
| Gated hard write | Reject reasons, SLA misses | Keep the gate |
| Capped auto | Cap trips, incident count | Re-gate that tool |
| Irreversible | Residual risk in writing | Dual control |
Monthly inventory — every tool in every catalog:
- Owner still at the company
- Last used date; disable after a written idle window
- Side-effect class still true (a “read” that grew a write is a new tool)
- Credential age vs rotation calendar
- Cap values still match the job
- Staging and prod identities still separate
Orphan tools with live credentials are unpaid attackers waiting for a prompt injection. Disable them. Do not “leave them in case we need them.”
Failure mode: an API key with opinions
The failure is not theoretical. It is a specific shape I keep seeing across 500+ automations and 20,000+ hours on agentic systems: a team wires a model to a platform key, pastes a system prompt that says “be careful,” and calls it a pilot.
What breaks. The model — confused, injected, or just wrong — calls a write that was sitting on the same key as the read. A merge, a send, a delete, a refund. Caps were a conversation. The gate was a Slack DM. Staging was prod with a different URL in a comment.
What it costs. A customer-visible action you cannot recall, a credential rotation under fire, and a week of explaining why “the AI” had admin. The restore is never the hard part. The trust is.
What you do instead.
| Anti-pattern | Replacement |
|---|---|
| “The agent has the Zapier key” | Per-tool identities, allowlisted actions |
| Production credentials in the prompt | Vault inject at the runner |
Sandbox theater (permissions: YAML nobody enforces) | Runner rejects with a typed error you can demo |
| Expanding scope mid-pilot | One job. New tools wait for the next slice |
| Shell “just for debugging” in prod | Purpose-built tools with schemas |
latest on a marketplace MCP | Pin, review, throwaway tenant first |
| Shared admin token across environments | Separate identities, break-glass with a clock |
| Prompt-only “do not delete” | Delete is not in the catalog |
Demo the reject path before you demo the happy path. If a vendor cannot show a rejected tool call, a cap trip, a dry-run, and a human gate, keep shopping.
How Spurlock Studios applies this in a pilot
In the $1,500 · 5-day pilot we pick the minimum tool set for one sentence-sized job. Reads first. Writes only if the job demands them, usually to an internal surface. Irreversible actions stay human-gated. You leave with a catalog and a runner you can keep operating.
A short example we actually scope:
- Job: enrich a lead record with firmographics and draft an internal note.
- Allowlist:
crm.get_lead,enrichment.lookup,crm.patch_internal_note. - Not allowlist:
crm.merge_leads,email.send,crm.delete. - Caps: one lead per run; enrichment API max 3 calls; patch only
internal_note. - Evaluator: note must cite enrichment fields present in the tool result; no invented revenue numbers when enrichment returned null.
That is sandboxed tool use. “Here’s our admin API key, go enrich everything” is not.
Three locks sit together in the operating manual:
- State machine — only
actmay call tools with side effects. - Sandbox — only allowlisted tools with caps.
- Evaluator — only passing artifacts proceed to hard writes.
Remove any lock and the system fails open in a different way. The sandbox is the lock this spoke owns.
Procurement RFP — ask vendors to demonstrate:
- A rejected unknown-tool call
- A blast-radius cap that stops a loop
- A dry-run that does not write
- A human gate on an irreversible class
- Separate staging and prod identities
- Redacted traces you can export
If they can only show a happy-path demo, you are buying an API key with opinions. Map and offer: /agentic.
FAQ
What are sandboxed AI tools?
Sandboxed AI tools are agent-callable functions wrapped in allowlists, least-privilege credentials, quantitative caps, and policies for irreversible actions. The model proposes a call. The runner enforces whether that call is legal. A description in a prompt is not a sandbox.
How do you implement safe tool use for agents?
Define a typed tool catalog, classify side effects, scope credentials per tool, and enforce caps and timeouts in the runner. Start irreversible actions behind human gates, fail closed on errors, and log redacted traces. Promote autonomy only after evaluation scores hold on that job — not after a quiet week.
Is MCP enough to be sandboxed?
No. MCP is a transport and interface pattern. Authorization in the spec is optional OAuth for HTTP. It does not automatically enforce least privilege or blast-radius caps. You still design the server’s capabilities, auth, and policy. A wide-open MCP server is not a sandbox.
Should agents have shell access?
Almost never in a business pilot. If you need shell for a devtools agent, isolate the environment, drop privileges, restrict network, and treat every command as high risk. Default to purpose-built tools with schemas. A shell is excessive functionality with a blinking cursor.
When can we remove human gates?
When offline and online evaluator scores meet the bar, blast-radius caps are proven under load, and the business accepts the residual risk in writing. Remove the gate per tool and per job type. Gates are a control, not an insult to the model. Irreversible sends and deletes often stay gated forever.
How does this relate to Spurlock Studios’ agentic offer?
Sandbox design is part of every pilot and build. We would rather ship a narrow, caged agent that works than a broad agent that can email your customers by accident. The five-day, $1,500 pilot on /agentic is the on-ramp; book via /contact?intent=agentic-pilot.
CTA
An agent that can act without a cage is an API key that talks. Build the catalog, the identities, the caps, and the gate — then let it earn a write.
Read the Agentic Systems Operating Manual, then use the agentic lane or book the pilot to put one job in a cage you can demo.
What questions does this article answer?
- What are sandboxed AI tools?
- Sandboxed AI tools are agent-callable functions wrapped in allowlists, least-privilege credentials, quantitative caps, and policies for irreversible actions. The model proposes a call. The runner enforces whether that call is legal. A description in a prompt is not a sandbox.
- How do you implement safe tool use for agents?
- Define a typed tool catalog, classify side effects, scope credentials per tool, and enforce caps and timeouts in the runner. Start irreversible actions behind human gates, fail closed on errors, and log redacted traces. Promote autonomy only after evaluation scores hold on that job — not after a quiet week.
- Is MCP enough to be sandboxed?
- No. MCP is a transport and interface pattern. Authorization in the spec is optional OAuth for HTTP. It does not automatically enforce least privilege or blast-radius caps. You still design the server’s capabilities, auth, and policy. A wide-open MCP server is not a sandbox.
- Should agents have shell access?
- Almost never in a business pilot. If you need shell for a devtools agent, isolate the environment, drop privileges, restrict network, and treat every command as high risk. Default to purpose-built tools with schemas. A shell is excessive functionality with a blinking cursor.
- When can we remove human gates?
- When offline and online evaluator scores meet the bar, blast-radius caps are proven under load, and the business accepts the residual risk in writing. Remove the gate per tool and per job type. Gates are a control, not an insult to the model. Irreversible sends and deletes often stay gated forever.
- How does this relate to Spurlock Studios’ agentic offer?
- Sandbox design is part of every pilot and build. We would rather ship a narrow, caged agent that works than a broad agent that can email your customers by accident. The five-day, $1,500 pilot on [/agentic](/agentic) is the on-ramp; book via [/contact?intent=agentic-pilot](/contact?intent=agentic-pilot).
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.