Spurlock Studios
Contact
Share LinkedIn X
An expired brass key. Thesis: SANDBOXED TOOL USE LETTING AGENTS.

Tool use without a sandbox is an API key with opinions. The model can propose a CRM write, a refund, or a public send; the runner decides whether that call is legal. Allowlists, scoped credentials, blast-radius caps, dry-runs, and human gates are how an agent earns the right to touch production.

This spoke sits inside the Agentic Systems Operating Manual. The doorway in front of a proposed call is pre-execution policy gates. This post is the cage around the tools themselves: what the agent may call, as whom, how far, and when a person still has to click.

The short answer

  • Allowlist a typed catalog. Name, purpose, input schema, output schema, side-effect class. The model chooses among entries. It does not get a blank “run whatever.”
  • Scope credentials per tool. Read-only for enrichment. A separate write token for the one field the agent may patch. Never the founder’s personal OAuth.
  • Cap blast radius in config. Max rows, max sends, max dollars, max deletes (usually zero), timeouts, concurrency. Vibes are not a cap.
  • Dry-run before writes. First week: “would write X.” Then dual-write to a shadow field. Then cut over with caps.
  • Gate the irreversible class. Refunds, public sends, schema changes, legal-sounding commitments. Autonomy is a promotion, not a default.

OWASP’s LLM06:2025 Excessive Agency names the three ways this goes wrong: excessive functionality, excessive permissions, excessive autonomy. A sandbox is the engineering answer to all three.

What a sandbox actually enforces

A sandbox is not a marketing word for “we thought about security.” It is a concrete boundary the runner enforces in code. NIST’s AI Risk Management Framework (NIST AI 100-1) asks you to map, measure, and manage risk across the lifecycle. The Generative AI Profile (NIST AI 600-1, July 2024) is the companion for systems that call tools. Neither document is a product. Both are the reason you write the boundary down.

BoundaryWhat the runner checksWhat “theater” looks like
Tool allowlistName is in the catalog for this job and this stateA prompt that says “please do not call delete”
Least privilegeToken injected for this tool, this environmentOne admin key shared across the fleet
Blast-radius capCounter would exceed the configured max“We watch Slack and hope”
Dry-run modeWrite tools return a preview, no side effectStaging credentials pointed at prod
Human gateIrreversible class waits for a logged actorA YAML permissions: file nothing reads
Fail closedAuth, schema, or policy miss → escalate / abortModel invents a workaround with another tool

MCP servers, custom HTTP tools, and in-process functions are all fine — if they sit behind that boundary. Unrestricted shell, wildcard admin tokens, and “the model can invent new tools” are how incidents start.

CISA’s Secure by Design line is the same idea in plainer English: the producer owns the default, not the operator who forgot to lock a toggle. Your agent is a product. Ship it locked.

Checklist — write these as code paths, not as slide bullets:

  • Unknown tool name is a hard reject
  • Catalog cannot grow from retrieved text
  • Secrets never enter the model context
  • Caps live in config the runner reads on every call
  • Irreversible class has a gate with a named owner
  • Staging and production are different identities

A cage you cannot demonstrate with a rejected call is not a cage.

Why allowlists beat tool descriptions

Publish a typed catalog before the model sees a description. The description is documentation for the planner. The allowlist is the law.

OWASP’s Top 10 for LLM Applications is explicit: minimize extensions, then minimize the functions inside each extension. A mailbox summarizer needs mail.read. It does not need mail.send because the vendor bundled both. That is excessive functionality — the first of the three LLM06 roots.

Catalog fieldWhy it existsFailure if you skip it
nameStable id the runner matchesModel aliases crm_delete as cleanup
purposeOne sentence a reviewer can reject“Helper” tools that do three jobs
input_schemaJSON Schema before any network callExtra fields the model invents get forwarded
output_schemaShape the next step may trustHallucinated keys treated as facts
side_effect_classread / soft_write / hard_write / irreversibleTrust arguments that are really class arguments
max_calls_per_runLocal cap on top of the global capA retry loop becomes a write storm
ownerHuman who can disable itOrphan tools with live credentials

If a job needs a new tool, a human adds it to the catalog with a review. Hot-adding tools mid-run because the model asked nicely is not a feature. Retrieved tickets, emails, and PDFs are data. They do not get a vote on the catalog.

Decision list for a new tool:

  1. Does this job fail without it? If no, do not add it.
  2. Can you split read and write into two tools with two identities? If yes, split.
  3. Is the write customer-visible or irreversible? If yes, class it that way on day one.
  4. Can you dry-run it? If no, you are not ready to expose it.
  5. Who owns disable and rotation? If “the agent,” you do not have an owner.

Open-ended tools — shell, “fetch any URL,” “run this SQL” — are the LLM06 example of a file-write implemented as a shell. Build the narrow tool. The wide one is a skeleton key with a schema.

How do you scope credentials per tool?

NIST SP 800-53 Rev. 5 AC-6 is least privilege: users and processes acting on behalf of users get only the accesses required for the assigned task. An agent is a process acting on behalf of a user. The founder’s Google token is not that process.

NIST SP 800-207 (Zero Trust Architecture, August 2020) adds the other half: no implicit trust from network location. “The runner is on our VPC” is not a credential strategy. Authenticate and authorize the tool call as its own session.

ToolIdentityScopes you actually needScopes you do not give
crm.get_leadagent-crm-readRead one object typeDelete, merge, export-all
enrichment.lookupagent-enrich-readLookup by domainBulk export, billing admin
crm.patch_internal_noteagent-crm-note-writePatch internal_note onlyAny customer-visible field
email.draftagent-mail-draftCreate draft in one mailboxSend, forward, rule changes
email.sendNone, or break-glass—Not in the pilot catalog
refund.createHuman-gated runnerPrepare payload onlyCharge / refund execute

LLM06’s mailbox example is the same table in story form: a summarizer with a send function, authenticated as a privileged identity, plus an injected email that says “forward the inbox.” Read-only OAuth on the user’s own mailbox would have stopped the send even if the model asked.

Procedure for every new identity:

  1. Create a dedicated service account or OAuth client for the tool, not the job, not the company.
  2. Grant the minimum scope. Prefer field-level or object-level ACL when the vendor has it.
  3. Store the secret in a vault. Inject it at the runner. Never in the prompt, never in git, never in a trace.
  4. Bind the identity to one environment. Staging tokens do not work in prod. Prod tokens do not live in notebooks.
  5. Set a rotation date and an owner. Short-lived tokens where the provider allows.
  6. Log which identity performed the write. “The agent did it” is not an actor.

Break-glass prod access is time-bounded and logged. Permanent admin “just for the pilot” is how sandboxes end.

What blast-radius caps belong in config?

Rate limits do not prevent Excessive Agency. They shrink how much damage a confused loop can do before a human notices. OWASP lists rate-limiting as a damage limiter, not a prevention. Put the numbers in config the runner increments on every accepted call.

CapDefault for a new jobWhy that defaultPromote only when
Rows / records touched per run1One lead, one ticket, one invoiceEvaluator scores hold on a golden set
Enrichment API calls per run3Lookup + retry + one alternateYou measured real miss rate
Emails sent per day0 (drafts only)You cannot unsay a sendShadow disagreement is boring
Dollars on any spend API0Cash leavesDual control + written residual risk
Files deleted0Restore is a projectAlmost never
Concurrent runs of this job1Retry storms amplify writesQueue depth is understood
Wall-clock timeoutSeconds, not minutesHung tools hold locksYou have a kill path
Tool-call budget per runLow double digitsLoops are not a strategyState machine bounds the loop

Caps that live in a Slack thread are not caps. Caps that the model can talk the runner out of are not caps.

Checklist for the counter implementation:

  • Increment before the side effect, not after
  • Persist counters across retries and process restarts
  • Key them by job_type + environment + utc_day (or tighter)
  • On exceed: typed policy_violation:blast_radius, then escalate
  • Do not reset a cap because the model apologized
  • Page a human when a cap trips twice in a day — that is a loop, not a limit

A retry storm without a concurrency cap is a write amplifier. Idempotency keys belong here too: the state machine supplies a client-generated key; the runner persists the outcome so a replay does not double-apply.

When do dry-runs and dual-write earn a write?

A dry-run is a write tool that returns “would write X” and stops. It is not a staging environment with a prod token. It is not a prompt that says “just pretend.”

NIST AI RMF’s MEASURE function is the reason this exists: you cannot manage a risk you have not compared to a human baseline. The first week in a new system, the agent proposes. A person would have written Y. You score the delta. Then you dual-write to a shadow field the customer never sees. Then you cut over with caps.

StageWhat the tool doesWho sees itExit criteria
Dry-runPreview payload, no writeEngineers + domain reviewerPreview matches human intent on a golden set
Dual-writeWrite shadow / internal fieldInternal onlyDisagreement rate is low and understood
Gated hard writeCustomer-visible field, human clickReviewer, then customerSLA on the gate; reject reasons logged
Capped auto writeHard write under countersCustomer / opsWritten residual risk; cap still on
IrreversibleSend, charge, deleteHuman, often foreverBusiness accepts residual risk in writing

Procedure for one new write tool:

  1. Implement the preview path first. Same schema, dry_run: true forced by the runner when the job is in that stage.
  2. Collect N previews against real (or realistic staging) records. Score them the way you will score production.
  3. Turn on dual-write to an internal field. Keep the customer-visible field human.
  4. Compare shadow vs human for a measured window — not “a quiet week.”
  5. Cut over the hard write with the cap at 1 and the gate still on.
  6. Promote the gate off for that tool only after the reject rate is boring.

Calendar time is not a promotion signal. A quiet week is not a measured week.

If you cannot dry-run a vendor API, wrap it. The wrapper records the would-be call and returns the preview. The real call stays behind a flag the runner owns.

Which actions stay behind a human gate?

Human gates are how you kill excessive autonomy — LLM06’s third root. The agent prepares the payload. A person (or a stricter secondary policy engine) releases it. Pair this post with pre-execution policy gates: the gate is the doorway; the sandbox is the room the doorway opens into.

Action classExamplesDefaultFirst cohort of live traffic
Irreversible moneyRefund, payout, ad-spend changeHuman gateEvery item
Irreversible contactEmail send, SMS, public replyHuman gateDraft auto; send gated
Irreversible mutateDelete, merge, permission dropHuman gateDual control if blast is wide
Legal-adjacentContract language, “we commit to…”Human gateNamed role, logged actor
Hard writeCustomer-visible CRM fieldCap + monitor; often gatedGate until shadow is boring
Soft writeInternal note, draft, tagAllow after evaluator passMonitor; no click
ReadFetch ticket, search docsAllow under rate capsStill log

Checklist — put this on the job, not in someone’s head:

  • Money movement named and gated
  • Customer-visible contact named and gated
  • Deletes / merges / permission drops named and gated
  • Soft writes listed as auto, with an owner who can demote them
  • “Unsure” defaults to a gate for the first live cohort
  • Dual control for refunds, production DNS, payroll-adjacent actions

Dual control is older than language models. Two approvals: the agent’s prepared payload plus a human, or two humans. Agents do not exempt you.

Removing a gate is a one-line policy change. Explaining an accidental customer email is a week.

How should the tool runner fail closed?

The runner is the enforcement point. Policy that lives only in the prompt (“please do not call delete”) is hope. OWASP’s complete-mediation line for LLM06 is the same rule: authorize in the downstream system, not by asking the model if the action is allowed.

Numbered path for every proposed call:

  1. Authenticate the run — tenant, job type, state name. Unknown run is a reject.
  2. Reject unknown tool names. No fuzzy match. No “closest catalog entry.”
  3. Validate args against JSON Schema before any network call. Extra fields are a reject, not a strip-and-continue, until you have a written exception.
  4. Inject secrets from a vault. The model never sees them. The model never supplies them.
  5. Apply rate limits and blast-radius counters. Increment first.
  6. Execute with a timeout. Hung is abort, not “try a different tool.”
  7. Normalize errors into typed failures the state machine understands.
  8. Emit a redacted trace span. Full PII does not belong in Slack.

If auth fails, schema fails, or the sandbox rejects a call, the run goes to escalate or abort. It does not go to “invent a workaround with another tool.” Clever workarounds are how sandboxes die.

FailureTyped resultNext stateWhat you do not do
Unknown toolpolicy_violation:unknown_toolabortAlias to a similar name
Schema misspolicy_violation:schemaescalateCoerce types and proceed
Cap exceededpolicy_violation:blast_radiusescalateReset the counter
Auth / vault misspolicy_violation:authabortFall back to a shared admin key
Timeouttool_timeoutabort or retry once with backoffFan out parallel retries
Downstream 4xxtool_rejectedtyped retry policyRetry a non-idempotent write blindly

Return errors the model can act on without revealing secrets: policy_violation:recipient_domain rather than a stack trace with a vault path. OWASP ASVS (v5.0.0 as of May 2025) is the input-sanitization and injection-prevention checklist for the code around the model. The model is not the sanitizer.

Idempotency belongs in this same path. For hard writes, accept a client-generated key from the state machine and persist outcomes so retries do not double-apply.

What prompt injection does to an unsandboxed agent

Any content the agent reads — tickets, emails, PDFs, web pages — can contain instructions. OWASP LLM01:2025 Prompt Injection is the entry. The community write-up is the same idea in attack language: system text and user text share one token stream.

Sandboxing does not solve injection. Without a sandbox, injection has a bigger blast radius. With a sandbox, injection can still ruin a summary. It should not be able to send the inbox.

ControlWhat it stopsWhat it does not stop
Tool allowlist that cannot grow from retrieved text“Please enable email.send” in a PDFA send tool you already allowed
Instruction channel ≠ document channelConcatenating a ticket as system policyThe model still reads the ticket
Strip requests for new tools or secretsCatalog mutation via contentA legal tool called with a bad arg
Runner-side policy on argsSSRF, wrong recipient domain, DROP SQLA well-formed bad send you allowed
Human gate on irreversibleInjected “send now”A gated payload a tired human rubber-stamps
Redaction in tracesSecret echo into SlackThe model seeing the secret in the first place

Treat untrusted text as data, not as system policy. Practical runner checks that belong next to the allowlist:

  • Email fields must match an allowlisted recipient domain for outbound drafts.
  • SQL tools — if you ever allow them — reject DROP / UPDATE / DELETE after a real parse, not a regex vibe.
  • URL fetchers block link-local, metadata, and private IP ranges. OWASP A10:2021 SSRF is the category. The SSRF Prevention Cheat Sheet is the allowlist-not-denylist rule: positive allow list for scheme, port, and destination; do not “block the bad IPs” and call it done.

“Read the webpage” is a powerful tool and a common exfiltration path under injection. If you need it, constrain egress: allowlisted hosts, size limits, content-type checks, no redirect-following into RFC1918.

Injection that cannot call a write tool is noise. That is the whole point of the cage.

What to log for every tool call

For ops and forensics you want a span you can replay without leaking a customer. Redaction is part of the sandbox design, not an afterthought. Agents increase the number of places secrets can leak: prompts, traces, tool args, “helpful” plugin logs.

FieldKeepRedact or drop
tool_nameAlways—
run_id / state_name / job_typeAlways—
side_effect_classAlways—
identity_idWhich service accountToken material
args_redactedKeys + types + allowlisted valuesPII, secrets, raw email bodies
result_classok / typed errorFull downstream body
duration_msAlways—
cost_attributionModel + tool vendorCustomer identifiers in the label
cap_countersValues after increment—
actorHuman id on gated releases—

Checklist:

  • Default deny on new fields in the span
  • PII dictionary applied before the span leaves the runner
  • Slack / chat gets a link to the record, not the record
  • Export path assumes a future incident-response dump
  • Plugin “debug logging” is off in prod or pointed at a sink you own

You do not want full PII in a Slack channel. You do want to answer “which identity patched this lead at 14:03” without opening a laptop in a parking lot.

Environment separation: staging is not a costume

NIST SP 800-207 treats every session as untrusted until authenticated. Copy-pasting the prod token into a notebook “just to test” is the opposite of that. Promoting a tool from staging to prod is a change-controlled event: review side-effect class, caps, and whether the evaluator covers the new failure modes.

EnvironmentCredentialsWritesAudienceAllowed surprises
DevMocks / fixturesFakeEngineersBroken previews
StagingStaging systemsReal staging dataDomain reviewersWrong field, still internal
ProdLeast-privilege prodCaps + gatesCustomers / opsNone you have not named

Promotion checklist:

  • Side-effect class reviewed by the tool owner
  • Caps set in prod config, not inherited as “unlimited”
  • Dry-run path still callable for support
  • Evaluator covers the new failure modes
  • Staging token revoked from any laptop that touched prod
  • Break-glass documented with a clock

Dev and staging may share a catalog shape. They do not share identities. If a vendor has only one tenant, you do not have staging. You have a prod experiment. Gate harder.

Third-party MCP and plugin risk

Marketplace tools arrive with someone else’s threat model. MCP is a transport and an interface. The authorization spec is optional OAuth for HTTP transports. It does not automatically enforce least privilege or blast-radius caps. A wide-open MCP server is not a sandbox.

Before a server enters an agent catalog:

CheckPass looks likeFail looks like
Scopes requestedRead-only, one resourceadmin, offline_access, “all mailboxes”
Throwaway tenantYou ran it against fake dataFirst run is your CRM
Logging / exfilYou read the egress and the log sink“Helpful” telemetry you cannot disable
Version pinDigest or exact versionlatest
Tool-list mutationDisabled at runtimeServer can add tools after connect
Schema honestyTools match the docsHidden execute next to search

Spurlock Studios would rather wrap two HTTP endpoints you own than enable twenty plugins you have not read. The operating manual treats sandboxes as a first-class layer for that reason.

Decision list:

  1. Do you own the server code? Prefer yes.
  2. Can you pin and review the tool list on every connect? If no, do not connect.
  3. Can the server see secrets other than its own? If yes, split the catalog.
  4. Can it reach the public internet beyond an allowlist? If yes, treat it as a fetch tool and apply SSRF rules.
  5. Can it send email or move money? If yes, it is irreversible even if the vendor named it notify.

STDIO MCP that inherits the operator’s environment inherits the operator’s keys. That is a founder laptop with a protocol.

Progressive autonomy, then inventory

Climb the ladder per job type using online evaluator scores and incident count — not calendar time.

  1. Dry-run only
  2. Soft writes to internal fields
  3. Hard writes with a human gate
  4. Hard writes auto under caps
  5. Irreversible actions still gated (often forever)
RungWhat you measureKill switch
Dry-runPreview vs human intentStay here
Soft writeShadow disagreementDemote to dry-run
Gated hard writeReject reasons, SLA missesKeep the gate
Capped autoCap trips, incident countRe-gate that tool
IrreversibleResidual risk in writingDual control

Monthly inventory — every tool in every catalog:

  • Owner still at the company
  • Last used date; disable after a written idle window
  • Side-effect class still true (a “read” that grew a write is a new tool)
  • Credential age vs rotation calendar
  • Cap values still match the job
  • Staging and prod identities still separate

Orphan tools with live credentials are unpaid attackers waiting for a prompt injection. Disable them. Do not “leave them in case we need them.”

Failure mode: an API key with opinions

The failure is not theoretical. It is a specific shape I keep seeing across 500+ automations and 20,000+ hours on agentic systems: a team wires a model to a platform key, pastes a system prompt that says “be careful,” and calls it a pilot.

What breaks. The model — confused, injected, or just wrong — calls a write that was sitting on the same key as the read. A merge, a send, a delete, a refund. Caps were a conversation. The gate was a Slack DM. Staging was prod with a different URL in a comment.

What it costs. A customer-visible action you cannot recall, a credential rotation under fire, and a week of explaining why “the AI” had admin. The restore is never the hard part. The trust is.

What you do instead.

Anti-patternReplacement
“The agent has the Zapier key”Per-tool identities, allowlisted actions
Production credentials in the promptVault inject at the runner
Sandbox theater (permissions: YAML nobody enforces)Runner rejects with a typed error you can demo
Expanding scope mid-pilotOne job. New tools wait for the next slice
Shell “just for debugging” in prodPurpose-built tools with schemas
latest on a marketplace MCPPin, review, throwaway tenant first
Shared admin token across environmentsSeparate identities, break-glass with a clock
Prompt-only “do not delete”Delete is not in the catalog

Demo the reject path before you demo the happy path. If a vendor cannot show a rejected tool call, a cap trip, a dry-run, and a human gate, keep shopping.

How Spurlock Studios applies this in a pilot

In the $1,500 · 5-day pilot we pick the minimum tool set for one sentence-sized job. Reads first. Writes only if the job demands them, usually to an internal surface. Irreversible actions stay human-gated. You leave with a catalog and a runner you can keep operating.

A short example we actually scope:

  • Job: enrich a lead record with firmographics and draft an internal note.
  • Allowlist: crm.get_lead, enrichment.lookup, crm.patch_internal_note.
  • Not allowlist: crm.merge_leads, email.send, crm.delete.
  • Caps: one lead per run; enrichment API max 3 calls; patch only internal_note.
  • Evaluator: note must cite enrichment fields present in the tool result; no invented revenue numbers when enrichment returned null.

That is sandboxed tool use. “Here’s our admin API key, go enrich everything” is not.

Three locks sit together in the operating manual:

  1. State machine — only act may call tools with side effects.
  2. Sandbox — only allowlisted tools with caps.
  3. Evaluator — only passing artifacts proceed to hard writes.

Remove any lock and the system fails open in a different way. The sandbox is the lock this spoke owns.

Procurement RFP — ask vendors to demonstrate:

  • A rejected unknown-tool call
  • A blast-radius cap that stops a loop
  • A dry-run that does not write
  • A human gate on an irreversible class
  • Separate staging and prod identities
  • Redacted traces you can export

If they can only show a happy-path demo, you are buying an API key with opinions. Map and offer: /agentic.

FAQ

What are sandboxed AI tools?

Sandboxed AI tools are agent-callable functions wrapped in allowlists, least-privilege credentials, quantitative caps, and policies for irreversible actions. The model proposes a call. The runner enforces whether that call is legal. A description in a prompt is not a sandbox.

How do you implement safe tool use for agents?

Define a typed tool catalog, classify side effects, scope credentials per tool, and enforce caps and timeouts in the runner. Start irreversible actions behind human gates, fail closed on errors, and log redacted traces. Promote autonomy only after evaluation scores hold on that job — not after a quiet week.

Is MCP enough to be sandboxed?

No. MCP is a transport and interface pattern. Authorization in the spec is optional OAuth for HTTP. It does not automatically enforce least privilege or blast-radius caps. You still design the server’s capabilities, auth, and policy. A wide-open MCP server is not a sandbox.

Should agents have shell access?

Almost never in a business pilot. If you need shell for a devtools agent, isolate the environment, drop privileges, restrict network, and treat every command as high risk. Default to purpose-built tools with schemas. A shell is excessive functionality with a blinking cursor.

When can we remove human gates?

When offline and online evaluator scores meet the bar, blast-radius caps are proven under load, and the business accepts the residual risk in writing. Remove the gate per tool and per job type. Gates are a control, not an insult to the model. Irreversible sends and deletes often stay gated forever.

How does this relate to Spurlock Studios’ agentic offer?

Sandbox design is part of every pilot and build. We would rather ship a narrow, caged agent that works than a broad agent that can email your customers by accident. The five-day, $1,500 pilot on /agentic is the on-ramp; book via /contact?intent=agentic-pilot.

CTA

An agent that can act without a cage is an API key that talks. Build the catalog, the identities, the caps, and the gate — then let it earn a write.

Read the Agentic Systems Operating Manual, then use the agentic lane or book the pilot to put one job in a cage you can demo.

FAQ

What questions does this article answer?

What are sandboxed AI tools?
Sandboxed AI tools are agent-callable functions wrapped in allowlists, least-privilege credentials, quantitative caps, and policies for irreversible actions. The model proposes a call. The runner enforces whether that call is legal. A description in a prompt is not a sandbox.
How do you implement safe tool use for agents?
Define a typed tool catalog, classify side effects, scope credentials per tool, and enforce caps and timeouts in the runner. Start irreversible actions behind human gates, fail closed on errors, and log redacted traces. Promote autonomy only after evaluation scores hold on that job — not after a quiet week.
Is MCP enough to be sandboxed?
No. MCP is a transport and interface pattern. Authorization in the spec is optional OAuth for HTTP. It does not automatically enforce least privilege or blast-radius caps. You still design the server’s capabilities, auth, and policy. A wide-open MCP server is not a sandbox.
Should agents have shell access?
Almost never in a business pilot. If you need shell for a devtools agent, isolate the environment, drop privileges, restrict network, and treat every command as high risk. Default to purpose-built tools with schemas. A shell is excessive functionality with a blinking cursor.
When can we remove human gates?
When offline and online evaluator scores meet the bar, blast-radius caps are proven under load, and the business accepts the residual risk in writing. Remove the gate per tool and per job type. Gates are a control, not an insult to the model. Irreversible sends and deletes often stay gated forever.
How does this relate to Spurlock Studios’ agentic offer?
Sandbox design is part of every pilot and build. We would rather ship a narrow, caged agent that works than a broad agent that can email your customers by accident. The five-day, $1,500 pilot on [/agentic](/agentic) is the on-ramp; book via [/contact?intent=agentic-pilot](/contact?intent=agentic-pilot).
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot