Spurlock Studios
Contact
Share LinkedIn X
A violet ring. Thesis: AGENTIC SYSTEMS OPERATING MANUAL MULTI.

An agentic system is not a chatbot with plugins. It is a production machine that plans, calls tools, checks its own work against criteria you defined, and stops when it should stop. If you cannot name the evaluator, the sandbox boundary, the state machine, and the kill switch, you do not have an agentic system. You have a demo.

This manual is how Spurlock Studios builds agentic work that founders and technical buyers can put on real data. It is the parent piece for the agentic lane. The spokes go deep on evaluators, policy gates, failed-tool loops, sandboxes, and the metrics that lie. Read this for the map. Use the spokes when one layer is the risk.

The short answer

  • An agentic system chooses steps under uncertainty, writes to systems you care about, and is judged by something other than the worker that produced the artifact.
  • Build the evaluator before the agent. Then put a policy gate in front of every side effect. Then cage the loop so failed tools cannot retry forever.
  • Pass rate alone will green-light a grind. Gate deploys on revision rate, trajectory, coverage, and cost per success — see why pass rate lies.
  • Default to automation when the path is known. Build an agent only when the path varies and you can still write pass/fail criteria. The brake pedal is when not to build an agent.
  • Prove one sentence-sized job in five days on your data. The on-ramp is /agentic.

What is an agentic system?

An agentic system is software that can choose steps toward a goal, use tools to change the world outside the model, and revise its path when evidence says the last step failed — under constraints you own.

Three properties separate it from a scripted automation:

  1. Choice under uncertainty. The system picks the next action from a set of allowed tools and states, not from a fixed graph of “always do A then B.”
  2. External effects. It can read and write systems you care about: tickets, CRMs, inboxes, code, calendars, knowledge stores.
  3. Judgement that is not the worker. Something other than the same context that produced the artifact decides whether the artifact is acceptable.

n8n fits here as the rail for deterministic glue — webhooks, queues, retries, human approvals — while models and tool runners sit inside bounded steps. The rail is boring on purpose. The agent lives in the steps where choice is required; the rail owns delivery, idempotency, and escalation.

If your “agent” is a single prompt that calls three APIs and always returns green, call it an automation. Language matters because budgets, risk reviews, and success metrics change when you admit you are shipping non-deterministic software.

Anthropic’s own engineering note on building effective agents lands in the same place: start simple, measure, and add multi-step agentic loops only when a cheaper pattern fails. That is not a slogan. It is the cheapest way to avoid a fleet you cannot operate.

You have…Call it…Primary control
Fixed path, rare judgementAutomationWorkflow tests + retries
Varying path, crisp criteriaAgentic systemEvaluator + gates + state machine
Varying path, mushy tasteWorkshop, not a buildCriteria first, then decide
High stakes, weak recoveryHuman + checklistDo not auto-write

What operating stack has to exist?

Every production agentic system at Spurlock Studios is built from the same stack. Skip a layer and you will pay for it in production, usually on a Tuesday.

LayerJobFailure mode if missing
Job contractOne sentence goal + acceptance criteriaInfinite scope, unmeasurable demos
EvaluatorIndependent pass/fail with evidenceSelf-grading theater
Policy gateAllow / deny / pending before side effectsPrompt-only “safety”
Tool sandboxAllowed actions, secrets, blast radiusAgents that email customers or delete rows
State machineExplicit states and transitionsLoops that never halt, duplicate writes
Memory policyWhat persists, what dies with the runContaminated context
Retrieval contractWhat may be cited as factRAG that invents policy
Handoff protocolWhat moves between agentsLost context, double work
Cost + kill switchesBudgets, caps, abortSurprise invoices
ObservabilityTraces, scores, operator dashboardYou cannot debug or trust it

You can implement these in different stacks. The stack is not the product. The contracts are.

NIST’s AI Risk Management Framework frames the same work as Govern, Map, Measure, and Manage. The AI RMF 1.0 is voluntary and use-case agnostic; it will not write your job contract. It will tell a buyer why “we shipped a chat UI” is not a risk program. The RMF Core is the part that maps onto this manual: you govern who can widen tools, you map what the job can break, you measure with an evaluator, and you manage abort and escalate as first-class states.

  • Job contract written with the buyer, including hard nos
  • Evaluator returns structured verdicts with evidence
  • Policy gate runs in code, fail-closed, before every write
  • Tool allowlist and scoped credentials
  • Terminal states: done, escalate, abort
  • Per-run budget and a kill switch you have actually tripped in staging
  • Trace fields an operator will read without opening a JSON dump

Why evaluators before agents?

Build the evaluator before the agent. That sentence is the whole strategy.

The evaluator is a separate component whose only job is to judge an artifact against criteria. It must not see the worker’s chain of thought. It must return a structured verdict: pass or fail, which criterion failed, evidence, and a next action when fail is recoverable.

Mechanical checks first. Schema validity, required fields, unit tests, allowlisted URLs, “ticket status is one of these enums,” “invoice total matches line items.” Models judge only what genuinely needs judgement: tone for a customer email, whether a summary omitted a material risk, whether a research brief answered the asked question.

Without an evaluator you are optimizing prompts in the dark. With one, every model swap, tool change, and prompt edit becomes a measured experiment.

OpenAI’s own evaluation best practices say the same thing in vendor language: write scoped tests early, make them task-specific, log everything, and automate scoring when you can. Their agent evals guide is explicit that a trace — model calls, tool calls, guardrails, handoffs — is what you grade when the unit of work is a loop, not a single completion. Working with evals is the API-shaped version: a dataset plus graders, not a vibe check after a demo.

Vendor eval products move. As of August 2026, OpenAI has published a deprecation window for its standalone Evals platform. That is a reason to own the golden set and the grader contract in your repo, not a reason to skip evaluation. The suite has to survive the logo on the dashboard.

Deep dive: Build the Evaluator Before the Agent.

Check typeOwnerExample
SchemaCodeJSON matches the contract
Business ruleCodeTotals equal line items
Citation ruleCode + light modelURL present or explicit no-match
JudgementModel evaluatorDid the brief answer the asked question
SafetyCode + allowlistsBanned promises, PII patterns

Pass rate is a vanity metric

A 94% pass rate can hide a system that revised four times, called the wrong tool, skipped half the real job shapes, and still needed a human rewrite. “Pass” is usually a thin binary on the final artifact. It ignores trajectory, coverage, and unit cost.

Gate deploys on a panel, not a single percentage:

MetricWhat it catchesDeploy veto if…
Golden-set pass rateObvious quality regressionsDrops past the agreed band
Revision rateGrind-to-greenMean revisions climb while pass holds
Trajectory scoreWrong tools, extra stepsTool-choice errors rise
Eval coverageUntested job shapesNew production clusters have no cases
Cost per successExpensive “wins”Dollars per passing run leave the band
Online / offline gapSilent production driftSampled live scores diverge from the suite
Escalate rateHidden human loadHumans become the real runtime

OpenAI’s agent-eval docs treat tool choice and handoff timing as first-class grades, not footnotes. If you only score the final blob, you will ship agents that wander and still “pass.” The spoke that owns the panel is why pass rate lies.

Do not celebrate latency alone. Fast wrong is still wrong. Do not celebrate a pass-rate jump after you loosened criteria. Version the evaluator. A criteria change is a release, not a rounding error.

How do pre-execution policy gates work?

A kill switch that lives in the system prompt is a suggestion. A pre-execution policy gate is a function in your runtime that sees the concrete tool name and arguments and returns allow, deny, or pending-approval before the tool runs. If the policy service is down, you fail closed.

That is the difference between “we told the model not to refund” and “the refund tool never fired.”

DecisionWhenWhat the trace must show
allowPayload matches policyTool, args hash, rule id
denyOut of allowlist, over cap, bad tenantReason code, no side effect
pending-approvalIrreversible or novel classQueue id, proposed payload
Fail closedPolicy timeout, unknown tool, bad parseAbort, no retry-as-allow

Gates sit after the model proposes and before the sandbox executes. Sandboxes limit damage if a call gets through. Gates decide whether the call happens at all. You need both.

  • Unknown tools cannot execute
  • Write tools require a tenant id that matches the run
  • Row / recipient / dollar caps enforced in code
  • Policy outage cannot be bypassed by the worker
  • Deny and pending leave a reason code an operator can filter

Failed-tool loops

Most “the agent got stuck” incidents are not mysterious. The tool returned an error, the model treated the error as more story, and the loop called the same tool with the same args until the budget died. The spoke that owns this failure is why agents loop on failed tools.

The operating fix is typed errors plus illegal transitions.

  1. Tool adapters return a typed result: ok, retryable, fatal, auth, not_found, rate_limited.
  2. The state machine, not the model, decides the next state.
  3. The same (tool, args_hash, error_class) pair cannot fire twice in one run.
  4. retryable gets a bounded backoff on the rail (n8n, queue, whatever you already trust).
  5. fatal and auth go to escalate or abort with the trace attached.
  6. The evaluator sees the tool outcome. A worker that “summarizes past” a failed write does not get to mark done.
Error classLegal next stateIllegal next state
okevaluateSilent second write
retryableact once more, then escalateInfinite retry
rate_limitedWait on the railImmediate re-call
authescalateGuess a new token
not_foundRevise query once, then escalateInvent the record
fatalabort“Try a different tool that deletes”

If you cannot draw that table for your tools, you do not have a loop. You have a hope.

How do you sandbox tool use?

An agent without a sandbox is a liability with an API key.

Sandbox means:

  • Allowlist of tools, not “whatever the model invents.”
  • Scoped credentials — read-only where possible, write scopes only for the tools that must write.
  • Blast-radius limits — rate caps, row caps, recipient caps, environment isolation (staging vs production).
  • Dry-run modes for first contact with a new tool.
  • Human gates on irreversible actions until the evaluator and error rates earn autonomy.

Anthropic’s computer-use documentation is unusually blunt for a vendor: dedicated VM or container, no sensitive logins in the environment, domain allowlists, and a human confirm on consequential actions. Their computer-use research note names the reason: screenshot and page content can carry prompt injection that overrides your instructions. If that is true for a desktop sandbox, it is true for a CRM tool that reads ticket text.

MCP servers and custom tool runners are fine. The Model Context Protocol is a way to expose tools and context over a standard. The 2026-07-28 spec even pushed toward stateless, self-contained requests so you can load-balance without a shared session store. That is transport. It is not a permission system. Unrestricted shell, unrestricted email send, and “admin” CRM tokens are still not fine for a pilot.

Deep dive: Sandboxed Tool Use.

Tool classPilot defaultAutonomy earned when
Read ticket / CRMAllow, scopedAlways, with redaction
Write internal draftAllow to internal fieldOnline scores hold
Send customer emailHuman gateRewrite rate and silent-fail samples stay low
Refund / billing changeDeny or pendingAlmost never in week one
Shell / code execIsolated runner or denyJob actually needs it

Where do state machines belong?

Agent loops need freedom inside a cage. The cage is a state machine.

Typical states for a business agent: intake → plan → act → evaluate → revise → done | escalate | abort. Transitions are explicit. Side effects only happen in act, and only after the policy gate. Evaluation never mutates production systems. Revision has a ceiling (usually three). Escalation packages the full trace for a human.

n8n is a natural home for the cage: each state can be a node or sub-workflow, with durable execution, retries, and a dead path for escalate. The model proposes; the machine decides whether the transition is legal.

If you run n8n in production, treat durability as part of the agent, not as hosting trivia. n8n’s durable scheduler exists so time-based work survives restarts and does not double-fire across instances. Queue mode is how production executions leave the editor process and land on workers. An agent that “works in the canvas” and vanishes on deploy is not an agentic system. It is a local demo with extra steps.

LangGraph’s persistence model is the same idea with different nouns. Checkpointers snapshot thread state so you can pause for a human, resume after a crash, and avoid re-running work that already succeeded. Persistence splits that short-term thread memory from long-term stores (preferences, facts). If you dump both into one transcript, you will re-inject failed reasoning as if it were policy.

StateMay write?May call model?Exit condition
intakeNoClassify onlyIn-scope contract attached
planNoYesTool plan within allowlist
actYes, after gateYesTyped tool result
evaluateNoJudge onlyVerdict + evidence
reviseNoYes, cappedNew artifact or ceiling
doneConfirm onlyNoEvaluator passed
escalateNoNoHuman package stored
abortNoNoKill reason stored

How should memory, retrieval, and handoffs be contracted?

Agent memory is not “stuff the whole transcript into the next call.” Memory is a policy. The policy itself is agent memory patterns.

Separate at least four stores:

  1. Ephemeral run context — dies when the run ends.
  2. Working scratch — intermediate artifacts for this job only.
  3. Durable facts — customer prefs, account IDs, approved SOPs — with ownership and TTL.
  4. Run history / traces — for ops and learning, not for raw re-injection into every prompt.

Persist preferences and identifiers. Forget raw intermediate reasoning. Never let a failed run’s bad conclusions become long-term “memory” without a promotion rule.

Retrieval-augmented generation fails in businesses for a boring reason: teams treat “retrieved” as “true.” Retrieval is a search result. Truth is a contract.

A retrieval contract answers:

  • Which corpora are authoritative for which question types?
  • What freshness rules apply?
  • Must citations be present for any factual claim?
  • What happens when retrieval returns nothing — refuse, ask, or fall back to a human?
  • How do you detect contradiction across chunks?

If the agent can invent policy when the index is empty, you do not have RAG. You have a confident liar with a vector database.

Multiple agents are useful when jobs naturally split: research vs draft vs compliance check; intake vs enrichment vs write-back. They are harmful when you multiply agents to look sophisticated.

A handoff is a typed package:

  • Goal and constraints
  • Artifacts produced so far
  • Open questions
  • Tools already tried and outcomes
  • Budget remaining
  • Evaluator criteria still unmet

Do not pass “the vibe.” Pass the package. The receiving agent should not need the sending agent’s private scratch.

StoreSurvives the run?May enter the next prompt?
Ephemeral contextNoThis run only
Working scratchNoThis run only, redacted
Durable factsYes, with TTLYes, if schema-valid
TracesYesNo — ops only
Retrieved chunksPer queryOnly with citation or no-match

How do you evaluate agents without fooling yourself?

Evaluation is a product discipline, not a vibe check after a demo.

Unit-level

  • Tool adapters: given fixture inputs, do they return typed outputs or typed errors?
  • Retrievers: precision/recall on a labeled query set for your corpus.
  • Schemas: every agent-facing JSON shape validates.
  • Policy gate: unknown tool, over-cap payload, and missing tenant all deny.

Task-level

Build a golden set of 30–100 real jobs (anonymized if needed). For each: input, required artifacts, pass criteria, known traps. Run the suite on every change that could affect behavior. Track pass rate, average revisions, cost per pass, escalate rate, and coverage of the job shapes you actually see.

Online

Sample production runs. Score with the same evaluator. Alert when online scores drift from offline. Drift is how quiet failures start.

What not to measure alone

Latency and token count without quality. “User thumbs up” without criteria. Self-reported confidence from the worker. A pass rate with no revision or coverage number next to it.

OpenAI’s eval guidance is useful here even if you never touch their dashboard: collect cases from production logs, keep humans in the loop to calibrate automated graders, and treat evaluation as continuous. That last point is the one teams skip. A golden set that never gains a case after an incident is a museum.

LayerQuestion it answersCadence
UnitDid this adapter lie?Every commit
Task / golden setDid this change regress the job?Every behavior change
TrajectoryDid it take a sane path?Every behavior change
Online sampleIs production drifting?Daily or weekly
Silent-fail sampleDid a “pass” later get rewritten?Weekly, painful, worth it

When should you not build an agent?

Default to automation when the path is known, the inputs are structured, and judgement is rare. Default to a human when stakes are high and criteria are contested. Build an agent when the path varies, tools are many, and you can still write acceptance criteria crisp enough to evaluate.

Skip agents (for now) when:

  1. The path is fully known. Same steps, same systems, rare exceptions.
  2. You cannot write pass/fail criteria. “Good” is still an argument.
  3. Stakes are high and recovery is hard, and you do not yet have gates and sandboxes.
  4. Credentials and data access are political. You will spend the month on access, not learning.
  5. Volume is tiny. Ten items a month may want a human and a template.
  6. Nobody owns the SOP. The agent becomes a scapegoat.
  7. You want a demo more than a metric.

Deep dive: When Not to Build an Agent.

SignalBuild this instead
Known path, structured I/On8n / automation
Criteria are mushProcess workshop
One irreversible writeHuman + checklist
Need a public chatbot, no criteriaWrong lane
Path varies, criteria existAgentic system

What multi-agent shape ships for a business?

Here is a shape that ships for small and mid-size teams without becoming a research project.

Roles

  • Router / intake — classifies the job, attaches the job contract, rejects out-of-scope work.
  • Worker — plans and acts inside the sandbox.
  • Evaluator — independent judgement; no tool writes.
  • Librarian (optional) — retrieval only; returns citations or “no hit.”
  • Operator surface — humans approve, abort, or re-scope.

Control flow

  1. Event or human request hits intake (often via n8n webhook).
  2. Job contract loaded; budget and tool allowlist attached.
  3. Worker enters plan → act loop under the state machine.
  4. Policy gate runs on every proposed write.
  5. After each material artifact, evaluator runs.
  6. Fail → revise until ceiling → escalate.
  7. Pass → write-back through allowlisted tools → done.
  8. Trace + cost + scores stored for ops.

What “done” means

Done is not “the model said done.” Done is: evaluator passed, side effects confirmed idempotently, and the run landed in a terminal state with a receipt the operator can audit.

Anthropic’s effective-agents writeup keeps repeating a useful constraint: successful teams were not the ones with the most elaborate graphs. They were the ones who could see the plan, keep the tool interface tight, and add complexity only when a simpler pattern failed. A business fleet that starts as intake + worker + evaluator is not “behind.” It is honest.

TemptationCostDo this first
Five specialists on day oneLost handoffs, no ownerOne worker, one evaluator
Shared chat as memoryContaminated contextTyped handoff package
Evaluator with write toolsSelf-dealingRead-only judge
“Latest” model as defaultSilent quality breaksPin, then re-run the suite

What does a five-day pilot actually prove?

Most teams do not need a twelve-week “AI transformation.” They need one narrow job proven on their data.

The Spurlock Studios agentic pilot is $1,500 · 5 days. One job, scoped tight enough to finish in a week. A working agent on your real data — not a slide deck. You keep it either way. The $1,500 credits toward a full build.

What you leave with:

  • A runnable agent for one sentence-sized job
  • An evaluator with explicit criteria
  • Sandboxed tools and a thin policy gate for that job
  • A state machine with a revision ceiling and an escalate path
  • A short build quote based on what we actually saw

What the week is not: a chatbot skin, a multi-agent org chart, or a promise that pass rate will hold after you add refunds and production sends.

Start on /agentic or go straight to /contact?intent=agentic-pilot.

DayOutcome you can point at
1Job contract, hard nos, evaluator shape
2Golden set v0 (real cases, including traps)
3Sandboxed tools + dry-run writes
4Loop + gate + escalate on your data
5Scores, cost, quote, keep-the-agent handoff

When the problem is architecture across a roadmap rather than a single agent, that is a different engagement shape — same principles, longer surface. The pilot still comes first because a roadmap without one green job is a slide.

What does a production walkthrough look like?

Job contract: “Given a new support ticket, classify severity, draft an internal summary with citations from the help center, and propose a reply — never send.”

Evaluator criteria (examples):

  • Severity is one of P1|P2|P3|P4
  • Summary includes at least one citation URL from retrieval or explicitly says “no doc match”
  • Proposed reply contains no promise of refund or SLA change unless those strings appear in retrieved policy
  • Schema validates

Tools in sandbox: ticket read API, help-center retriever, draft write to internal field. Not in sandbox: send reply, issue refund, change billing.

Policy gate: deny send_reply and issue_refund by name. Cap retrieval calls. Require tenant_id on every tool payload.

State machine: intake → retrieve → draft → evaluate → revise (max 3) → escalate or done.

Memory: customer ID and prior ticket IDs may persist; raw model scratch does not.

Cost: hard cap on retrieval calls and revisions; abort to human queue if exceeded.

That system is agentic. A Zap that posts “new ticket” into Slack is not. Both can be valuable. Only one needs this manual.

StepWho decidesWhat gets written
IntakeRouter + contractNothing external
RetrieveLibrarian / workerNothing external
DraftWorkerInternal field only
EvaluateEvaluatorVerdict record
ReviseWorker, if ceiling remainsInternal field overwrite
EscalateState machineHuman package
SendHuman, laterCustomer channel

What fails after the demo?

Agent theater. Fancy UI, no evaluator, no sandbox, no budget. Demo day works. Week three does not.

Prompt as policy. Rules living only in natural language. Policies belong in code checks and allowlists; language fills gaps. OWASP’s LLM Top 10 still leads with prompt injection for a reason: the model will follow instructions found in content. The community writeup on prompt injection is the plain-language version. The OWASP GenAI project is the living index. None of those pages say “add a nicer system prompt” as the whole fix.

Unbounded loops. No revision ceiling. No typed tool errors. Cost and chaos grow together.

RAG without refuse. Empty retrieval still produces “facts.”

Too many agents too early. Three agents before one job is green. Split only after the single-worker path is measured.

No human path. Escalation is a first-class state, not an apology.

Pass-rate theater. One green percentage, no revision or coverage number, no online sample.

Shared enrichment keys. One credential that can see every tenant. That is a breach design, not a shortcut.

FailureWhat it costsWhat you do instead
Self-grading workerSilent wrongnessIndependent evaluator
Prompt-only denyInjection bypassPolicy gate, fail closed
Retry-on-text-errorToken burn, duplicate writesTyped errors + illegal transitions
Empty-index answersInvented policyRefuse or escalate
Latest-model defaultUnexplained score dropsPin, re-run golden set

What should ops actually read?

If the only “observability” is a provider dashboard, you will not catch silent wrongness. Traces must show: state, tool calls, inputs/outputs (redacted), gate decisions, evaluator verdicts, cost, latency, and escalation reason. Scores from your evaluator suite should land on a dashboard a human checks weekly — not a graveyard of JSON in object storage.

OpenTelemetry now keeps GenAI conventions in a dedicated repo: semantic-conventions-genai. The older opentelemetry.io GenAI pages redirect there. You do not have to export every prompt. You do have to agree on span names and attributes so “what tool fired, for which tenant, at what cost” is not a custom folklore per engineer. The attribute registry is the boring list that makes that possible.

Token spend is a product feature. Treat it like one.

Per-run budgets, per-day budgets, max tool calls, max revisions, model tiers by state (plan on a cheaper model, evaluate on a stricter one when needed), and hard kill switches when spend or error rate crosses a line. Log cost on every transition. Ops should see dollars next to failure rates.

CapTypical pilot defaultWhat it stops
Max tool calls / runLow double digitsWandering tool spam
Max revisions3Grind-to-green
Per-run spendSet with the buyerOne runaway loop
Per-day spendSet with the buyerQuiet overnight burn
Error-rate killTrip after repeated fatal / authRetry storms

A kill switch you have never tripped in staging is a rumor. Force a budget abort on purpose before the first soft launch. The trace should show abort and a reason code, not a hung act.

Human gates fail when every run waits on a busy founder. Design queues:

  • Batch review for soft writes (internal notes) twice a day
  • Immediate review only for irreversible classes
  • Auto-promote when online pass rate holds for a defined window on that job type
  • Spot checks forever — autonomy is not absence of audit

The operator surface should show the same trace fields ops already use: criteria failures, cost, and the proposed write payload. Asking a human to re-read the whole chat is how gates get muted.

RoleOwnsWeekly artifact
Job ownerCriteria and riskCriteria changelog
Systems ownerCredentials, schemas, capsAccess review
Agent engineerPrompts, tools, state machineSuite diff
Ops reviewerScores and incidentsDrift + silent-fail sample

One person can wear multiple hats at a small company. Zero people wearing the ops hat is how silent failure becomes culture.

Tenancy and injection

If more than one customer or department shares infrastructure, tenancy is an agent feature. Every run carries tenant_id. Tool credentials are bound to that tenant. Retrieval ACLs filter before ranking. Memory keys are prefixed. Logs are partitioned. A “shared enrichment key” that can see every CRM is a data-breach design.

Prompt injection is a tenancy problem too. Ticket text, email bodies, and retrieved pages are untrusted. OWASP treats that as LLM01 for a reason. Content from Tenant A must never expand tools or memory for Tenant B. Sandboxes and allowlists are the first wall. Evaluator checks for cross-tenant identifiers in artifacts are a useful second wall. The policy gate is the wall that actually stops the write.

BoundaryEnforce inFail mode if skipped
CredentialPer-tenant secret, not a shared admin tokenCross-tenant read
RetrievalACL filter before rankPolicy leak across accounts
Memory key{tenant_id}:{entity}Contaminated prefs
Tool argsGate rejects foreign idsWrite to the wrong account
LogsPartition + redactionOne export holds everyone
  • No tool credential can see more than one tenant
  • Retrieval returns zero hits rather than another tenant’s doc
  • Memory promotion requires a tenant-scoped schema
  • Gate denies payloads whose ids do not match the run
  • Operator exports cannot dump another tenant’s traces by default

This is not “enterprise later.” It is the minimum if two teams share a runtime. Agencies already know the version of this story from n8n: isolate credentials or inherit incidents.

Pin models against drift

“Latest” as a default is an availability choice that often breaks quality silently. Pin the model id on every call. When the provider ships a new version, run the golden set in a side-by-side before you promote. A pass-rate bump after a silent upgrade is not a win until you know which criteria moved.

Drift shows up in three places:

  1. Provider drift — the pinned id still exists, behavior changed, or the pin was ignored.
  2. Prompt drift — someone edited the worker or the evaluator and did not version it.
  3. World drift — the job changed (new ticket types, new policy docs) and the golden set did not.

Treat each as a release. OpenAI’s eval docs call this continuous evaluation: grow the set from production logs, keep humans calibrating the grader. You do not need their product to do that. You need a suite that fails CI when scores leave the band.

ChangeRequired before promote
New model idFull golden set + cost band
Prompt or tool schemaGolden set + trajectory slice
Evaluator criteriaVersion bump + note that pass rate is not comparable
New production failureNew case within 48 hours
Widened tool allowlistDeny-path tests for the new tool

If you cannot say which model id wrote last week’s traces, you cannot debug last week’s traces.

Idempotent writes

Agents retry. Rails retry. Humans click twice. If act is not idempotent, a “successful” run can create two tickets, two drafts, or two charges. The state machine can forbid a second transition and still lose if the first write’s ack never came back.

Give every side effect an idempotency key derived from run_id + tool + args_hash (or a business key the downstream already understands). The tool adapter must be safe to call twice with the same key. The evaluator must not treat “record already exists with this key” as a hard fail if the payload matches.

WriteKeySafe retry looks like
Internal draftrun_id:draftOverwrite same field
Ticket commentrun_id:commentSame comment id returned
CRM noterun_id:note or external idNo second note
Email sendDo not auto-retryHuman or outbox with key
Charge / refundDownstream idempotency keyOne money movement

Dry-run modes should exercise the key path, not only the happy path. A sandbox that cannot show you a duplicate-suppressed write is a sandbox that will surprise you on the first timeout.

What order do you build in?

  1. Write the job contract and acceptance criteria with the buyer.
  2. Build the evaluator and a tiny golden set.
  3. Implement tools behind a sandbox with dry-run.
  4. Put the policy gate in front of every write. Fail closed.
  5. Wire the state machine (n8n or equivalent) with budgets, typed errors, and escalate.
  6. Add retrieval and memory only if the job needs them — with contracts.
  7. Run the golden set until pass rate and revision rate and cost are acceptable.
  8. Soft-launch with human gates on writes.
  9. Widen autonomy only when online scores hold.
  10. Split agents only after the single-worker path is measured.

Skipping to step 7 because a vendor demo looked good is how you buy regret.

Treat each layer as a module with an owner and a test. The evaluator exports judge(artifact, criteria) -> Verdict. The sandbox exports callTool(name, args, ctx) -> Result. The gate exports decide(tool, args, ctx) -> allow|deny|pending. The state machine exports transition(state, event, ctx) -> State. Memory and RAG export read/write functions with schemas. Observability wraps all of the above.

Integration tests should freeze a run through intake to terminal with fixture tools. Contract tests should freeze golden-set scores on CI. Load tests should freeze budget trips. You do not need a research lab. You need the same hygiene you already use for payments and auth.

Document the hard nos in the same repo as the code. Hard nos that live only in Slack will be rediscovered after an incident.

When a vendor sells you “agents,” ask:

  1. Show the evaluator on our sample cases, not yours.
  2. Show the tool allowlist and how new tools are added.
  3. Show the policy gate deny a real payload.
  4. Show the state machine or equivalent control flow.
  5. Show per-run budgets and a kill switch demo.
  6. Show a trace with redaction.
  7. Show what happens on empty retrieval.
  8. Show what happens when the same tool fails twice.
  9. Show who owns prompts after go-live.
  10. Show exit: can we export and run without you?

If answers are slides without receipts, you are buying theater.

WordMeaning in this manual
Job contractGoal, audience, criteria, hard nos
EvaluatorIndependent verdict with evidence
Policy gateAllow / deny / pending before side effects
SandboxAllowlisted tools + caps + least privilege
State machineLegal transitions and terminals
Handoff packageTyped relay between agents or humans
Kill switchAutomatic stop on spend, error, or policy
Golden setLabeled jobs for regression

Use the words precisely. Language drift recreates agent theater under new names.

Founders and technical buyers who need work done — triage, research briefs, enrichment, internal ops agents, content drafts with hard constraints — and who will not accept “trust the model.” If you want a public chatbot with no criteria, this is the wrong lane.

Spurlock Studios ships agentic systems with explicit state machines, sandboxed tool runners, and reflection loops that self-correct. Builds typically land in 2 to 10 weeks after a pilot proves the job. That range is packaging, not a promise that your CRM is a two-week problem.

I have spent 20,000+ hours architecting agentic systems and shipped 500+ automations. The hours do not make the stack optional. They are why I refuse to start at the chatbot.

Read the spokes in the order your risk demands. Most teams should start with evaluators, then sandboxes and policy gates, then prove the job. Come back to this manual when you need the full map.

Cluster map

This manual is the hub. Use the spoke that matches the failure, not the one with the trendiest demo.

Evaluate before you scale

Control the tools

Run it in production

FAQ

What is an agentic system in plain terms?

An agentic system is software that can choose tools and steps toward a goal, change external systems, and revise when checks fail — under budgets and rules you define. It is not a chat UI. The difference from automation is meaningful choice under uncertainty plus independent evaluation.

How do you evaluate AI agents without fooling yourself?

Separate the evaluator from the worker. Use mechanical checks first, then model judgement only where needed. Maintain a golden set of real jobs and run it on every meaningful change. Track pass rate, revisions, cost per pass, coverage, and escalate rate. Never trust the worker’s self-score as the primary metric.

When should we not build an agent?

Skip the agent when the path is already known, when you cannot write pass/fail criteria, or when the honest fix is a checklist and a webhook. High-stakes writes without gates and sandboxes are another no. If you are unsure, start with automation or a scoping conversation — the longer argument is when not to build an agent.

How do we stop agents from doing dangerous things?

Allowlist tools, scope credentials, cap blast radius, and put a policy gate in front of every side effect. Require human approval for irreversible actions until scores earn autonomy. Put kill switches on spend and error rate. Sandboxes and gates are not optional for production tool use.

How much does an agentic pilot cost at Spurlock Studios?

The pilot is $1,500 for five business days: one narrow job on your real data, a working agent you keep, and a build quote based on what we saw. Details and packaging live on /agentic.

What is the difference between an agent and an automation?

Automation follows a known path with rare judgement. An agent chooses among tools and paths under uncertainty and must be evaluated. If you can draw the flowchart completely, you probably want automation. If the path varies but criteria are clear, you may want an agent.

CTA

Ready to prove one job in five days? /agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

What is an agentic system in plain terms?
An agentic system is software that can choose tools and steps toward a goal, change external systems, and revise when checks fail — under budgets and rules you define. It is not a chat UI. The difference from automation is meaningful choice under uncertainty plus independent evaluation.
How do you evaluate AI agents without fooling yourself?
Separate the evaluator from the worker. Use mechanical checks first, then model judgement only where needed. Maintain a golden set of real jobs and run it on every meaningful change. Track pass rate, revisions, cost per pass, coverage, and escalate rate. Never trust the worker’s self-score as the primary metric.
When should we not build an agent?
Skip the agent when the path is already known, when you cannot write pass/fail criteria, or when the honest fix is a checklist and a webhook. High-stakes writes without gates and sandboxes are another no. If you are unsure, start with automation or a scoping conversation — the longer argument is [when not to build an agent](/blog/when-not-to-build-an-agent).
How do we stop agents from doing dangerous things?
Allowlist tools, scope credentials, cap blast radius, and put a policy gate in front of every side effect. Require human approval for irreversible actions until scores earn autonomy. Put kill switches on spend and error rate. Sandboxes and gates are not optional for production tool use.
How much does an agentic pilot cost at Spurlock Studios?
The pilot is $1,500 for five business days: one narrow job on your real data, a working agent you keep, and a build quote based on what we saw. Details and packaging live on [/agentic](/agentic).
What is the difference between an agent and an automation?
Automation follows a known path with rare judgement. An agent chooses among tools and paths under uncertainty and must be evaluated. If you can draw the flowchart completely, you probably want automation. If the path varies but criteria are clear, you may want an agent.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot