Agentic Systems: An Operating Manual for Multi-Agent Work That Ships
An agentic system is evaluators, policy gates, sandboxes, and kill switches — not a chat window — so production tool work survives contact with real data.
William Spurlock Founder — Spurlock Studios Updated 32 MIN
An agentic system is not a chatbot with plugins. It is a production machine that plans, calls tools, checks its own work against criteria you defined, and stops when it should stop. If you cannot name the evaluator, the sandbox boundary, the state machine, and the kill switch, you do not have an agentic system. You have a demo.
This manual is how Spurlock Studios builds agentic work that founders and technical buyers can put on real data. It is the parent piece for the agentic lane. The spokes go deep on evaluators, policy gates, failed-tool loops, sandboxes, and the metrics that lie. Read this for the map. Use the spokes when one layer is the risk.
The short answer
- An agentic system chooses steps under uncertainty, writes to systems you care about, and is judged by something other than the worker that produced the artifact.
- Build the evaluator before the agent. Then put a policy gate in front of every side effect. Then cage the loop so failed tools cannot retry forever.
- Pass rate alone will green-light a grind. Gate deploys on revision rate, trajectory, coverage, and cost per success — see why pass rate lies.
- Default to automation when the path is known. Build an agent only when the path varies and you can still write pass/fail criteria. The brake pedal is when not to build an agent.
- Prove one sentence-sized job in five days on your data. The on-ramp is /agentic.
What is an agentic system?
An agentic system is software that can choose steps toward a goal, use tools to change the world outside the model, and revise its path when evidence says the last step failed — under constraints you own.
Three properties separate it from a scripted automation:
- Choice under uncertainty. The system picks the next action from a set of allowed tools and states, not from a fixed graph of “always do A then B.”
- External effects. It can read and write systems you care about: tickets, CRMs, inboxes, code, calendars, knowledge stores.
- Judgement that is not the worker. Something other than the same context that produced the artifact decides whether the artifact is acceptable.
n8n fits here as the rail for deterministic glue — webhooks, queues, retries, human approvals — while models and tool runners sit inside bounded steps. The rail is boring on purpose. The agent lives in the steps where choice is required; the rail owns delivery, idempotency, and escalation.
If your “agent” is a single prompt that calls three APIs and always returns green, call it an automation. Language matters because budgets, risk reviews, and success metrics change when you admit you are shipping non-deterministic software.
Anthropic’s own engineering note on building effective agents lands in the same place: start simple, measure, and add multi-step agentic loops only when a cheaper pattern fails. That is not a slogan. It is the cheapest way to avoid a fleet you cannot operate.
| You have… | Call it… | Primary control |
|---|---|---|
| Fixed path, rare judgement | Automation | Workflow tests + retries |
| Varying path, crisp criteria | Agentic system | Evaluator + gates + state machine |
| Varying path, mushy taste | Workshop, not a build | Criteria first, then decide |
| High stakes, weak recovery | Human + checklist | Do not auto-write |
What operating stack has to exist?
Every production agentic system at Spurlock Studios is built from the same stack. Skip a layer and you will pay for it in production, usually on a Tuesday.
| Layer | Job | Failure mode if missing |
|---|---|---|
| Job contract | One sentence goal + acceptance criteria | Infinite scope, unmeasurable demos |
| Evaluator | Independent pass/fail with evidence | Self-grading theater |
| Policy gate | Allow / deny / pending before side effects | Prompt-only “safety” |
| Tool sandbox | Allowed actions, secrets, blast radius | Agents that email customers or delete rows |
| State machine | Explicit states and transitions | Loops that never halt, duplicate writes |
| Memory policy | What persists, what dies with the run | Contaminated context |
| Retrieval contract | What may be cited as fact | RAG that invents policy |
| Handoff protocol | What moves between agents | Lost context, double work |
| Cost + kill switches | Budgets, caps, abort | Surprise invoices |
| Observability | Traces, scores, operator dashboard | You cannot debug or trust it |
You can implement these in different stacks. The stack is not the product. The contracts are.
NIST’s AI Risk Management Framework frames the same work as Govern, Map, Measure, and Manage. The AI RMF 1.0 is voluntary and use-case agnostic; it will not write your job contract. It will tell a buyer why “we shipped a chat UI” is not a risk program. The RMF Core is the part that maps onto this manual: you govern who can widen tools, you map what the job can break, you measure with an evaluator, and you manage abort and escalate as first-class states.
- Job contract written with the buyer, including hard nos
- Evaluator returns structured verdicts with evidence
- Policy gate runs in code, fail-closed, before every write
- Tool allowlist and scoped credentials
- Terminal states:
done,escalate,abort - Per-run budget and a kill switch you have actually tripped in staging
- Trace fields an operator will read without opening a JSON dump
Why evaluators before agents?
Build the evaluator before the agent. That sentence is the whole strategy.
The evaluator is a separate component whose only job is to judge an artifact against criteria. It must not see the worker’s chain of thought. It must return a structured verdict: pass or fail, which criterion failed, evidence, and a next action when fail is recoverable.
Mechanical checks first. Schema validity, required fields, unit tests, allowlisted URLs, “ticket status is one of these enums,” “invoice total matches line items.” Models judge only what genuinely needs judgement: tone for a customer email, whether a summary omitted a material risk, whether a research brief answered the asked question.
Without an evaluator you are optimizing prompts in the dark. With one, every model swap, tool change, and prompt edit becomes a measured experiment.
OpenAI’s own evaluation best practices say the same thing in vendor language: write scoped tests early, make them task-specific, log everything, and automate scoring when you can. Their agent evals guide is explicit that a trace — model calls, tool calls, guardrails, handoffs — is what you grade when the unit of work is a loop, not a single completion. Working with evals is the API-shaped version: a dataset plus graders, not a vibe check after a demo.
Vendor eval products move. As of August 2026, OpenAI has published a deprecation window for its standalone Evals platform. That is a reason to own the golden set and the grader contract in your repo, not a reason to skip evaluation. The suite has to survive the logo on the dashboard.
Deep dive: Build the Evaluator Before the Agent.
| Check type | Owner | Example |
|---|---|---|
| Schema | Code | JSON matches the contract |
| Business rule | Code | Totals equal line items |
| Citation rule | Code + light model | URL present or explicit no-match |
| Judgement | Model evaluator | Did the brief answer the asked question |
| Safety | Code + allowlists | Banned promises, PII patterns |
Pass rate is a vanity metric
A 94% pass rate can hide a system that revised four times, called the wrong tool, skipped half the real job shapes, and still needed a human rewrite. “Pass” is usually a thin binary on the final artifact. It ignores trajectory, coverage, and unit cost.
Gate deploys on a panel, not a single percentage:
| Metric | What it catches | Deploy veto if… |
|---|---|---|
| Golden-set pass rate | Obvious quality regressions | Drops past the agreed band |
| Revision rate | Grind-to-green | Mean revisions climb while pass holds |
| Trajectory score | Wrong tools, extra steps | Tool-choice errors rise |
| Eval coverage | Untested job shapes | New production clusters have no cases |
| Cost per success | Expensive “wins” | Dollars per passing run leave the band |
| Online / offline gap | Silent production drift | Sampled live scores diverge from the suite |
| Escalate rate | Hidden human load | Humans become the real runtime |
OpenAI’s agent-eval docs treat tool choice and handoff timing as first-class grades, not footnotes. If you only score the final blob, you will ship agents that wander and still “pass.” The spoke that owns the panel is why pass rate lies.
Do not celebrate latency alone. Fast wrong is still wrong. Do not celebrate a pass-rate jump after you loosened criteria. Version the evaluator. A criteria change is a release, not a rounding error.
How do pre-execution policy gates work?
A kill switch that lives in the system prompt is a suggestion. A pre-execution policy gate is a function in your runtime that sees the concrete tool name and arguments and returns allow, deny, or pending-approval before the tool runs. If the policy service is down, you fail closed.
That is the difference between “we told the model not to refund” and “the refund tool never fired.”
| Decision | When | What the trace must show |
|---|---|---|
allow | Payload matches policy | Tool, args hash, rule id |
deny | Out of allowlist, over cap, bad tenant | Reason code, no side effect |
pending-approval | Irreversible or novel class | Queue id, proposed payload |
| Fail closed | Policy timeout, unknown tool, bad parse | Abort, no retry-as-allow |
Gates sit after the model proposes and before the sandbox executes. Sandboxes limit damage if a call gets through. Gates decide whether the call happens at all. You need both.
- Unknown tools cannot execute
- Write tools require a tenant id that matches the run
- Row / recipient / dollar caps enforced in code
- Policy outage cannot be bypassed by the worker
- Deny and pending leave a reason code an operator can filter
Failed-tool loops
Most “the agent got stuck” incidents are not mysterious. The tool returned an error, the model treated the error as more story, and the loop called the same tool with the same args until the budget died. The spoke that owns this failure is why agents loop on failed tools.
The operating fix is typed errors plus illegal transitions.
- Tool adapters return a typed result:
ok,retryable,fatal,auth,not_found,rate_limited. - The state machine, not the model, decides the next state.
- The same
(tool, args_hash, error_class)pair cannot fire twice in one run. retryablegets a bounded backoff on the rail (n8n, queue, whatever you already trust).fatalandauthgo toescalateorabortwith the trace attached.- The evaluator sees the tool outcome. A worker that “summarizes past” a failed write does not get to mark
done.
| Error class | Legal next state | Illegal next state |
|---|---|---|
ok | evaluate | Silent second write |
retryable | act once more, then escalate | Infinite retry |
rate_limited | Wait on the rail | Immediate re-call |
auth | escalate | Guess a new token |
not_found | Revise query once, then escalate | Invent the record |
fatal | abort | “Try a different tool that deletes” |
If you cannot draw that table for your tools, you do not have a loop. You have a hope.
How do you sandbox tool use?
An agent without a sandbox is a liability with an API key.
Sandbox means:
- Allowlist of tools, not “whatever the model invents.”
- Scoped credentials — read-only where possible, write scopes only for the tools that must write.
- Blast-radius limits — rate caps, row caps, recipient caps, environment isolation (staging vs production).
- Dry-run modes for first contact with a new tool.
- Human gates on irreversible actions until the evaluator and error rates earn autonomy.
Anthropic’s computer-use documentation is unusually blunt for a vendor: dedicated VM or container, no sensitive logins in the environment, domain allowlists, and a human confirm on consequential actions. Their computer-use research note names the reason: screenshot and page content can carry prompt injection that overrides your instructions. If that is true for a desktop sandbox, it is true for a CRM tool that reads ticket text.
MCP servers and custom tool runners are fine. The Model Context Protocol is a way to expose tools and context over a standard. The 2026-07-28 spec even pushed toward stateless, self-contained requests so you can load-balance without a shared session store. That is transport. It is not a permission system. Unrestricted shell, unrestricted email send, and “admin” CRM tokens are still not fine for a pilot.
Deep dive: Sandboxed Tool Use.
| Tool class | Pilot default | Autonomy earned when |
|---|---|---|
| Read ticket / CRM | Allow, scoped | Always, with redaction |
| Write internal draft | Allow to internal field | Online scores hold |
| Send customer email | Human gate | Rewrite rate and silent-fail samples stay low |
| Refund / billing change | Deny or pending | Almost never in week one |
| Shell / code exec | Isolated runner or deny | Job actually needs it |
Where do state machines belong?
Agent loops need freedom inside a cage. The cage is a state machine.
Typical states for a business agent: intake → plan → act → evaluate → revise → done | escalate | abort. Transitions are explicit. Side effects only happen in act, and only after the policy gate. Evaluation never mutates production systems. Revision has a ceiling (usually three). Escalation packages the full trace for a human.
n8n is a natural home for the cage: each state can be a node or sub-workflow, with durable execution, retries, and a dead path for escalate. The model proposes; the machine decides whether the transition is legal.
If you run n8n in production, treat durability as part of the agent, not as hosting trivia. n8n’s durable scheduler exists so time-based work survives restarts and does not double-fire across instances. Queue mode is how production executions leave the editor process and land on workers. An agent that “works in the canvas” and vanishes on deploy is not an agentic system. It is a local demo with extra steps.
LangGraph’s persistence model is the same idea with different nouns. Checkpointers snapshot thread state so you can pause for a human, resume after a crash, and avoid re-running work that already succeeded. Persistence splits that short-term thread memory from long-term stores (preferences, facts). If you dump both into one transcript, you will re-inject failed reasoning as if it were policy.
| State | May write? | May call model? | Exit condition |
|---|---|---|---|
intake | No | Classify only | In-scope contract attached |
plan | No | Yes | Tool plan within allowlist |
act | Yes, after gate | Yes | Typed tool result |
evaluate | No | Judge only | Verdict + evidence |
revise | No | Yes, capped | New artifact or ceiling |
done | Confirm only | No | Evaluator passed |
escalate | No | No | Human package stored |
abort | No | No | Kill reason stored |
How should memory, retrieval, and handoffs be contracted?
Agent memory is not “stuff the whole transcript into the next call.” Memory is a policy. The policy itself is agent memory patterns.
Separate at least four stores:
- Ephemeral run context — dies when the run ends.
- Working scratch — intermediate artifacts for this job only.
- Durable facts — customer prefs, account IDs, approved SOPs — with ownership and TTL.
- Run history / traces — for ops and learning, not for raw re-injection into every prompt.
Persist preferences and identifiers. Forget raw intermediate reasoning. Never let a failed run’s bad conclusions become long-term “memory” without a promotion rule.
Retrieval-augmented generation fails in businesses for a boring reason: teams treat “retrieved” as “true.” Retrieval is a search result. Truth is a contract.
A retrieval contract answers:
- Which corpora are authoritative for which question types?
- What freshness rules apply?
- Must citations be present for any factual claim?
- What happens when retrieval returns nothing — refuse, ask, or fall back to a human?
- How do you detect contradiction across chunks?
If the agent can invent policy when the index is empty, you do not have RAG. You have a confident liar with a vector database.
Multiple agents are useful when jobs naturally split: research vs draft vs compliance check; intake vs enrichment vs write-back. They are harmful when you multiply agents to look sophisticated.
A handoff is a typed package:
- Goal and constraints
- Artifacts produced so far
- Open questions
- Tools already tried and outcomes
- Budget remaining
- Evaluator criteria still unmet
Do not pass “the vibe.” Pass the package. The receiving agent should not need the sending agent’s private scratch.
| Store | Survives the run? | May enter the next prompt? |
|---|---|---|
| Ephemeral context | No | This run only |
| Working scratch | No | This run only, redacted |
| Durable facts | Yes, with TTL | Yes, if schema-valid |
| Traces | Yes | No — ops only |
| Retrieved chunks | Per query | Only with citation or no-match |
How do you evaluate agents without fooling yourself?
Evaluation is a product discipline, not a vibe check after a demo.
Unit-level
- Tool adapters: given fixture inputs, do they return typed outputs or typed errors?
- Retrievers: precision/recall on a labeled query set for your corpus.
- Schemas: every agent-facing JSON shape validates.
- Policy gate: unknown tool, over-cap payload, and missing tenant all deny.
Task-level
Build a golden set of 30–100 real jobs (anonymized if needed). For each: input, required artifacts, pass criteria, known traps. Run the suite on every change that could affect behavior. Track pass rate, average revisions, cost per pass, escalate rate, and coverage of the job shapes you actually see.
Online
Sample production runs. Score with the same evaluator. Alert when online scores drift from offline. Drift is how quiet failures start.
What not to measure alone
Latency and token count without quality. “User thumbs up” without criteria. Self-reported confidence from the worker. A pass rate with no revision or coverage number next to it.
OpenAI’s eval guidance is useful here even if you never touch their dashboard: collect cases from production logs, keep humans in the loop to calibrate automated graders, and treat evaluation as continuous. That last point is the one teams skip. A golden set that never gains a case after an incident is a museum.
| Layer | Question it answers | Cadence |
|---|---|---|
| Unit | Did this adapter lie? | Every commit |
| Task / golden set | Did this change regress the job? | Every behavior change |
| Trajectory | Did it take a sane path? | Every behavior change |
| Online sample | Is production drifting? | Daily or weekly |
| Silent-fail sample | Did a “pass” later get rewritten? | Weekly, painful, worth it |
When should you not build an agent?
Default to automation when the path is known, the inputs are structured, and judgement is rare. Default to a human when stakes are high and criteria are contested. Build an agent when the path varies, tools are many, and you can still write acceptance criteria crisp enough to evaluate.
Skip agents (for now) when:
- The path is fully known. Same steps, same systems, rare exceptions.
- You cannot write pass/fail criteria. “Good” is still an argument.
- Stakes are high and recovery is hard, and you do not yet have gates and sandboxes.
- Credentials and data access are political. You will spend the month on access, not learning.
- Volume is tiny. Ten items a month may want a human and a template.
- Nobody owns the SOP. The agent becomes a scapegoat.
- You want a demo more than a metric.
Deep dive: When Not to Build an Agent.
| Signal | Build this instead |
|---|---|
| Known path, structured I/O | n8n / automation |
| Criteria are mush | Process workshop |
| One irreversible write | Human + checklist |
| Need a public chatbot, no criteria | Wrong lane |
| Path varies, criteria exist | Agentic system |
What multi-agent shape ships for a business?
Here is a shape that ships for small and mid-size teams without becoming a research project.
Roles
- Router / intake — classifies the job, attaches the job contract, rejects out-of-scope work.
- Worker — plans and acts inside the sandbox.
- Evaluator — independent judgement; no tool writes.
- Librarian (optional) — retrieval only; returns citations or “no hit.”
- Operator surface — humans approve, abort, or re-scope.
Control flow
- Event or human request hits intake (often via n8n webhook).
- Job contract loaded; budget and tool allowlist attached.
- Worker enters
plan→actloop under the state machine. - Policy gate runs on every proposed write.
- After each material artifact, evaluator runs.
- Fail → revise until ceiling → escalate.
- Pass → write-back through allowlisted tools →
done. - Trace + cost + scores stored for ops.
What “done” means
Done is not “the model said done.” Done is: evaluator passed, side effects confirmed idempotently, and the run landed in a terminal state with a receipt the operator can audit.
Anthropic’s effective-agents writeup keeps repeating a useful constraint: successful teams were not the ones with the most elaborate graphs. They were the ones who could see the plan, keep the tool interface tight, and add complexity only when a simpler pattern failed. A business fleet that starts as intake + worker + evaluator is not “behind.” It is honest.
| Temptation | Cost | Do this first |
|---|---|---|
| Five specialists on day one | Lost handoffs, no owner | One worker, one evaluator |
| Shared chat as memory | Contaminated context | Typed handoff package |
| Evaluator with write tools | Self-dealing | Read-only judge |
| “Latest” model as default | Silent quality breaks | Pin, then re-run the suite |
What does a five-day pilot actually prove?
Most teams do not need a twelve-week “AI transformation.” They need one narrow job proven on their data.
The Spurlock Studios agentic pilot is $1,500 · 5 days. One job, scoped tight enough to finish in a week. A working agent on your real data — not a slide deck. You keep it either way. The $1,500 credits toward a full build.
What you leave with:
- A runnable agent for one sentence-sized job
- An evaluator with explicit criteria
- Sandboxed tools and a thin policy gate for that job
- A state machine with a revision ceiling and an escalate path
- A short build quote based on what we actually saw
What the week is not: a chatbot skin, a multi-agent org chart, or a promise that pass rate will hold after you add refunds and production sends.
Start on /agentic or go straight to /contact?intent=agentic-pilot.
| Day | Outcome you can point at |
|---|---|
| 1 | Job contract, hard nos, evaluator shape |
| 2 | Golden set v0 (real cases, including traps) |
| 3 | Sandboxed tools + dry-run writes |
| 4 | Loop + gate + escalate on your data |
| 5 | Scores, cost, quote, keep-the-agent handoff |
When the problem is architecture across a roadmap rather than a single agent, that is a different engagement shape — same principles, longer surface. The pilot still comes first because a roadmap without one green job is a slide.
What does a production walkthrough look like?
Job contract: “Given a new support ticket, classify severity, draft an internal summary with citations from the help center, and propose a reply — never send.”
Evaluator criteria (examples):
- Severity is one of
P1|P2|P3|P4 - Summary includes at least one citation URL from retrieval or explicitly says “no doc match”
- Proposed reply contains no promise of refund or SLA change unless those strings appear in retrieved policy
- Schema validates
Tools in sandbox: ticket read API, help-center retriever, draft write to internal field. Not in sandbox: send reply, issue refund, change billing.
Policy gate: deny send_reply and issue_refund by name. Cap retrieval calls. Require tenant_id on every tool payload.
State machine: intake → retrieve → draft → evaluate → revise (max 3) → escalate or done.
Memory: customer ID and prior ticket IDs may persist; raw model scratch does not.
Cost: hard cap on retrieval calls and revisions; abort to human queue if exceeded.
That system is agentic. A Zap that posts “new ticket” into Slack is not. Both can be valuable. Only one needs this manual.
| Step | Who decides | What gets written |
|---|---|---|
| Intake | Router + contract | Nothing external |
| Retrieve | Librarian / worker | Nothing external |
| Draft | Worker | Internal field only |
| Evaluate | Evaluator | Verdict record |
| Revise | Worker, if ceiling remains | Internal field overwrite |
| Escalate | State machine | Human package |
| Send | Human, later | Customer channel |
What fails after the demo?
Agent theater. Fancy UI, no evaluator, no sandbox, no budget. Demo day works. Week three does not.
Prompt as policy. Rules living only in natural language. Policies belong in code checks and allowlists; language fills gaps. OWASP’s LLM Top 10 still leads with prompt injection for a reason: the model will follow instructions found in content. The community writeup on prompt injection is the plain-language version. The OWASP GenAI project is the living index. None of those pages say “add a nicer system prompt” as the whole fix.
Unbounded loops. No revision ceiling. No typed tool errors. Cost and chaos grow together.
RAG without refuse. Empty retrieval still produces “facts.”
Too many agents too early. Three agents before one job is green. Split only after the single-worker path is measured.
No human path. Escalation is a first-class state, not an apology.
Pass-rate theater. One green percentage, no revision or coverage number, no online sample.
Shared enrichment keys. One credential that can see every tenant. That is a breach design, not a shortcut.
| Failure | What it costs | What you do instead |
|---|---|---|
| Self-grading worker | Silent wrongness | Independent evaluator |
| Prompt-only deny | Injection bypass | Policy gate, fail closed |
| Retry-on-text-error | Token burn, duplicate writes | Typed errors + illegal transitions |
| Empty-index answers | Invented policy | Refuse or escalate |
| Latest-model default | Unexplained score drops | Pin, re-run golden set |
What should ops actually read?
If the only “observability” is a provider dashboard, you will not catch silent wrongness. Traces must show: state, tool calls, inputs/outputs (redacted), gate decisions, evaluator verdicts, cost, latency, and escalation reason. Scores from your evaluator suite should land on a dashboard a human checks weekly — not a graveyard of JSON in object storage.
OpenTelemetry now keeps GenAI conventions in a dedicated repo: semantic-conventions-genai. The older opentelemetry.io GenAI pages redirect there. You do not have to export every prompt. You do have to agree on span names and attributes so “what tool fired, for which tenant, at what cost” is not a custom folklore per engineer. The attribute registry is the boring list that makes that possible.
Token spend is a product feature. Treat it like one.
Per-run budgets, per-day budgets, max tool calls, max revisions, model tiers by state (plan on a cheaper model, evaluate on a stricter one when needed), and hard kill switches when spend or error rate crosses a line. Log cost on every transition. Ops should see dollars next to failure rates.
| Cap | Typical pilot default | What it stops |
|---|---|---|
| Max tool calls / run | Low double digits | Wandering tool spam |
| Max revisions | 3 | Grind-to-green |
| Per-run spend | Set with the buyer | One runaway loop |
| Per-day spend | Set with the buyer | Quiet overnight burn |
| Error-rate kill | Trip after repeated fatal / auth | Retry storms |
A kill switch you have never tripped in staging is a rumor. Force a budget abort on purpose before the first soft launch. The trace should show abort and a reason code, not a hung act.
Human gates fail when every run waits on a busy founder. Design queues:
- Batch review for soft writes (internal notes) twice a day
- Immediate review only for irreversible classes
- Auto-promote when online pass rate holds for a defined window on that job type
- Spot checks forever — autonomy is not absence of audit
The operator surface should show the same trace fields ops already use: criteria failures, cost, and the proposed write payload. Asking a human to re-read the whole chat is how gates get muted.
| Role | Owns | Weekly artifact |
|---|---|---|
| Job owner | Criteria and risk | Criteria changelog |
| Systems owner | Credentials, schemas, caps | Access review |
| Agent engineer | Prompts, tools, state machine | Suite diff |
| Ops reviewer | Scores and incidents | Drift + silent-fail sample |
One person can wear multiple hats at a small company. Zero people wearing the ops hat is how silent failure becomes culture.
Tenancy and injection
If more than one customer or department shares infrastructure, tenancy is an agent feature. Every run carries tenant_id. Tool credentials are bound to that tenant. Retrieval ACLs filter before ranking. Memory keys are prefixed. Logs are partitioned. A “shared enrichment key” that can see every CRM is a data-breach design.
Prompt injection is a tenancy problem too. Ticket text, email bodies, and retrieved pages are untrusted. OWASP treats that as LLM01 for a reason. Content from Tenant A must never expand tools or memory for Tenant B. Sandboxes and allowlists are the first wall. Evaluator checks for cross-tenant identifiers in artifacts are a useful second wall. The policy gate is the wall that actually stops the write.
| Boundary | Enforce in | Fail mode if skipped |
|---|---|---|
| Credential | Per-tenant secret, not a shared admin token | Cross-tenant read |
| Retrieval | ACL filter before rank | Policy leak across accounts |
| Memory key | {tenant_id}:{entity} | Contaminated prefs |
| Tool args | Gate rejects foreign ids | Write to the wrong account |
| Logs | Partition + redaction | One export holds everyone |
- No tool credential can see more than one tenant
- Retrieval returns zero hits rather than another tenant’s doc
- Memory promotion requires a tenant-scoped schema
- Gate denies payloads whose ids do not match the run
- Operator exports cannot dump another tenant’s traces by default
This is not “enterprise later.” It is the minimum if two teams share a runtime. Agencies already know the version of this story from n8n: isolate credentials or inherit incidents.
Pin models against drift
“Latest” as a default is an availability choice that often breaks quality silently. Pin the model id on every call. When the provider ships a new version, run the golden set in a side-by-side before you promote. A pass-rate bump after a silent upgrade is not a win until you know which criteria moved.
Drift shows up in three places:
- Provider drift — the pinned id still exists, behavior changed, or the pin was ignored.
- Prompt drift — someone edited the worker or the evaluator and did not version it.
- World drift — the job changed (new ticket types, new policy docs) and the golden set did not.
Treat each as a release. OpenAI’s eval docs call this continuous evaluation: grow the set from production logs, keep humans calibrating the grader. You do not need their product to do that. You need a suite that fails CI when scores leave the band.
| Change | Required before promote |
|---|---|
| New model id | Full golden set + cost band |
| Prompt or tool schema | Golden set + trajectory slice |
| Evaluator criteria | Version bump + note that pass rate is not comparable |
| New production failure | New case within 48 hours |
| Widened tool allowlist | Deny-path tests for the new tool |
If you cannot say which model id wrote last week’s traces, you cannot debug last week’s traces.
Idempotent writes
Agents retry. Rails retry. Humans click twice. If act is not idempotent, a “successful” run can create two tickets, two drafts, or two charges. The state machine can forbid a second transition and still lose if the first write’s ack never came back.
Give every side effect an idempotency key derived from run_id + tool + args_hash (or a business key the downstream already understands). The tool adapter must be safe to call twice with the same key. The evaluator must not treat “record already exists with this key” as a hard fail if the payload matches.
| Write | Key | Safe retry looks like |
|---|---|---|
| Internal draft | run_id:draft | Overwrite same field |
| Ticket comment | run_id:comment | Same comment id returned |
| CRM note | run_id:note or external id | No second note |
| Email send | Do not auto-retry | Human or outbox with key |
| Charge / refund | Downstream idempotency key | One money movement |
Dry-run modes should exercise the key path, not only the happy path. A sandbox that cannot show you a duplicate-suppressed write is a sandbox that will surprise you on the first timeout.
What order do you build in?
- Write the job contract and acceptance criteria with the buyer.
- Build the evaluator and a tiny golden set.
- Implement tools behind a sandbox with dry-run.
- Put the policy gate in front of every write. Fail closed.
- Wire the state machine (n8n or equivalent) with budgets, typed errors, and escalate.
- Add retrieval and memory only if the job needs them — with contracts.
- Run the golden set until pass rate and revision rate and cost are acceptable.
- Soft-launch with human gates on writes.
- Widen autonomy only when online scores hold.
- Split agents only after the single-worker path is measured.
Skipping to step 7 because a vendor demo looked good is how you buy regret.
Treat each layer as a module with an owner and a test. The evaluator exports judge(artifact, criteria) -> Verdict. The sandbox exports callTool(name, args, ctx) -> Result. The gate exports decide(tool, args, ctx) -> allow|deny|pending. The state machine exports transition(state, event, ctx) -> State. Memory and RAG export read/write functions with schemas. Observability wraps all of the above.
Integration tests should freeze a run through intake to terminal with fixture tools. Contract tests should freeze golden-set scores on CI. Load tests should freeze budget trips. You do not need a research lab. You need the same hygiene you already use for payments and auth.
Document the hard nos in the same repo as the code. Hard nos that live only in Slack will be rediscovered after an incident.
When a vendor sells you “agents,” ask:
- Show the evaluator on our sample cases, not yours.
- Show the tool allowlist and how new tools are added.
- Show the policy gate deny a real payload.
- Show the state machine or equivalent control flow.
- Show per-run budgets and a kill switch demo.
- Show a trace with redaction.
- Show what happens on empty retrieval.
- Show what happens when the same tool fails twice.
- Show who owns prompts after go-live.
- Show exit: can we export and run without you?
If answers are slides without receipts, you are buying theater.
| Word | Meaning in this manual |
|---|---|
| Job contract | Goal, audience, criteria, hard nos |
| Evaluator | Independent verdict with evidence |
| Policy gate | Allow / deny / pending before side effects |
| Sandbox | Allowlisted tools + caps + least privilege |
| State machine | Legal transitions and terminals |
| Handoff package | Typed relay between agents or humans |
| Kill switch | Automatic stop on spend, error, or policy |
| Golden set | Labeled jobs for regression |
Use the words precisely. Language drift recreates agent theater under new names.
Founders and technical buyers who need work done — triage, research briefs, enrichment, internal ops agents, content drafts with hard constraints — and who will not accept “trust the model.” If you want a public chatbot with no criteria, this is the wrong lane.
Spurlock Studios ships agentic systems with explicit state machines, sandboxed tool runners, and reflection loops that self-correct. Builds typically land in 2 to 10 weeks after a pilot proves the job. That range is packaging, not a promise that your CRM is a two-week problem.
I have spent 20,000+ hours architecting agentic systems and shipped 500+ automations. The hours do not make the stack optional. They are why I refuse to start at the chatbot.
Read the spokes in the order your risk demands. Most teams should start with evaluators, then sandboxes and policy gates, then prove the job. Come back to this manual when you need the full map.
Cluster map
This manual is the hub. Use the spoke that matches the failure, not the one with the trendiest demo.
Evaluate before you scale
- Evaluators before agents
- The evaluator is the product
- Why pass rate lies
- When not to build an agent
- Why agent demos fail in production
Control the tools
- Tool-use sandboxes
- Pre-execution policy gates
- Prompt-injection defense for agents
- Idempotent agent tool writes
- State machines for agent loops
Run it in production
- Agent memory patterns
- Observability for agents
- Cost controls for agent fleets
- Single vs multi-agent
- Why agents loop on failed tools
FAQ
What is an agentic system in plain terms?
An agentic system is software that can choose tools and steps toward a goal, change external systems, and revise when checks fail — under budgets and rules you define. It is not a chat UI. The difference from automation is meaningful choice under uncertainty plus independent evaluation.
How do you evaluate AI agents without fooling yourself?
Separate the evaluator from the worker. Use mechanical checks first, then model judgement only where needed. Maintain a golden set of real jobs and run it on every meaningful change. Track pass rate, revisions, cost per pass, coverage, and escalate rate. Never trust the worker’s self-score as the primary metric.
When should we not build an agent?
Skip the agent when the path is already known, when you cannot write pass/fail criteria, or when the honest fix is a checklist and a webhook. High-stakes writes without gates and sandboxes are another no. If you are unsure, start with automation or a scoping conversation — the longer argument is when not to build an agent.
How do we stop agents from doing dangerous things?
Allowlist tools, scope credentials, cap blast radius, and put a policy gate in front of every side effect. Require human approval for irreversible actions until scores earn autonomy. Put kill switches on spend and error rate. Sandboxes and gates are not optional for production tool use.
How much does an agentic pilot cost at Spurlock Studios?
The pilot is $1,500 for five business days: one narrow job on your real data, a working agent you keep, and a build quote based on what we saw. Details and packaging live on /agentic.
What is the difference between an agent and an automation?
Automation follows a known path with rare judgement. An agent chooses among tools and paths under uncertainty and must be evaluated. If you can draw the flowchart completely, you probably want automation. If the path varies but criteria are clear, you may want an agent.
CTA
Ready to prove one job in five days? /agentic · /contact?intent=agentic-pilot
What questions does this article answer?
- What is an agentic system in plain terms?
- An agentic system is software that can choose tools and steps toward a goal, change external systems, and revise when checks fail — under budgets and rules you define. It is not a chat UI. The difference from automation is meaningful choice under uncertainty plus independent evaluation.
- How do you evaluate AI agents without fooling yourself?
- Separate the evaluator from the worker. Use mechanical checks first, then model judgement only where needed. Maintain a golden set of real jobs and run it on every meaningful change. Track pass rate, revisions, cost per pass, coverage, and escalate rate. Never trust the worker’s self-score as the primary metric.
- When should we not build an agent?
- Skip the agent when the path is already known, when you cannot write pass/fail criteria, or when the honest fix is a checklist and a webhook. High-stakes writes without gates and sandboxes are another no. If you are unsure, start with automation or a scoping conversation — the longer argument is [when not to build an agent](/blog/when-not-to-build-an-agent).
- How do we stop agents from doing dangerous things?
- Allowlist tools, scope credentials, cap blast radius, and put a policy gate in front of every side effect. Require human approval for irreversible actions until scores earn autonomy. Put kill switches on spend and error rate. Sandboxes and gates are not optional for production tool use.
- How much does an agentic pilot cost at Spurlock Studios?
- The pilot is $1,500 for five business days: one narrow job on your real data, a working agent you keep, and a build quote based on what we saw. Details and packaging live on [/agentic](/agentic).
- What is the difference between an agent and an automation?
- Automation follows a known path with rare judgement. An agent chooses among tools and paths under uncertainty and must be evaluated. If you can draw the flowchart completely, you probably want automation. If the path varies but criteria are clear, you may want an agent.
- anthropic.com
- nist.gov
- nvlpubs.nist.gov
- airc.nist.gov
- developers.openai.com
- developers.openai.com
- developers.openai.com
- platform.claude.com
- anthropic.com
- modelcontextprotocol.io
- blog.modelcontextprotocol.io
- docs.n8n.io
- docs.n8n.io
- docs.langchain.com
- docs.langchain.com
- owasp.org
- owasp.org
- genai.owasp.org
- github.com
- opentelemetry.io
- opentelemetry.io
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
AI Agents Why Agents Loop on Failed Tools: No-Progress Detection Beats Longer Prompts
Agents loop on failed tools because the harness never detects no-progress. Fingerprint calls, honor retryable:false, cap turns, terminate with a reason code.
AI Agents Why Agent Demos Die in Production: Control Gaps, Not Model IQ
Demo success proves a staged happy path. Production fails when the control loop — schemas, auth, evaluators, and kill switches — was never in the harness.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.