Spurlock Studios
Contact
Tag

agents

This is the agents tag archive on Will's Journal: every published post that shares this tag, listed in one place.

It is a collection page, not a topic essay — scan the cluster here instead of filtering the full journal index.

Which posts are tagged agents?

Two clipped paper packets. Thesis: SHOULD AGENT ESCALATE INSTEAD RETRYING. AI Agents

When should the agent escalate instead of retrying

Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.

26 MIN
A small text-file card with no glyphs. Thesis: AGENT 10 MORE EXPENSIVE THAN. AI Agents

Why is my agent 10× more expensive than the chatbot demo

Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.

34 MIN
A violet ring. Thesis: WHO ACCOUNTABLE AGENT ACTS REFUNDS. AI Agents

Who is accountable when an agent acts (refunds, emails, writes)

A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.

31 MIN
A small stack of coins. Thesis: PASS RATE LIES REVISION RATE. AI Agents

Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage

Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.

25 MIN
A scuffed work smartphone with a blank glowing circular button. Thesis: AGENTS LOOP FAILED TOOLS PROGRESS. AI Agents

Why Agents Loop on Failed Tools: No-Progress Detection Beats Longer Prompts

Agents loop on failed tools because the harness never detects no-progress. Fingerprint calls, honor retryable:false, cap turns, terminate with a reason code.

22 MIN
A violet ring. Thesis: AGENTIC SYSTEMS OPERATING MANUAL MULTI. AI Agents

Agentic Systems: An Operating Manual for Multi-Agent Work That Ships

An agentic system is evaluators, policy gates, sandboxes, and kill switches — not a chat window — so production tool work survives contact with real data.

32 MIN
Nested brass frames. Thesis: AGENT DEMOS DIE PRODUCTION CONTROL. AI Agents

Why Agent Demos Die in Production: Control Gaps, Not Model IQ

Demo success proves a staged happy path. Production fails when the control loop — schemas, auth, evaluators, and kill switches — was never in the harness.

24 MIN
A cracked amber fuse. Thesis: BROKE TRIED EVALUATE AI AGENT. AI Agents

What broke when I tried to evaluate an AI agent in production

Production eval failed on four fronts: no gold labels, biased samples, judge drift, eval tools that wrote. Isolate the harness before you trust the live score.

29 MIN
A violet ring. Thesis: BUILD EVALUATOR BEFORE AGENT. AI Agents

Build the Evaluator Before the Agent

If the judge shares the worker's context, you are grading your own homework. Build criteria, evidence, and ceilings first — then grant real tool autonomy.

22 MIN
A violet ring. Thesis: EVALUATOR PRODUCT. AI Agents

The Evaluator Is the Product

Agent accuracy did not come from a better prompt or a bigger model. It came from separating the thing that does the work from the thing that judges it.

12 MIN
A small stack of coins. Thesis: KNOW IF AI AGENT ACTUALLY. AI Agents

How do I know if my AI agent is actually working

You know an AI agent is working when rewrite rate, policy denials, cost, and time-to-done hold. Thumbs and CSAT hide unpaid human editors on real writes.

25 MIN
Nested brass frames. Thesis: WRITE GOOD TOOL SCHEMAS AI. AI Agents

How do I write good tool schemas for AI agents

Write tool schemas as agent UX: honest required fields, enums for closed sets, descriptions that constrain, and one non-overlapping tool per side effect.

22 MIN
A violet ring. Thesis: STOP AGENT DOING SOMETHING DESTRUCTIVE. AI Agents

How do I stop an agent from doing something destructive

Stop destructive agent actions with a blast-radius table, an allowlisted tool set, and dual control on money and delete, enforced in code before the tool runs.

28 MIN
A small stack of coins. Thesis: TEST AI AGENTS BEFORE THEY. AI Agents

How do I test AI agents before they ship

Test AI agents before they ship on a golden set, staged tools, and shadow mode. Pass rate is not the gate — score cost, escalate, and known-bad fails too.

26 MIN
An expired brass key. Thesis: MANAGE MULTIPLE AI AGENTS PRODUCTION. AI Agents

How do I manage multiple AI agents in production

Isolate each production agent: credentials, write surface, named owner, and evals. Handoff with typed contracts. Do not share one god-agent across jobs.

27 MIN
A violet ring. Thesis: KEEP AGENT EMAILING CUSTOMERS INJECTED. AI Agents

How do I keep the agent from emailing customers when injected via a ticket

Treat ticket text as untrusted data. Keep send-email off the reader, gate the proposed send in code, and require a human before any customer outbound fires.

28 MIN
A cracked amber fuse. Thesis: SET POLICIES AGENTS CANNOT. AI Agents

How do I set policies for what agents can and cannot do

Write a named-owner catalog that returns allow, deny, or pending-approval before every tool, then fail closed if that catalog cannot load. Not legal advice.

25 MIN
A small stack of coins. Thesis: SINGLE AGENT FIRST SPLIT ONLY. AI Agents

Single Agent First: Split Only When Trust, Audience, or Timing Conflicts

Start with one agent and many tools. Split only when trust, audience, or timing conflict—and prove that split with pass rate, cost, and escalate rate.

26 MIN
Nested brass frames. Thesis: DID LOADING ALL MCP TOOLS. AI Agents

Why did loading all my MCP tools blow the context window

Every MCP tool schema you attach is prompt tokens. Loading the whole catalog fills the window before the task starts—send a job-sized subset or a router.

32 MIN
Two clipped paper packets. Thesis: MONITOR AI AGENTS PRODUCTION. AI Agents

How do I monitor AI agents in production

Monitor production agents on traces, tool errors, cost per turn, human rewrite rate, and policy denials — not chatbot thumbs. That is the ops scoreboard.

28 MIN
A cracked amber fuse. Thesis: BIG SHOULD GOLDEN SET BEFORE. AI Agents

How big should my golden set be before soft-launch

Before soft-launch, size the golden set by coverage — happy, auth fail, empty tool, policy deny, money, ambiguous — not a magic N or vanity pass rate.

30 MIN
A scuffed work smartphone with a blank glowing circular button. Thesis: AGENT KEEP CALLING SAME FAILED. AI Agents

Why does my agent keep calling the same failed tool

Your agent retries one failed tool because nothing fingerprints the call or caps same-tool executes. Escalate auth and policy on first hit — do not grind.

25 MIN
Two clipped paper packets. Thesis: STATE MACHINES AGENT LOOPS DETERMINISM. AI Agents

State Machines for Agent Loops: Determinism Where It Matters

Give the model freedom inside a named state. Legal transitions, revision ceilings, and escalate paths keep operations deterministic when tokens are not.

26 MIN
A scuffed work smartphone with a blank glowing circular button. Thesis: MCP NATIVE FUNCTION CALLING PORTABILITY. AI Agents

MCP vs Native Function Calling: Portability Tax vs Shortest Loop

Native function calling wins for one app's short tool loop; MCP earns the tax when tools must be shared and governed across hosts—not a LangChain swap.

18 MIN
A lime beam hitting a small brass nameplate. Thesis: RAG LIE RETRIEVAL CONTRACTS BUSINESS. AI Agents

RAG That Does Not Lie: Retrieval Contracts for Business Knowledge

Retrieval is search, not truth. Production RAG needs corpus rules, citations, refuse-on-empty, and contradiction handling — or the agent invents policy.

18 MIN
A small stack of coins. Thesis: LANGGRAPH CREWAI CUSTOM LOOP CHOOSE. AI Agents

LangGraph vs CrewAI vs a Custom Loop: Choose Control, Not Fashion

Choose LangGraph, CrewAI, or a custom loop by control and durability, then prove the pick on one golden set and a fixed cost band—not fashion rankings.

20 MIN
A brushed metal coupon. Thesis: AGENT MEMORY PATTERNS PERSIST FORGET. AI Agents

Agent Memory Patterns: What to Persist, What to Forget

Agent memory is three stores, not a bigger window: working dies with the run, episodic is searchable history, and the store holds approved facts only.

19 MIN
A folded lab sheet with no readable lines. Thesis: GOLDEN SETS PRODUCTION FAILURES TURN. AI Agents

Golden Sets from Production Failures: Turn Bad Runs into Regression Fuel

Turn a bad production agent run into a regression test: harvest the trace, stub the tools, write the expected terminal, and gate deploys on that fixture.

18 MIN
A scuffed work smartphone with a blank glowing circular button. Thesis: PROMPT INJECTION TOOL AGENTS STOP. AI Agents

Prompt Injection for Tool Agents: Stop Text from Becoming Actions

Stop prompt injection by isolating untrusted email and tickets from write tools, then gating every proposed call in code — never in the system prompt.

18 MIN
A small stack of coins. Thesis: COST CONTROLS AGENT FLEETS BUDGETS. AI Agents

Cost Controls for Agent Fleets: Budgets, Caps, and Kill Switches

Control agent-fleet cost with runner-enforced caps, policy-versioned prompt caches, and state-based model routing — not a prompt that says to be frugal.

25 MIN
A cracked amber fuse. Thesis: IDEMPOTENT AGENT TOOL WRITES RETRIES. AI Agents

Idempotent Agent Tool Writes: Retries Without Double Emails or Double Charges

Timeouts are unknown, not failed. Mint one runtime idempotency key per write intent and reuse it across retries so you cannot double-email or double-charge.

16 MIN
A small text-file card with no glyphs. Thesis: FRACTIONAL AI CTO MODEL NEED. AI Agents

The Fractional AI CTO Model: When You Need Architecture, Not Another Chatbot

A fractional AI CTO owns architecture, evaluation standards, and a deploy gate — not a chatbot retainer. Hire one when colliding initiatives lack rails.

24 MIN
A violet ring. Thesis: SCOPING AGENTIC PILOT PROVES VALUE. AI Agents

Scoping an Agentic Pilot That Proves Value in Five Days

Scope a 2–6 week agent pilot as one job, real data, an evaluator, and a cage. Five days proves the spike; extra weeks buy access, labels, and a second measure.

20 MIN
A violet ring. Thesis: PIN MODEL GATE UPGRADE CATCH. AI Agents

Pin the Model, Gate the Upgrade: Catch Agent Drift Before Customers Do

Yes—pin production agents to explicit model IDs. Floating aliases change behavior with no deploy. Upgrade only through a golden-set gate you actually run.

20 MIN
Two clipped paper packets. Thesis: OBSERVABILITY AGENTS TRACES SCORES DASHBOARD. AI Agents

Observability for Agents: Traces, Scores, and the Dashboard Ops Actually Reads

Provider dashboards miss silent wrongness. Agent observability is traces with redacted tool I/O, evaluator scores, and cost — the screen ops actually reads.

18 MIN
Nested brass frames. Thesis: TOOL SCHEMAS AGENTS FOLLOW DESCRIPTIONS. AI Agents

Tool Schemas Agents Follow: Descriptions, Enums, and Killing the Omnibus Tool

Agents invent arguments when schemas are vague. Write JSON Schema like agent UX—enums, required fields, property descriptions—and kill the do_anything tool.

16 MIN
Nested brass frames. Thesis: BUILD AGENT BUILD INSTEAD. AI Agents

When Not to Build an Agent (And What to Build Instead)

Skip the agent when the path is known or criteria are mush. A workflow plus one schema-checked LLM step is the default. Build the loop only after that fails.

28 MIN
A scuffed work smartphone with a blank glowing circular button. Thesis: DURABLE AGENT RUNTIMES SURVIVE RESTARTS. AI Agents

Durable Agent Runtimes: Survive Restarts Without Calling It "Memory"

A durable agent runtime stores resume-correct execution progress outside the process so a crash or human wait continues the same named run—not a new one.

20 MIN
A cracked amber fuse. Thesis: PRE EXECUTION POLICY GATES KILL. AI Agents

Pre-Execution Policy Gates: The Kill Switch That Lives Outside the Prompt

Your agent’s kill switch is a pre-execution policy gate outside the prompt: allow, deny, or pending-approval before side effects — fail closed on outages.

16 MIN
A violet scale. Thesis: LLM AS JUDGE RELIABILITY CALIBRATE. AI Agents

LLM-as-Judge Reliability: Calibrate the Scorer Before You Trust the Score

An LLM judge is a noisy instrument, not ground truth. Calibrate against human labels, kill position and verbosity bias, re-check after model or rubric changes.

19 MIN
Amber node beads on a dark rail. Thesis: AGENT LOOP LLM WORKFLOW PICK. AI Agents

Agent Loop vs LLM-in-Workflow: Pick the Shape That Matches the Uncertainty

Most jobs need a workflow with one LLM step, not an agent loop. Use a three-tier test — certainty, branching, blast radius — then prove hybrid is enough.

18 MIN
FAQ

What should you know about this tag archive?

What is this tag archive?

This is the agents tag archive on Will's Journal: every published post that shares this tag, listed in one place.

Is this a guide to the topic, or a list of posts?

A list of posts. This page is a collection, not a topic essay — the 41 posts below are the agents cluster on Will's Journal.