agents
This is the agents tag archive on Will's Journal: every published post that shares this tag, listed in one place.
It is a collection page, not a topic essay — scan the cluster here instead of filtering the full journal index.
Which posts are tagged agents?
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
AI Agents Why Agents Loop on Failed Tools: No-Progress Detection Beats Longer Prompts
Agents loop on failed tools because the harness never detects no-progress. Fingerprint calls, honor retryable:false, cap turns, terminate with a reason code.
AI Agents Agentic Systems: An Operating Manual for Multi-Agent Work That Ships
An agentic system is evaluators, policy gates, sandboxes, and kill switches — not a chat window — so production tool work survives contact with real data.
AI Agents Why Agent Demos Die in Production: Control Gaps, Not Model IQ
Demo success proves a staged happy path. Production fails when the control loop — schemas, auth, evaluators, and kill switches — was never in the harness.
AI Agents What broke when I tried to evaluate an AI agent in production
Production eval failed on four fronts: no gold labels, biased samples, judge drift, eval tools that wrote. Isolate the harness before you trust the live score.
AI Agents Build the Evaluator Before the Agent
If the judge shares the worker's context, you are grading your own homework. Build criteria, evidence, and ceilings first — then grant real tool autonomy.
AI Agents The Evaluator Is the Product
Agent accuracy did not come from a better prompt or a bigger model. It came from separating the thing that does the work from the thing that judges it.
AI Agents How do I know if my AI agent is actually working
You know an AI agent is working when rewrite rate, policy denials, cost, and time-to-done hold. Thumbs and CSAT hide unpaid human editors on real writes.
AI Agents How do I write good tool schemas for AI agents
Write tool schemas as agent UX: honest required fields, enums for closed sets, descriptions that constrain, and one non-overlapping tool per side effect.
AI Agents How do I stop an agent from doing something destructive
Stop destructive agent actions with a blast-radius table, an allowlisted tool set, and dual control on money and delete, enforced in code before the tool runs.
AI Agents How do I test AI agents before they ship
Test AI agents before they ship on a golden set, staged tools, and shadow mode. Pass rate is not the gate — score cost, escalate, and known-bad fails too.
AI Agents How do I manage multiple AI agents in production
Isolate each production agent: credentials, write surface, named owner, and evals. Handoff with typed contracts. Do not share one god-agent across jobs.
AI Agents How do I keep the agent from emailing customers when injected via a ticket
Treat ticket text as untrusted data. Keep send-email off the reader, gate the proposed send in code, and require a human before any customer outbound fires.
AI Agents How do I set policies for what agents can and cannot do
Write a named-owner catalog that returns allow, deny, or pending-approval before every tool, then fail closed if that catalog cannot load. Not legal advice.
AI Agents Single Agent First: Split Only When Trust, Audience, or Timing Conflicts
Start with one agent and many tools. Split only when trust, audience, or timing conflict—and prove that split with pass rate, cost, and escalate rate.
AI Agents Why did loading all my MCP tools blow the context window
Every MCP tool schema you attach is prompt tokens. Loading the whole catalog fills the window before the task starts—send a job-sized subset or a router.
AI Agents How do I monitor AI agents in production
Monitor production agents on traces, tool errors, cost per turn, human rewrite rate, and policy denials — not chatbot thumbs. That is the ops scoreboard.
AI Agents How big should my golden set be before soft-launch
Before soft-launch, size the golden set by coverage — happy, auth fail, empty tool, policy deny, money, ambiguous — not a magic N or vanity pass rate.
AI Agents Why does my agent keep calling the same failed tool
Your agent retries one failed tool because nothing fingerprints the call or caps same-tool executes. Escalate auth and policy on first hit — do not grind.
AI Agents State Machines for Agent Loops: Determinism Where It Matters
Give the model freedom inside a named state. Legal transitions, revision ceilings, and escalate paths keep operations deterministic when tokens are not.
AI Agents MCP vs Native Function Calling: Portability Tax vs Shortest Loop
Native function calling wins for one app's short tool loop; MCP earns the tax when tools must be shared and governed across hosts—not a LangChain swap.
AI Agents RAG That Does Not Lie: Retrieval Contracts for Business Knowledge
Retrieval is search, not truth. Production RAG needs corpus rules, citations, refuse-on-empty, and contradiction handling — or the agent invents policy.
AI Agents LangGraph vs CrewAI vs a Custom Loop: Choose Control, Not Fashion
Choose LangGraph, CrewAI, or a custom loop by control and durability, then prove the pick on one golden set and a fixed cost band—not fashion rankings.
AI Agents Agent Memory Patterns: What to Persist, What to Forget
Agent memory is three stores, not a bigger window: working dies with the run, episodic is searchable history, and the store holds approved facts only.
AI Agents Golden Sets from Production Failures: Turn Bad Runs into Regression Fuel
Turn a bad production agent run into a regression test: harvest the trace, stub the tools, write the expected terminal, and gate deploys on that fixture.
AI Agents Prompt Injection for Tool Agents: Stop Text from Becoming Actions
Stop prompt injection by isolating untrusted email and tickets from write tools, then gating every proposed call in code — never in the system prompt.
AI Agents Cost Controls for Agent Fleets: Budgets, Caps, and Kill Switches
Control agent-fleet cost with runner-enforced caps, policy-versioned prompt caches, and state-based model routing — not a prompt that says to be frugal.
AI Agents Idempotent Agent Tool Writes: Retries Without Double Emails or Double Charges
Timeouts are unknown, not failed. Mint one runtime idempotency key per write intent and reuse it across retries so you cannot double-email or double-charge.
AI Agents The Fractional AI CTO Model: When You Need Architecture, Not Another Chatbot
A fractional AI CTO owns architecture, evaluation standards, and a deploy gate — not a chatbot retainer. Hire one when colliding initiatives lack rails.
AI Agents Scoping an Agentic Pilot That Proves Value in Five Days
Scope a 2–6 week agent pilot as one job, real data, an evaluator, and a cage. Five days proves the spike; extra weeks buy access, labels, and a second measure.
AI Agents Pin the Model, Gate the Upgrade: Catch Agent Drift Before Customers Do
Yes—pin production agents to explicit model IDs. Floating aliases change behavior with no deploy. Upgrade only through a golden-set gate you actually run.
AI Agents Observability for Agents: Traces, Scores, and the Dashboard Ops Actually Reads
Provider dashboards miss silent wrongness. Agent observability is traces with redacted tool I/O, evaluator scores, and cost — the screen ops actually reads.
AI Agents Tool Schemas Agents Follow: Descriptions, Enums, and Killing the Omnibus Tool
Agents invent arguments when schemas are vague. Write JSON Schema like agent UX—enums, required fields, property descriptions—and kill the do_anything tool.
AI Agents When Not to Build an Agent (And What to Build Instead)
Skip the agent when the path is known or criteria are mush. A workflow plus one schema-checked LLM step is the default. Build the loop only after that fails.
AI Agents Durable Agent Runtimes: Survive Restarts Without Calling It "Memory"
A durable agent runtime stores resume-correct execution progress outside the process so a crash or human wait continues the same named run—not a new one.
AI Agents Pre-Execution Policy Gates: The Kill Switch That Lives Outside the Prompt
Your agent’s kill switch is a pre-execution policy gate outside the prompt: allow, deny, or pending-approval before side effects — fail closed on outages.
AI Agents LLM-as-Judge Reliability: Calibrate the Scorer Before You Trust the Score
An LLM judge is a noisy instrument, not ground truth. Calibrate against human labels, kill position and verbosity bias, re-check after model or rubric changes.
AI Agents Agent Loop vs LLM-in-Workflow: Pick the Shape That Matches the Uncertainty
Most jobs need a workflow with one LLM step, not an agent loop. Use a three-tier test — certainty, branching, blast radius — then prove hybrid is enough.
What should you know about this tag archive?
What is this tag archive?
This is the agents tag archive on Will's Journal: every published post that shares this tag, listed in one place.
Is this a guide to the topic, or a list of posts?
A list of posts. This page is a collection, not a topic essay — the 41 posts below are the agents cluster on Will's Journal.