Agentic posts from Will's Journal.
These are the Agentic posts from Will's Journal — writing from this practice, not the Agentic service page.
The Agentic practice: The night shift that does not invent a price — page 2 of this archive.
Which Agentic posts are in this archive?
AI Agents How do I write good tool schemas for AI agents
Write tool schemas as agent UX: honest required fields, enums for closed sets, descriptions that constrain, and one non-overlapping tool per side effect.
AI Agents How do I stop an agent from doing something destructive
Stop destructive agent actions with a blast-radius table, an allowlisted tool set, and dual control on money and delete, enforced in code before the tool runs.
AI Agents How do I test AI agents before they ship
Test AI agents before they ship on a golden set, staged tools, and shadow mode. Pass rate is not the gate — score cost, escalate, and known-bad fails too.
AI Agents How do I manage multiple AI agents in production
Isolate each production agent: credentials, write surface, named owner, and evals. Handoff with typed contracts. Do not share one god-agent across jobs.
AI Agents How do I keep the agent from emailing customers when injected via a ticket
Treat ticket text as untrusted data. Keep send-email off the reader, gate the proposed send in code, and require a human before any customer outbound fires.
AI Agents Sandboxed Tool Use: Letting Agents Act Without Letting Them Loose
Tool use without a sandbox is an API key with opinions. Allowlists, scoped credentials, blast-radius caps, dry-runs, and human gates earn a production write.
AI Agents How do I set policies for what agents can and cannot do
Write a named-owner catalog that returns allow, deny, or pending-approval before every tool, then fail closed if that catalog cannot load. Not legal advice.
AI Agents Single Agent First: Split Only When Trust, Audience, or Timing Conflicts
Start with one agent and many tools. Split only when trust, audience, or timing conflict—and prove that split with pass rate, cost, and escalate rate.
AI Agents Why did quality drop when we didn’t change our prompts
Quality dropped because a pin, tool schema, retrieval corpus, eval set, or traffic mix moved — not because you edited prompts. Isolate, then pin versions.
AI Agents Why did loading all my MCP tools blow the context window
Every MCP tool schema you attach is prompt tokens. Loading the whole catalog fills the window before the task starts—send a job-sized subset or a router.
AI Agents How do I monitor AI agents in production
Monitor production agents on traces, tool errors, cost per turn, human rewrite rate, and policy denials — not chatbot thumbs. That is the ops scoreboard.
AI Agents How big should my golden set be before soft-launch
Before soft-launch, size the golden set by coverage — happy, auth fail, empty tool, policy deny, money, ambiguous — not a magic N or vanity pass rate.
What should I know about this Agentic archive?
- What is this Agentic archive?
- These are the Agentic posts from Will's Journal — writing from this practice, not the Agentic service page.
- Is this the Agentic service page?
- No. This is a journal lane archive. The Agentic practice is at /agentic. The night shift that does not invent a price.