Agentic posts from Will's Journal.
These are the Agentic posts from Will's Journal — writing from this practice, not the Agentic service page.
The Agentic practice: The night shift that does not invent a price — page 4 of this archive.
Which Agentic posts are in this archive?
AI Agents Scoping an Agentic Pilot That Proves Value in Five Days
Scope a 2–6 week agent pilot as one job, real data, an evaluator, and a cage. Five days proves the spike; extra weeks buy access, labels, and a second measure.
AI Agents Pin the Model, Gate the Upgrade: Catch Agent Drift Before Customers Do
Yes—pin production agents to explicit model IDs. Floating aliases change behavior with no deploy. Upgrade only through a golden-set gate you actually run.
AI Agents Observability for Agents: Traces, Scores, and the Dashboard Ops Actually Reads
Provider dashboards miss silent wrongness. Agent observability is traces with redacted tool I/O, evaluator scores, and cost — the screen ops actually reads.
AI Agents Tool Schemas Agents Follow: Descriptions, Enums, and Killing the Omnibus Tool
Agents invent arguments when schemas are vague. Write JSON Schema like agent UX—enums, required fields, property descriptions—and kill the do_anything tool.
AI Agents When Not to Build an Agent (And What to Build Instead)
Skip the agent when the path is known or criteria are mush. A workflow plus one schema-checked LLM step is the default. Build the loop only after that fails.
AI Agents Durable Agent Runtimes: Survive Restarts Without Calling It "Memory"
A durable agent runtime stores resume-correct execution progress outside the process so a crash or human wait continues the same named run—not a new one.
AI Agents Pre-Execution Policy Gates: The Kill Switch That Lives Outside the Prompt
Your agent’s kill switch is a pre-execution policy gate outside the prompt: allow, deny, or pending-approval before side effects — fail closed on outages.
AI Agents LLM-as-Judge Reliability: Calibrate the Scorer Before You Trust the Score
An LLM judge is a noisy instrument, not ground truth. Calibrate against human labels, kill position and verbosity bias, re-check after model or rubric changes.
AI Agents Agent Loop vs LLM-in-Workflow: Pick the Shape That Matches the Uncertainty
Most jobs need a workflow with one LLM step, not an agent loop. Use a three-tier test — certainty, branching, blast radius — then prove hybrid is enough.
What should I know about this Agentic archive?
- What is this Agentic archive?
- These are the Agentic posts from Will's Journal — writing from this practice, not the Agentic service page.
- Is this the Agentic service page?
- No. This is a journal lane archive. The Agentic practice is at /agentic. The night shift that does not invent a price.