evaluation
This is the evaluation tag archive on Will's Journal: every published post that shares this tag, listed in one place.
It is a collection page, not a topic essay — scan the cluster here instead of filtering the full journal index.
Which posts are tagged evaluation?
AI Agents The Evaluator Is the Product
Agent accuracy did not come from a better prompt or a bigger model. It came from separating the thing that does the work from the thing that judges it.
AI Agents The Fractional AI CTO Model: When You Need Architecture, Not Another Chatbot
A fractional AI CTO owns architecture, evaluation standards, and a deploy gate — not a chatbot retainer. Hire one when colliding initiatives lack rails.
AI Agents Scoping an Agentic Pilot That Proves Value in Five Days
Scope a 2–6 week agent pilot as one job, real data, an evaluator, and a cage. Five days proves the spike; extra weeks buy access, labels, and a second measure.
What should you know about this tag archive?
What is this tag archive?
This is the evaluation tag archive on Will's Journal: every published post that shares this tag, listed in one place.
Is this a guide to the topic, or a list of posts?
A list of posts. This page is a collection, not a topic essay — the 3 posts below are the evaluation cluster on Will's Journal.