Scoping an Agentic Pilot That Proves Value in Five Days
Scope a 2–6 week agent pilot as one job, real data, an evaluator, and a cage. Five days proves the spike; extra weeks buy access, labels, and a second measure.
William Spurlock Founder — Spurlock Studios Updated 20 MIN
A 2–6 week agent pilot proves one job on real data against written criteria, inside a cage that can stop. Five days is the spike that answers whether the job even works. The extra weeks absorb access, golden-set labels, legal review, and a second measurement pass — not a second job. If you cannot say the work in one sentence, you are not scoping a pilot. You are scoping a platform.
This spoke sits under the Agentic Systems Operating Manual. If the path is already known, stop and read when not to build an agent before you book engineering days.
The short answer
- One sentence job, 20–50 real cases, pass/fail criteria, a tool allowlist, and terminal states (
done,escalate,abort) written before week one. - Five days proves the thin path on that slice. Two to six weeks is the honest calendar once credentials, labeling, and a second score land.
- Extra weeks do not buy extra jobs. They buy uglier cases, a frozen scope, and receipts you can defend.
- Success is a scorecard, not applause. Pick two or three metrics and refuse “it felt magical.”
- Three honest exits: ship path, rescope to automation, or park. All three beat a zombie POC.
What a 2–6 week agent pilot is (and is not)
Is: a time-boxed build that puts one agentic job on a runnable path with an evaluator, a sandbox, and an escalate lane, measured on real inputs.
Is not: a strategy workshop with no artifact; a chat wrapper with your logo; a twelve-integration platform; a promise that Friday ships AGI.
Anthropic’s Building effective agents essay (Dec 19, 2024) still draws the line most teams need: workflows are “systems where LLMs and tools are orchestrated through predefined code paths”; agents are systems where the model “dynamically direct[s] [its] own processes and tool usage.” They tell you to find the simplest solution first — “This might mean not building agentic systems at all.” A pilot that cannot say which of those two you are proving is already lying.
LangGraph’s docs repeat the same physics: workflows have predetermined code paths; agents “define their own processes and tool usage.” If every edge is known before the first ticket arrives, you do not have an agent pilot. You have an automation spike wearing a costume.
| Shape | Fits a 2–6 week pilot | Does not |
|---|---|---|
| One job, one primary system, 1–2 tools | Yes | “Own support” |
| Evaluator + golden slice | Yes | Slide deck + live demo only |
| Internal drafts / staging writes | Yes | Autonomous public sends |
| Kill switch + revision cap | Yes | Open-ended tool loop |
| Decision at the end | Yes | “We’ll keep iterating” with no exit |
If you need architecture across many initiatives, that is a longer engagement. If the path is fully known, buy automation. The brake pedal is when not to build an agent.
Why five days and two-to-six weeks share one scope
Spurlock Studios runs a productized spike at $1,500 · 5 business days on /agentic: one job, your real data, you keep the agent, the fee credits toward a build. That week answers “does this sentence work?” It does not replace a 2–6 week calendar inside a company that still needs legal, a domain reviewer, and a second score on a wider slice.
Treat the clocks as layers, not competing products:
| Clock | What it is for | What it is not for |
|---|---|---|
| Five-day spike | Thin path, first golden slice, first cost sheet | Multi-team rollout |
| 2-week window | Access + labels + one hardening pass | Second job, second system of record |
| 4-week window | Second measure on held-out cases, cage tightening | Company-wide “brain” |
| 6-week window | Residual risks, runbook, go / no-go for a build | Fleet, multi-tenant billing, fine-tunes |
I have spent 20,000+ hours on agentic systems and shipped 500+ automations. The weeks that taught us something true were the ones that refused to grow the job sentence. The weeks that lied were the ones that used “we have more time” as permission to add tools.
OpenAI’s practical guide to building agents is blunt about the climb: customers do better with an incremental approach, a single agent first, and a deterministic solution when that is enough. Their building-agents track starts the same way — one focused agent, then tools, then networks. A 2–6 week pilot that opens with a multi-agent org chart has already skipped the evidence.
- The job sentence is identical in week one and week six
- Extra weeks are booked for access, labels, or a second measure — named in writing
- Nobody can add a tool after the freeze without swapping something of equal size
- The five-day spike, if you run one, feeds the longer window instead of restarting it
How to write the one-sentence job
If you need a paragraph, you have two jobs. Use this template and do not decorate it:
“Given [trigger], produce [artifact] for [audience], such that [criteria], using [systems], and never [hard no].”
Jobs that fit a 2–6 week window:
- “Given a new Tier-2 support ticket, produce an internal triage summary for the on-call lead, such that severity is enum-valid and citations-or-no_match hold, using the help desk plus help center, and never email the customer.”
- “Given a new inbound lead, produce an enriched internal note for sales, such that firmographic fields are null-safe and sourced, using the CRM plus one enrichment API, and never merge or delete leads.”
- “Given a meeting transcript, produce a task list in the tracker with owner guesses, such that each task has a verb and a due-date guess, and never assign or notify until a human confirms.”
Jobs that do not:
- “Own customer support.”
- “Be our sales team.”
- “Replace the ops department.”
- “Stand up an agent platform the whole company can extend.”
| Test | Pass | Fail |
|---|---|---|
| Length | One sentence | A paragraph or a slide |
| Artifact | Named file, row, or draft | “Better experience” |
| Audience | One role | “The business” |
| Hard no | Explicit | “We’ll be careful” |
| Systems | One primary + 1–2 tools | Every SaaS you own |
Run a 45-minute scoping call with a shared doc:
- Each person writes a job sentence silently.
- Compare and merge to one.
- List hard nos.
- List systems.
- Draft five criteria.
- Pick twenty to fifty sample IDs for the golden slice.
If step 2 fails, do not book engineering days. You do not have a pilot. You have a disagreement.
How long should the 2–6 week calendar be?
Pick the shortest window that can finish the sentence honestly. Longer is not safer if the extra days are unstructured.
| Signal | 2 weeks | 4 weeks | 6 weeks |
|---|---|---|---|
| Access already live on staging | Default | Only if the slice is ugly | Rare |
| Legal has not seen a sandbox diagram | Too tight | Default | If counsel needs a second pass |
| Domain reviewer has <2 hours/week | Do not start | Borderline | Still risky |
| Golden slice is 20 clean cases | Fits | Use the extra time for held-out | Waste |
| Golden slice is 50 messy cases | Too tight | Default | If you need a second measure |
| Irreversible actions in scope | Do not start | Still no | Still no — cut them |
Decision rule:
- If credentials and twenty labeled cases exist on day one, book two weeks.
- If legal or IT will take a week to land read access, book four and spend week one on access only.
- If you need a held-out second measure after the first score, book six and freeze scope at the end of week two.
- If you cannot name which of those three you are in, you are not ready to pick a date.
NIST’s AI Risk Management Framework is Govern → Map → Measure → Manage. The July 2024 Generative AI Profile (NIST AI 600-1) applies that loop to generative systems. A 2–6 week pilot that skips Map (what the job is, who is harmed if it is wrong) and Measure (how you will know) is a demo with a longer invoice.
Criteria before tools — or you only have a demo
Write pass/fail acceptance lines on day one. No criteria, no pilot.
Microsoft Foundry’s own how-to on evaluating an agent treats evaluation as the thing you do during development so you can set an acceptance threshold — they use “an 85% task adherence passing rate” as the example, not as a universal law. Google’s agent evaluation docs split the same work into rapid eval (while you change logic), test-case eval (regression on a fixed set), and online monitoring (after you ship). A pilot that only does a live demo is none of those three.
Write criteria as lines a stranger could score:
| Criterion | Type | Pass | Fail |
|---|---|---|---|
| Severity is one of {sev1, sev2, sev3, unknown} | Mechanical | Enum holds | Free text or missing |
Every claim has a citation or no_match | Mechanical | All claims tagged | Bare assertion |
| Internal only — no customer email field populated | Mechanical | Field empty | Send attempted |
| Domain reviewer would file the same severity | Human | Agree or escalate | Silent disagreement |
| Cost per passing run under the written ceiling | Mechanical | At or under | Unbounded retries |
- Five to ten criteria exist as a checklist, not a vibe
- At least half are mechanical (schema, enum, citation, deny)
- A human score is reserved for the taste that machines should not launder
- The worker never marks the run
done— the evaluator does - Targets are written before the first model call
“Feels magical” is not a metric. Neither is “the room clapped.”
Real data, thin slice
Ten to fifty real examples beat a thousand synthetic ones. Anonymize if you must. Keep the ugly edge cases. Pilots on toy data prove toy performance.
| Slice | Use it for | Do not use it for |
|---|---|---|
| 10 cases | Day-one fixture smoke | A go / no-go |
| 20–30 cases | First golden set | Claiming production readiness |
| 40–50 cases | 2–6 week window with a held-out half | Still not a fleet |
| “We’ll find cases later” | Nothing | A booked calendar |
Split the slice on purpose:
- Label 20–30 cases before the worker exists.
- Hold out 10–20 cases the builder does not tune against.
- Score the held-out set in the last week.
- If held-out collapses and the train set looks green, you overfit the demo.
Google’s eval docs treat a fixed test-case set as the regression object. If you keep adding “just one more example” after every fail, you do not have a set. You have a moving target.
Day-one data checklist:
- Read credentials to staging or a prod read replica
- Written list of fields the agent may write
- PII rules in one page
- Rate limits known
- A backup human path if the agent is down
- Sample IDs pulled, not “we’ll query live and see”
Day-one blockers live here. Send the checklist before the window starts. Access delayed is the usual reason “two weeks” becomes five.
Minimum tools and the cage
Allowlist the smallest set that can complete the sentence. Prefer drafts and internal fields over customer-visible sends.
OWASP’s LLM06:2025 Excessive Agency is the security name for a fat toolbelt: damaging actions from unexpected, ambiguous, or manipulated model output. Their mitigations are scope rules, not poetry — minimize extensions, minimize permissions, require user approval on high-impact actions. The December 2025 OWASP Top 10 for Agentic Applications names the same failure in agent language: tool misuse, identity and privilege abuse, unexpected code execution. A pilot that hands the model a shell and a production write is not brave. It is unscoped.
| Tool class | In a 2–6 week pilot | Out |
|---|---|---|
| Read help center / CRM / ticket | Yes, scoped fields | Dump the whole tenant |
| Write an internal note / draft | Yes, with schema | Silent customer email |
| Create a task pending confirm | Yes | Auto-assign and notify |
| Refund, delete, public post | No | No — even at week six |
| Open-ended web browse | Usually no | Yes only if the job is research and the sink is internal |
| Shell / arbitrary code | No | No |
Cage rules that must exist on paper:
- Every tool is named, versioned, and allowlisted.
- Write tools run as a different identity than read tools.
- High-impact actions require a human gate outside the model.
- Revision cap (example: two) and a run budget exist before the first live call.
abortis a first-class terminal state, not an exception log.
Microsoft Foundry’s observability write-up puts evaluation, red teaming, and post-production monitoring in one lifecycle. You do not need their product. You do need the order: measure in development, probe the cage, then watch live traffic. A pilot that only watches the happy-path demo skips the middle.
Who must be in the room before week one
Missing people are how you discover on the last Friday that “severity” meant something else.
| Role | Job | If missing |
|---|---|---|
| Sponsor | Declares the sentence and accepts metrics | The pilot becomes a tour |
| System owner | Grants sandbox credentials | Week one is email |
| Domain reviewer | Labels golden cases and judges edges | Day-last surprise |
| Builder | Implements the thin path | You have a workshop |
| Legal / security (as needed) | Approves drafts-only + staging | The window stalls mid-build |
Hours, not titles:
- Sponsor: one 45-minute kickoff, one mid-window freeze, one readout.
- System owner: access in week one, then on-call for rate limits.
- Domain reviewer: two hours a week to label and disagree.
- Builder: the rest of the calendar.
If a company cannot find a domain reviewer, they are not ready for agents regardless of model hype. I will not pretend a committee of executives can substitute for the person who does the job today.
When legal is nervous, do not argue philosophy. Offer drafts-only, staging credentials, redacted traces, and human gates on writes. Bring a one-page sandbox diagram. Nervous counsel is often unprotected counsel — show the cage.
Week-by-week shape for 2, 4, and 6 weeks
Steal the shape. Do not invent a new methodology mid-window.
Two-week window
| Week | Work | Exit |
|---|---|---|
| 1 | Contract, access, evaluator, fixtures, thin path | First ten cases score |
| 2 | Harden on the remaining slice, cost caps, readout | Scorecard + keep/kill |
Four-week window
| Week | Work | Exit |
|---|---|---|
| 1 | Access, contract, criteria, sample IDs | Scope freeze candidate |
| 2 | Evaluator + thin path on 20 cases | First score |
| 3 | Edge cases, deny rules, logging | Scope freeze (hard) |
| 4 | Held-out score, runbook, readout | Three-way decision |
Six-week window
| Week | Work | Exit |
|---|---|---|
| 1 | Access and legal diagram | Credentials live |
| 2 | Contract, criteria, golden slice | Scope freeze candidate |
| 3 | Thin path + first score | Residual risk list |
| 4 | Cage tightening, cost sheet | Hard freeze |
| 5 | Held-out measure + fail harvest | No new tools |
| 6 | Runbook, build options, readout | Ship / rescope / park |
The five-day Spurlock spike, when you run it, maps onto “thin path + first score.” It is Day 1 contract, Day 2 evaluator, Day 3 worker, Day 4 harden, Day 5 receipts. Timelines assume access lands on day one. Access delayed is how five days becomes eight and two weeks becomes four.
Daily async note, every window:
- What passed
- What failed
- What is blocked
- Whether the job sentence moved (it must not)
Mid-window scope freeze: no new tools after the freeze unless the sentence itself was wrong. End-of-window readout: metrics table, residual risks, next step grounded in what you saw.
Anthropic’s multi-agent research write-up is honest about the last mile: the gap from prototype to production is where most of the work hides, and small errors compound. A 2–6 week pilot that spends week five adding a second agent has learned the wrong lesson from that paper.
In scope vs out of scope
Print this. Argue about it once. Then stop.
| In scope | Out of scope for the pilot |
|---|---|
| One job | Multi-department platform |
| One primary system + 1–2 tools | Every SaaS you own |
| Evaluator harness | Perfect model fine-tunes |
| Internal drafts | Autonomous public sends |
| Thin memory fields | Company-wide “brain” |
| Kill switch + revision cap | Fleet multi-tenant billing |
| Cost sheet from the window | Unlimited token budget |
| Held-out score (4–6 weeks) | “We’ll know it when we see it” |
Out-of-scope items can land on a later build. They should not land mid-pilot as “quick adds.”
Write anti-goals in the same doc:
- Not replacing the team
- Not sending customer email
- Not building a company brain
- Not proving a model vendor
- Not standing up a multi-agent mesh
Anti-goals protect the window when excitement spikes mid-build.
What success looks like on paper
Copy this scorecard. Fill the targets before the first run. Fill actuals at the readout.
| Metric | Example target | Actual | Notes |
|---|---|---|---|
| Golden-set pass rate | ≥ 85% on the written criteria | Microsoft’s evaluate-agent page uses 85% task adherence as an example gate, not a law. Write your own. | |
| Held-out pass rate (4–6 wk) | Within 10 points of golden | Collapse means you overfit | |
| Median revisions to pass | ≤ 2 | Cap exists in the runner | |
| Cost per passing run | ≤ written ceiling | From finance comfort, not a vendor slide | |
| Escalate rate | 10–25% early is fine | Zero escalate often means the cage is fake | |
| Irreversible actions auto-sent | 0 | Non-negotiable | |
| Domain-reviewer agreement | Sampled, written | Disagreement is a criterion bug |
Two or three metrics are enough. Five is fine if they are independent. Twelve is a dashboard looking for a purpose.
NIST AI 600-1’s Measure function is the adult version of this table: you decide what “good” means, you collect evidence, you do not outsource the definition to the model that produced the artifact.
End-state artifacts — if these are missing, the window was a demo:
- Job contract markdown
- Evaluator criteria + golden slice (+ held-out IDs if 4–6 weeks)
- Tool catalog (name, auth, side-effect class)
- State table: intake → act → evaluate → revise → done / escalate / abort
- Runbook for escalate
- Cost sheet from the window
- Build options with ranges grounded in what you saw — not a fantasy deck
“You keep it” means the workflow or runner export, credentials documentation, criteria doc, and escalate runbook. You can run without the builder. Support after the window is a separate conversation. Packaging for the Spurlock spike lives on /agentic.
Red flags that the pilot will lie
Decline or rescope. A false-green pilot is worse than no pilot.
| Red flag | What it produces | What you do |
|---|---|---|
| Success = executive enthusiasm | A clap, then a production incident | Require the scorecard |
| Refusal to allow real data | Toy pass rate | Pause until a slice exists |
| Irreversible actions in week one | Real damage | Cut the tool |
| Job sentence expands daily | A platform, no proof | Freeze or walk |
| No domain reviewer | Last-day surprise | Do not start |
| New tools after the freeze | Unscored surface | Swap, do not grow |
| “We’ll add eval later” | A demo | Stop the clock |
| Multi-job “while we’re in there” | Two half-pilots | Sequence or refuse |
I have watched 35,000+ hours of client busywork disappear when the job was narrow and the criteria were honest. I have also watched pilots burn a quarter because nobody could say no. The second kind does not get a trophy for effort.
Failure mode: the job sentence that grows
This is the failure that looks like progress.
Week one: “Draft an internal triage summary.” Week two: “Also suggest a reply.” Week three: “Also file it in the right queue.” Week four: “Also email the customer if severity is low.” Week six: “Why is the agent emailing the wrong people?”
What breaks: the golden set no longer matches the job. Criteria rot. The escalate lane is bypassed because “the new thing is the point.” Cost per run climbs and nobody can say which add caused it. You leave with a story, not a score.
What it costs: the calendar you already paid for, plus a production-shaped mess you now have to unwind, plus a sponsor who thinks agents “don’t work.”
What you do instead:
- Park the new request on a written lot the same day it appears.
- If the change is required for the original sentence to make sense, swap it for something of equal size. Document the swap in the daily note.
- If the change is a second job, it waits for a second window.
- If the sponsor will not freeze, end the pilot early and keep the artifact. That is a cheaper honesty than a false green.
Growth mid-window is how pilots lie. Freeze after the named freeze point unless the sentence itself was wrong.
After the pilot: three honest outcomes
- Ship path — metrics met on the written slice; quote a build to harden and widen. Do not widen inside the pilot.
- Rescope — the agent was the wrong shape; automation or a human process wins. That is a successful pilot. You learned it on a bounded clock.
- Park — value unclear; you still keep the artifact and the criteria. Revisit when the job or the data changes.
All three beat a zombie POC that never decides.
The parent map for what a build must install next — evaluator, sandbox, state machine, cost caps, escalate — is the operating manual. The five-day spike and the 2–6 week calendar are how you earn the right to open that manual for a real job instead of a slide.
FAQ
What makes a good AI agent pilot project?
One sentence job, real data, explicit evaluator criteria, a minimal sandboxed tool list, revision and budget caps, and written success metrics — delivered as a runnable system, not slides. A good pilot can fail honestly. A bad one can only succeed theatrically.
How do you scope a 2–6 week agent pilot?
Cut to one job, freeze the sentence, build the evaluator first, implement a thin state path, measure on a golden slice, and hold out cases if you have four to six weeks. Use the extra calendar for access, labels, legal, and a second score — not for a second job. Access and a domain reviewer must be available before week one.
Why does Spurlock Studios price the five-day spike at $1,500?
It is enough commitment to use real data and real criteria, bounded enough to decide quickly, and structured so the artifact remains yours with credit toward a full build. The five days prove the thin path. A 2–6 week internal calendar, if you need one, is scoped the same way. Packaging lives on /agentic.
Can we pilot multiple jobs in one 2–6 week window?
Not honestly. Sequence pilots or move to a longer build. Parallel jobs in one window recreate the scope failure: two half-harnesses, one confused scorecard, and no clean exit. One sentence per window.
What do we need ready before week one?
A job-sentence draft, a sample of real inputs (or a date when sample IDs will exist), an API access plan, a domain reviewer with hours on the calendar, and agreement that public sends, refunds, deletes, and other irreversible actions stay out of scope. If legal is in the loop, they need the sandbox diagram before credentials move.
How does pilot scope connect to the operating manual?
The pilot installs the minimum viable stack from the manual: evaluator, sandbox, state machine, cost caps, escalate. Platform concerns, fleets, and multi-agent meshes come after proof. If you cannot name those layers at the readout, you ran a demo.
CTA
Scope the sentence. Then prove it on a clock that cannot hide.
What questions does this article answer?
- What makes a good AI agent pilot project?
- One sentence job, real data, explicit evaluator criteria, a minimal sandboxed tool list, revision and budget caps, and written success metrics — delivered as a runnable system, not slides. A good pilot can fail honestly. A bad one can only succeed theatrically.
- How do you scope a 2–6 week agent pilot?
- Cut to one job, freeze the sentence, build the evaluator first, implement a thin state path, measure on a golden slice, and hold out cases if you have four to six weeks. Use the extra calendar for access, labels, legal, and a second score — not for a second job. Access and a domain reviewer must be available before week one.
- Why does Spurlock Studios price the five-day spike at $1,500?
- It is enough commitment to use real data and real criteria, bounded enough to decide quickly, and structured so the artifact remains yours with credit toward a full build. The five days prove the thin path. A 2–6 week internal calendar, if you need one, is scoped the same way. Packaging lives on [/agentic](/agentic).
- Can we pilot multiple jobs in one 2–6 week window?
- Not honestly. Sequence pilots or move to a longer build. Parallel jobs in one window recreate the scope failure: two half-harnesses, one confused scorecard, and no clean exit. One sentence per window.
- What do we need ready before week one?
- A job-sentence draft, a sample of real inputs (or a date when sample IDs will exist), an API access plan, a domain reviewer with hours on the calendar, and agreement that public sends, refunds, deletes, and other irreversible actions stay out of scope. If legal is in the loop, they need the sandbox diagram before credentials move.
- How does pilot scope connect to the operating manual?
- The pilot installs the minimum viable stack from the [manual](/blog/agentic-systems-operating-manual): evaluator, sandbox, state machine, cost caps, escalate. Platform concerns, fleets, and multi-agent meshes come after proof. If you cannot name those layers at the readout, you ran a demo.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.