Spurlock Studios
Contact
Share LinkedIn X
A violet ring. Thesis: SCOPING AGENTIC PILOT PROVES VALUE.

A 2–6 week agent pilot proves one job on real data against written criteria, inside a cage that can stop. Five days is the spike that answers whether the job even works. The extra weeks absorb access, golden-set labels, legal review, and a second measurement pass — not a second job. If you cannot say the work in one sentence, you are not scoping a pilot. You are scoping a platform.

This spoke sits under the Agentic Systems Operating Manual. If the path is already known, stop and read when not to build an agent before you book engineering days.

The short answer

  • One sentence job, 20–50 real cases, pass/fail criteria, a tool allowlist, and terminal states (done, escalate, abort) written before week one.
  • Five days proves the thin path on that slice. Two to six weeks is the honest calendar once credentials, labeling, and a second score land.
  • Extra weeks do not buy extra jobs. They buy uglier cases, a frozen scope, and receipts you can defend.
  • Success is a scorecard, not applause. Pick two or three metrics and refuse “it felt magical.”
  • Three honest exits: ship path, rescope to automation, or park. All three beat a zombie POC.

What a 2–6 week agent pilot is (and is not)

Is: a time-boxed build that puts one agentic job on a runnable path with an evaluator, a sandbox, and an escalate lane, measured on real inputs.

Is not: a strategy workshop with no artifact; a chat wrapper with your logo; a twelve-integration platform; a promise that Friday ships AGI.

Anthropic’s Building effective agents essay (Dec 19, 2024) still draws the line most teams need: workflows are “systems where LLMs and tools are orchestrated through predefined code paths”; agents are systems where the model “dynamically direct[s] [its] own processes and tool usage.” They tell you to find the simplest solution first — “This might mean not building agentic systems at all.” A pilot that cannot say which of those two you are proving is already lying.

LangGraph’s docs repeat the same physics: workflows have predetermined code paths; agents “define their own processes and tool usage.” If every edge is known before the first ticket arrives, you do not have an agent pilot. You have an automation spike wearing a costume.

ShapeFits a 2–6 week pilotDoes not
One job, one primary system, 1–2 toolsYes“Own support”
Evaluator + golden sliceYesSlide deck + live demo only
Internal drafts / staging writesYesAutonomous public sends
Kill switch + revision capYesOpen-ended tool loop
Decision at the endYes“We’ll keep iterating” with no exit

If you need architecture across many initiatives, that is a longer engagement. If the path is fully known, buy automation. The brake pedal is when not to build an agent.

Why five days and two-to-six weeks share one scope

Spurlock Studios runs a productized spike at $1,500 · 5 business days on /agentic: one job, your real data, you keep the agent, the fee credits toward a build. That week answers “does this sentence work?” It does not replace a 2–6 week calendar inside a company that still needs legal, a domain reviewer, and a second score on a wider slice.

Treat the clocks as layers, not competing products:

ClockWhat it is forWhat it is not for
Five-day spikeThin path, first golden slice, first cost sheetMulti-team rollout
2-week windowAccess + labels + one hardening passSecond job, second system of record
4-week windowSecond measure on held-out cases, cage tighteningCompany-wide “brain”
6-week windowResidual risks, runbook, go / no-go for a buildFleet, multi-tenant billing, fine-tunes

I have spent 20,000+ hours on agentic systems and shipped 500+ automations. The weeks that taught us something true were the ones that refused to grow the job sentence. The weeks that lied were the ones that used “we have more time” as permission to add tools.

OpenAI’s practical guide to building agents is blunt about the climb: customers do better with an incremental approach, a single agent first, and a deterministic solution when that is enough. Their building-agents track starts the same way — one focused agent, then tools, then networks. A 2–6 week pilot that opens with a multi-agent org chart has already skipped the evidence.

  • The job sentence is identical in week one and week six
  • Extra weeks are booked for access, labels, or a second measure — named in writing
  • Nobody can add a tool after the freeze without swapping something of equal size
  • The five-day spike, if you run one, feeds the longer window instead of restarting it

How to write the one-sentence job

If you need a paragraph, you have two jobs. Use this template and do not decorate it:

“Given [trigger], produce [artifact] for [audience], such that [criteria], using [systems], and never [hard no].”

Jobs that fit a 2–6 week window:

  • “Given a new Tier-2 support ticket, produce an internal triage summary for the on-call lead, such that severity is enum-valid and citations-or-no_match hold, using the help desk plus help center, and never email the customer.”
  • “Given a new inbound lead, produce an enriched internal note for sales, such that firmographic fields are null-safe and sourced, using the CRM plus one enrichment API, and never merge or delete leads.”
  • “Given a meeting transcript, produce a task list in the tracker with owner guesses, such that each task has a verb and a due-date guess, and never assign or notify until a human confirms.”

Jobs that do not:

  • “Own customer support.”
  • “Be our sales team.”
  • “Replace the ops department.”
  • “Stand up an agent platform the whole company can extend.”
TestPassFail
LengthOne sentenceA paragraph or a slide
ArtifactNamed file, row, or draft“Better experience”
AudienceOne role“The business”
Hard noExplicit“We’ll be careful”
SystemsOne primary + 1–2 toolsEvery SaaS you own

Run a 45-minute scoping call with a shared doc:

  1. Each person writes a job sentence silently.
  2. Compare and merge to one.
  3. List hard nos.
  4. List systems.
  5. Draft five criteria.
  6. Pick twenty to fifty sample IDs for the golden slice.

If step 2 fails, do not book engineering days. You do not have a pilot. You have a disagreement.

How long should the 2–6 week calendar be?

Pick the shortest window that can finish the sentence honestly. Longer is not safer if the extra days are unstructured.

Signal2 weeks4 weeks6 weeks
Access already live on stagingDefaultOnly if the slice is uglyRare
Legal has not seen a sandbox diagramToo tightDefaultIf counsel needs a second pass
Domain reviewer has <2 hours/weekDo not startBorderlineStill risky
Golden slice is 20 clean casesFitsUse the extra time for held-outWaste
Golden slice is 50 messy casesToo tightDefaultIf you need a second measure
Irreversible actions in scopeDo not startStill noStill no — cut them

Decision rule:

  1. If credentials and twenty labeled cases exist on day one, book two weeks.
  2. If legal or IT will take a week to land read access, book four and spend week one on access only.
  3. If you need a held-out second measure after the first score, book six and freeze scope at the end of week two.
  4. If you cannot name which of those three you are in, you are not ready to pick a date.

NIST’s AI Risk Management Framework is Govern → Map → Measure → Manage. The July 2024 Generative AI Profile (NIST AI 600-1) applies that loop to generative systems. A 2–6 week pilot that skips Map (what the job is, who is harmed if it is wrong) and Measure (how you will know) is a demo with a longer invoice.

Criteria before tools — or you only have a demo

Write pass/fail acceptance lines on day one. No criteria, no pilot.

Microsoft Foundry’s own how-to on evaluating an agent treats evaluation as the thing you do during development so you can set an acceptance threshold — they use “an 85% task adherence passing rate” as the example, not as a universal law. Google’s agent evaluation docs split the same work into rapid eval (while you change logic), test-case eval (regression on a fixed set), and online monitoring (after you ship). A pilot that only does a live demo is none of those three.

Write criteria as lines a stranger could score:

CriterionTypePassFail
Severity is one of {sev1, sev2, sev3, unknown}MechanicalEnum holdsFree text or missing
Every claim has a citation or no_matchMechanicalAll claims taggedBare assertion
Internal only — no customer email field populatedMechanicalField emptySend attempted
Domain reviewer would file the same severityHumanAgree or escalateSilent disagreement
Cost per passing run under the written ceilingMechanicalAt or underUnbounded retries
  • Five to ten criteria exist as a checklist, not a vibe
  • At least half are mechanical (schema, enum, citation, deny)
  • A human score is reserved for the taste that machines should not launder
  • The worker never marks the run done — the evaluator does
  • Targets are written before the first model call

“Feels magical” is not a metric. Neither is “the room clapped.”

Real data, thin slice

Ten to fifty real examples beat a thousand synthetic ones. Anonymize if you must. Keep the ugly edge cases. Pilots on toy data prove toy performance.

SliceUse it forDo not use it for
10 casesDay-one fixture smokeA go / no-go
20–30 casesFirst golden setClaiming production readiness
40–50 cases2–6 week window with a held-out halfStill not a fleet
“We’ll find cases later”NothingA booked calendar

Split the slice on purpose:

  1. Label 20–30 cases before the worker exists.
  2. Hold out 10–20 cases the builder does not tune against.
  3. Score the held-out set in the last week.
  4. If held-out collapses and the train set looks green, you overfit the demo.

Google’s eval docs treat a fixed test-case set as the regression object. If you keep adding “just one more example” after every fail, you do not have a set. You have a moving target.

Day-one data checklist:

  • Read credentials to staging or a prod read replica
  • Written list of fields the agent may write
  • PII rules in one page
  • Rate limits known
  • A backup human path if the agent is down
  • Sample IDs pulled, not “we’ll query live and see”

Day-one blockers live here. Send the checklist before the window starts. Access delayed is the usual reason “two weeks” becomes five.

Minimum tools and the cage

Allowlist the smallest set that can complete the sentence. Prefer drafts and internal fields over customer-visible sends.

OWASP’s LLM06:2025 Excessive Agency is the security name for a fat toolbelt: damaging actions from unexpected, ambiguous, or manipulated model output. Their mitigations are scope rules, not poetry — minimize extensions, minimize permissions, require user approval on high-impact actions. The December 2025 OWASP Top 10 for Agentic Applications names the same failure in agent language: tool misuse, identity and privilege abuse, unexpected code execution. A pilot that hands the model a shell and a production write is not brave. It is unscoped.

Tool classIn a 2–6 week pilotOut
Read help center / CRM / ticketYes, scoped fieldsDump the whole tenant
Write an internal note / draftYes, with schemaSilent customer email
Create a task pending confirmYesAuto-assign and notify
Refund, delete, public postNoNo — even at week six
Open-ended web browseUsually noYes only if the job is research and the sink is internal
Shell / arbitrary codeNoNo

Cage rules that must exist on paper:

  1. Every tool is named, versioned, and allowlisted.
  2. Write tools run as a different identity than read tools.
  3. High-impact actions require a human gate outside the model.
  4. Revision cap (example: two) and a run budget exist before the first live call.
  5. abort is a first-class terminal state, not an exception log.

Microsoft Foundry’s observability write-up puts evaluation, red teaming, and post-production monitoring in one lifecycle. You do not need their product. You do need the order: measure in development, probe the cage, then watch live traffic. A pilot that only watches the happy-path demo skips the middle.

Who must be in the room before week one

Missing people are how you discover on the last Friday that “severity” meant something else.

RoleJobIf missing
SponsorDeclares the sentence and accepts metricsThe pilot becomes a tour
System ownerGrants sandbox credentialsWeek one is email
Domain reviewerLabels golden cases and judges edgesDay-last surprise
BuilderImplements the thin pathYou have a workshop
Legal / security (as needed)Approves drafts-only + stagingThe window stalls mid-build

Hours, not titles:

  • Sponsor: one 45-minute kickoff, one mid-window freeze, one readout.
  • System owner: access in week one, then on-call for rate limits.
  • Domain reviewer: two hours a week to label and disagree.
  • Builder: the rest of the calendar.

If a company cannot find a domain reviewer, they are not ready for agents regardless of model hype. I will not pretend a committee of executives can substitute for the person who does the job today.

When legal is nervous, do not argue philosophy. Offer drafts-only, staging credentials, redacted traces, and human gates on writes. Bring a one-page sandbox diagram. Nervous counsel is often unprotected counsel — show the cage.

Week-by-week shape for 2, 4, and 6 weeks

Steal the shape. Do not invent a new methodology mid-window.

Two-week window

WeekWorkExit
1Contract, access, evaluator, fixtures, thin pathFirst ten cases score
2Harden on the remaining slice, cost caps, readoutScorecard + keep/kill

Four-week window

WeekWorkExit
1Access, contract, criteria, sample IDsScope freeze candidate
2Evaluator + thin path on 20 casesFirst score
3Edge cases, deny rules, loggingScope freeze (hard)
4Held-out score, runbook, readoutThree-way decision

Six-week window

WeekWorkExit
1Access and legal diagramCredentials live
2Contract, criteria, golden sliceScope freeze candidate
3Thin path + first scoreResidual risk list
4Cage tightening, cost sheetHard freeze
5Held-out measure + fail harvestNo new tools
6Runbook, build options, readoutShip / rescope / park

The five-day Spurlock spike, when you run it, maps onto “thin path + first score.” It is Day 1 contract, Day 2 evaluator, Day 3 worker, Day 4 harden, Day 5 receipts. Timelines assume access lands on day one. Access delayed is how five days becomes eight and two weeks becomes four.

Daily async note, every window:

  • What passed
  • What failed
  • What is blocked
  • Whether the job sentence moved (it must not)

Mid-window scope freeze: no new tools after the freeze unless the sentence itself was wrong. End-of-window readout: metrics table, residual risks, next step grounded in what you saw.

Anthropic’s multi-agent research write-up is honest about the last mile: the gap from prototype to production is where most of the work hides, and small errors compound. A 2–6 week pilot that spends week five adding a second agent has learned the wrong lesson from that paper.

In scope vs out of scope

Print this. Argue about it once. Then stop.

In scopeOut of scope for the pilot
One jobMulti-department platform
One primary system + 1–2 toolsEvery SaaS you own
Evaluator harnessPerfect model fine-tunes
Internal draftsAutonomous public sends
Thin memory fieldsCompany-wide “brain”
Kill switch + revision capFleet multi-tenant billing
Cost sheet from the windowUnlimited token budget
Held-out score (4–6 weeks)“We’ll know it when we see it”

Out-of-scope items can land on a later build. They should not land mid-pilot as “quick adds.”

Write anti-goals in the same doc:

  • Not replacing the team
  • Not sending customer email
  • Not building a company brain
  • Not proving a model vendor
  • Not standing up a multi-agent mesh

Anti-goals protect the window when excitement spikes mid-build.

What success looks like on paper

Copy this scorecard. Fill the targets before the first run. Fill actuals at the readout.

MetricExample targetActualNotes
Golden-set pass rate≥ 85% on the written criteriaMicrosoft’s evaluate-agent page uses 85% task adherence as an example gate, not a law. Write your own.
Held-out pass rate (4–6 wk)Within 10 points of goldenCollapse means you overfit
Median revisions to pass≤ 2Cap exists in the runner
Cost per passing run≤ written ceilingFrom finance comfort, not a vendor slide
Escalate rate10–25% early is fineZero escalate often means the cage is fake
Irreversible actions auto-sent0Non-negotiable
Domain-reviewer agreementSampled, writtenDisagreement is a criterion bug

Two or three metrics are enough. Five is fine if they are independent. Twelve is a dashboard looking for a purpose.

NIST AI 600-1’s Measure function is the adult version of this table: you decide what “good” means, you collect evidence, you do not outsource the definition to the model that produced the artifact.

End-state artifacts — if these are missing, the window was a demo:

  • Job contract markdown
  • Evaluator criteria + golden slice (+ held-out IDs if 4–6 weeks)
  • Tool catalog (name, auth, side-effect class)
  • State table: intake → act → evaluate → revise → done / escalate / abort
  • Runbook for escalate
  • Cost sheet from the window
  • Build options with ranges grounded in what you saw — not a fantasy deck

“You keep it” means the workflow or runner export, credentials documentation, criteria doc, and escalate runbook. You can run without the builder. Support after the window is a separate conversation. Packaging for the Spurlock spike lives on /agentic.

Red flags that the pilot will lie

Decline or rescope. A false-green pilot is worse than no pilot.

Red flagWhat it producesWhat you do
Success = executive enthusiasmA clap, then a production incidentRequire the scorecard
Refusal to allow real dataToy pass ratePause until a slice exists
Irreversible actions in week oneReal damageCut the tool
Job sentence expands dailyA platform, no proofFreeze or walk
No domain reviewerLast-day surpriseDo not start
New tools after the freezeUnscored surfaceSwap, do not grow
“We’ll add eval later”A demoStop the clock
Multi-job “while we’re in there”Two half-pilotsSequence or refuse

I have watched 35,000+ hours of client busywork disappear when the job was narrow and the criteria were honest. I have also watched pilots burn a quarter because nobody could say no. The second kind does not get a trophy for effort.

Failure mode: the job sentence that grows

This is the failure that looks like progress.

Week one: “Draft an internal triage summary.” Week two: “Also suggest a reply.” Week three: “Also file it in the right queue.” Week four: “Also email the customer if severity is low.” Week six: “Why is the agent emailing the wrong people?”

What breaks: the golden set no longer matches the job. Criteria rot. The escalate lane is bypassed because “the new thing is the point.” Cost per run climbs and nobody can say which add caused it. You leave with a story, not a score.

What it costs: the calendar you already paid for, plus a production-shaped mess you now have to unwind, plus a sponsor who thinks agents “don’t work.”

What you do instead:

  1. Park the new request on a written lot the same day it appears.
  2. If the change is required for the original sentence to make sense, swap it for something of equal size. Document the swap in the daily note.
  3. If the change is a second job, it waits for a second window.
  4. If the sponsor will not freeze, end the pilot early and keep the artifact. That is a cheaper honesty than a false green.

Growth mid-window is how pilots lie. Freeze after the named freeze point unless the sentence itself was wrong.

After the pilot: three honest outcomes

  1. Ship path — metrics met on the written slice; quote a build to harden and widen. Do not widen inside the pilot.
  2. Rescope — the agent was the wrong shape; automation or a human process wins. That is a successful pilot. You learned it on a bounded clock.
  3. Park — value unclear; you still keep the artifact and the criteria. Revisit when the job or the data changes.

All three beat a zombie POC that never decides.

The parent map for what a build must install next — evaluator, sandbox, state machine, cost caps, escalate — is the operating manual. The five-day spike and the 2–6 week calendar are how you earn the right to open that manual for a real job instead of a slide.

FAQ

What makes a good AI agent pilot project?

One sentence job, real data, explicit evaluator criteria, a minimal sandboxed tool list, revision and budget caps, and written success metrics — delivered as a runnable system, not slides. A good pilot can fail honestly. A bad one can only succeed theatrically.

How do you scope a 2–6 week agent pilot?

Cut to one job, freeze the sentence, build the evaluator first, implement a thin state path, measure on a golden slice, and hold out cases if you have four to six weeks. Use the extra calendar for access, labels, legal, and a second score — not for a second job. Access and a domain reviewer must be available before week one.

Why does Spurlock Studios price the five-day spike at $1,500?

It is enough commitment to use real data and real criteria, bounded enough to decide quickly, and structured so the artifact remains yours with credit toward a full build. The five days prove the thin path. A 2–6 week internal calendar, if you need one, is scoped the same way. Packaging lives on /agentic.

Can we pilot multiple jobs in one 2–6 week window?

Not honestly. Sequence pilots or move to a longer build. Parallel jobs in one window recreate the scope failure: two half-harnesses, one confused scorecard, and no clean exit. One sentence per window.

What do we need ready before week one?

A job-sentence draft, a sample of real inputs (or a date when sample IDs will exist), an API access plan, a domain reviewer with hours on the calendar, and agreement that public sends, refunds, deletes, and other irreversible actions stay out of scope. If legal is in the loop, they need the sandbox diagram before credentials move.

How does pilot scope connect to the operating manual?

The pilot installs the minimum viable stack from the manual: evaluator, sandbox, state machine, cost caps, escalate. Platform concerns, fleets, and multi-agent meshes come after proof. If you cannot name those layers at the readout, you ran a demo.

CTA

Scope the sentence. Then prove it on a clock that cannot hide.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

What makes a good AI agent pilot project?
One sentence job, real data, explicit evaluator criteria, a minimal sandboxed tool list, revision and budget caps, and written success metrics — delivered as a runnable system, not slides. A good pilot can fail honestly. A bad one can only succeed theatrically.
How do you scope a 2–6 week agent pilot?
Cut to one job, freeze the sentence, build the evaluator first, implement a thin state path, measure on a golden slice, and hold out cases if you have four to six weeks. Use the extra calendar for access, labels, legal, and a second score — not for a second job. Access and a domain reviewer must be available before week one.
Why does Spurlock Studios price the five-day spike at $1,500?
It is enough commitment to use real data and real criteria, bounded enough to decide quickly, and structured so the artifact remains yours with credit toward a full build. The five days prove the thin path. A 2–6 week internal calendar, if you need one, is scoped the same way. Packaging lives on [/agentic](/agentic).
Can we pilot multiple jobs in one 2–6 week window?
Not honestly. Sequence pilots or move to a longer build. Parallel jobs in one window recreate the scope failure: two half-harnesses, one confused scorecard, and no clean exit. One sentence per window.
What do we need ready before week one?
A job-sentence draft, a sample of real inputs (or a date when sample IDs will exist), an API access plan, a domain reviewer with hours on the calendar, and agreement that public sends, refunds, deletes, and other irreversible actions stay out of scope. If legal is in the loop, they need the sandbox diagram before credentials move.
How does pilot scope connect to the operating manual?
The pilot installs the minimum viable stack from the [manual](/blog/agentic-systems-operating-manual): evaluator, sandbox, state machine, cost caps, escalate. Platform concerns, fleets, and multi-agent meshes come after proof. If you cannot name those layers at the readout, you ran a demo.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot