LangGraph vs CrewAI vs a Custom Loop: Choose Control, Not Fashion
Choose LangGraph, CrewAI, or a custom loop by control and durability, then prove the pick on one golden set and a fixed cost band—not fashion rankings.
William Spurlock Founder — Spurlock Studios Updated 20 MIN
LangGraph vs CrewAI vs writing the agent loop yourself is not a beauty contest. For production, pick the abstraction that matches how much control, durability, and auditability you need — then prove the choice on the same golden set and the same cost band.
This spoke belongs to the Agentic Systems Operating Manual. It assumes you already know when not to build an agent. If the path is known, you do not need this comparison. You need a workflow.
The short answer
- LangGraph wins when control flow must be explicit: branches, cycles, checkpoints, and human interrupts you can evidence.
- CrewAI wins when the work maps cleanly to roles and tasks and you need a working multi-agent shape fast — including production crews when the metaphor fits, with Flows when order must be fixed.
- Custom loop wins when the job is a thin tool loop with your own policy, eval, and persistence — and you refuse framework tax you will not use.
- Never rank by GitHub vibes. Rank by pass rate, cost per pass, escalate rate, and time-to-debug on your golden set.
- MCP does not replace any of these. MCP is a tool protocol. These options are orchestration choices.
What problem does each abstraction solve?
LangChain’s own agent docs define the job without a brand: an agent is a model calling tools in a loop until a task is complete, and a harness is everything around that loop (LangChain agents). LangGraph, CrewAI, and a custom loop are three harnesses. They are not three religions.
| Option | Core metaphor | You get | You pay |
|---|---|---|---|
| LangGraph | Explicit state graph | Nodes, edges, typed state, checkpointers, interrupt() HITL | You design the graph; resume rules are strict |
| CrewAI | Roles, tasks, crews (+ Flows) | Fast multi-agent collaboration; optional deterministic Flows | Higher-level magic; extra manager/planner calls if you turn them on |
| Custom loop | Your code owns the loop | Minimal deps; exact policy, eval, and budget wiring | You build persistence, HITL, and resume yourself |
As of August 2026, neither LangGraph nor CrewAI is “dead” or demo-only. Both ship actively maintained docs and production primitives. The failure mode is picking the wrong posture, not picking a corpse.
- I can name the durability or HITL feature I need this month
- I can stub tools and run offline evals without a vendor cloud
- I can emit run, tool, and eval spans an operator can read
- I can pin versions and re-run a golden set after upgrades
If every box stays unchecked, write the loop. Framework fashion is expensive.
How much control do you actually need?
Ask four questions before you open a tutorial:
- Must a regulator, auditor, or ops lead see the exact branch taken?
- Must a run pause for a human and resume hours later without losing state?
- Are there cycles (revise → evaluate → act) that are part of the product, not a hack?
- Will you outgrow a role/task metaphor within one quarter?
| Need | Lean toward |
|---|---|
| Yes to 1–3 | LangGraph, or custom with an equivalent checkpoint and HITL contract |
| Mostly collaboration, deadline pressure, role mapping is natural | CrewAI (Crews for open work; Flows when order must be fixed) |
| No to all four; one agent, few tools, short runs | Custom loop |
LangGraph’s Graph API is a state schema plus nodes plus edges (LangGraph Graph API). CrewAI’s intro tells you to start production apps with a Flow and drop a Crew in only when a step needs autonomous collaboration (CrewAI introduction). A custom loop is the same state names in your code.
Bravery is not a framework. Control is a product requirement.
When does framework tax exceed the benefit?
Framework tax shows up as extra LLM calls, opaque mid-run state, upgrade churn, and engineers debugging the library instead of the job.
CrewAI makes some of that tax explicit. A hierarchical process requires a manager_llm or manager_agent (CrewAI crews). Turn on planning and the library sends crew data to an AgentPlanner before each iteration and injects that plan into every task description — another model call you did not budget. Those are not bugs. They are the price of a manager metaphor.
LangGraph’s tax is different: you own the graph. If you compile without a checkpointer and then claim human-in-the-loop, you bought a drawing, not a resume contract (LangGraph checkpointers).
| Tax | Where it hides | What you measure |
|---|---|---|
| Manager / planner tokens | Hierarchical crews, planning=True | Extra calls per run vs a single-agent baseline |
| Resume fiction | Graph with no checkpointer; Flow that blocks on console input | Hours-later resume on a killed process |
| Upgrade drift | Unpinned framework + unpinned model | Golden-set delta after pip |
| Debug theater | Pretty crew transcript, no tool/eval spans | Minutes from “CRM write wrong” to the offending node |
Use this gate before adopting anything heavier than a thin harness:
- Name the control feature you cannot ship without (checkpoint, interrupt, role split).
- Price the extra tokens that feature will spend on the golden set.
- Confirm you can stub tools and fail closed.
- If you cannot do 1–3 this week, stay custom.
What does LangGraph give you in production?
LangGraph models the agent as a state machine you can draw. Official docs split the machine into three parts: shared State, Nodes that update it, and Edges that pick the next node (LangGraph Graph API). Production teams care about three primitives that are first-class in current docs:
- Checkpointers — snapshot state after each super-step; threads keyed by
thread_id(checkpointers) interrupt()— pause inside a node, surface a JSON-serializable payload, resume withCommand(resume=…)(interrupts)- Persistence split — checkpointers for thread-scoped short-term state; stores for cross-thread facts (persistence)
That combination is why LangGraph shows up when runs must survive crashes, human waits, or audit questions. Official checkpointer docs list the jobs they exist for: human-in-the-loop, memory between turns, time travel, and fault tolerance. Time travel is not a slogan. You replay from a prior checkpoint or fork state to try an alternate path (time travel).
| Primitive | What the docs actually say | What breaks if you skip it |
|---|---|---|
thread_id | Required in configurable when a checkpointer is on | Resume cannot load state |
| Super-step checkpoint | Snapshot after each graph tick | You can only resume at those boundaries |
| Pending writes | Successful sibling nodes stay durable if another node fails | You re-run work that already succeeded |
interrupt() + same thread | Resume value becomes the return of interrupt() | HITL is a slide, not a contract |
Resume is stricter than most demos admit. After an interrupt, the node restarts from the top. Code and side effects before the pause run again (Graph API resume notes). Put CRM writes after the interrupt, or make them idempotent. Otherwise a human “approve” double-sends the email.
LangChain’s create_agent factory is a configurable harness on top of this world (LangChain agents). It is a fast start, not a substitute for naming your states. Think in intake, act, evaluate, revise, terminal even if you stay custom. The graph library is optional. The state names are not.
A plain request/response agent often should not use LangGraph. You bought a graph runtime for a one-shot function.
What does CrewAI give you in production?
CrewAI’s posture is role-based collaboration: agents with roles and goals, tasks with expected outputs, crews that run sequentially or hierarchically (CrewAI crews). Tasks require a description and an expected output; they can carry human_input, guardrails, and explicit context from prior tasks (CrewAI tasks). For research → analyze → write → review shapes, the metaphor is productive and you get a working system quickly.
CrewAI’s own intro is blunt about production shape: start with a Flow. Use a Flow for structure, state, and logic. Use a Crew inside a Flow step when a task needs autonomy (CrewAI introduction). Flows are event-driven: @start marks an entry, @listen chains the next method, and state rides on the Flow object (CrewAI Flows).
| CrewAI piece | Official job | When it is the wrong piece |
|---|---|---|
| Sequential crew | Tasks run in listed order | You needed a manager you did not budget |
| Hierarchical crew | Manager assigns and validates; manager_llm or manager_agent required | One write path with one policy gate |
| Flow | Outer state, branches, loops | You only needed one agent and two tools |
human_input on a task | Human reviews the agent’s final answer | You needed mid-node approval, not end-of-task review |
@human_feedback (Flows, 1.8.0+) | Pause a Flow, collect feedback, route on outcomes | You left the default console block in production |
HITL is documented, not imaginary. Task-level human_input can pause a crew in Pending Human Input and resume via webhook in enterprise deploys (CrewAI human-in-the-loop). Flow-level @human_feedback can emit outcomes like approved/rejected and route @listen methods. The same page says the default decorator blocks on console input; production needs an async HumanFeedbackProvider (human feedback in Flows).
Crews also grew a first-party checkpoint: checkpoint=True saves after events such as task_completed, defaulting to .checkpoints/ JSON, with Crew.from_checkpoint() to resume (CrewAI crews — checkpointing). That is real. It is still a different contract than LangGraph’s per-super-step thread_id snapshots and interrupt() inside a node. Verify pause/resume on your deploy before you tell an auditor you have HITL.
Teams that “hate CrewAI in production” often stayed in pure Crews when regulation required a fixed order — then blamed the library for a metaphor mismatch. CrewAI is not “only for demos.” Treat it as a velocity-first abstraction with a control ceiling. Hit the ceiling → migrate the control plane, not your company identity.
When does a custom loop win?
A custom loop is usually:
intake → plan (optional) → tool calls → evaluate → revise or terminal
plus your policy gate, budgets, and traces. Direct provider tool use lives here. LangChain would still call that a harness (LangChain agents). You just own the file.
| Signal | Meaning |
|---|---|
| One agent, ≤8 tools | Framework graph is optional |
| Runs finish in one request window | Checkpoint tax may not pay |
| You already own durable jobs (queues, Durable Objects, Temporal) | Do not buy a second runtime |
| Policy and eval are non-negotiable | Keep them in your code, not buried |
| You can name every state on a whiteboard | You do not need a role metaphor |
Choose custom when the job is a thin tool loop and you already have persistence. Do not choose custom because you want to feel clever. Choose it because the golden set already passes and a graph library would add nodes you will not use.
Custom does not mean careless. It means you own the boring parts on purpose.
- Pin the model and the tool schemas.
- Put the evaluator in the loop before the write.
- Persist run id, tool args, and eval scores yourself.
- Add a queue + resume only when a human wait appears in the real job.
If step 4 never appears, you never needed LangGraph. If roles never appear, you never needed a crew.
How do you compare them on the same golden set?
Fashion rankings invent benchmarks. You should not. Run this bake-off on one frozen set.
| Step | What you lock |
|---|---|
| 1 | Same job types and golden cases (pass/fail criteria frozen) |
| 2 | Same tool stubs or sandboxed tools |
| 3 | Same model pin and temperature policy |
| 4 | Same max turns, budget, and kill switch |
| 5 | Report pass rate, cost per pass, escalate rate, p95 latency, debug minutes per failure |
Decision rule we use on pilots:
- If custom clears the bar, ship custom.
- If LangGraph clears the bar and you need HITL or checkpointing you do not want to rebuild, ship LangGraph.
- If CrewAI clears the bar faster and the job is collaboration-shaped, ship CrewAI — with Flows where order must be proven.
- If two options tie on quality, pick the one with lower cost per pass and faster incident debug.
Do not “average” three frameworks into a chimera. One control plane ships. The losers stay in a branch until the next golden-set miss forces a rematch.
| Metric | Why it is the referee | How it lies if you skip the lock |
|---|---|---|
| Pass rate | Quality on your cases | A vendor demo used different tools |
| Cost per pass | Tokens including manager/planner | You compared a crew to a one-shot |
| Escalate rate | How often a human must finish | You scored “looks done” as a pass |
| Debug minutes | Time from bad write to the node | Pretty logs, no tool ids |
Intuition-only framework merges are how regressions ship. I have built 500+ automations and spent 20,000+ hours on agentic systems. The bake-offs that changed a pick were boring: same stubs, same cases, one spreadsheet. The bake-offs that wasted a week were three repos and no frozen scorer.
- Golden cases frozen before the first framework import
- Tool stubs identical across candidates
- Model pin identical across candidates
- Writes disabled until the eval gate is green
- One person scores; no “it felt better”
Does MCP replace LangGraph or CrewAI?
No. MCP (Model Context Protocol) is an open standard for connecting AI applications to external systems — tools, data sources, and prompt-shaped workflows (MCP intro). The architecture overview is explicit: MCP focuses solely on the protocol for context exchange — it does not dictate how AI applications use LLMs or manage the provided context (MCP architecture).
LangChain’s MCP page matches that split. MCP standardizes how hosts discover and call tools. langchain-mcp-adapters turns those tools into LangChain tools you can hang on an agent (LangChain MCP). You can put the same MCP servers behind LangGraph, CrewAI, or a custom loop. Choosing MCP does not choose your orchestration layer.
| Layer | Job | Not the job |
|---|---|---|
| MCP | Discover and call tools, resources, prompts | Branch, persist, stop, approve |
| LangGraph / CrewAI / custom | When to call, how to branch, how to stop | Being a portable tool bus |
| Evaluator | Pass/fail on the golden set | Picking a vendor |
If someone says “we standardized on MCP, so we do not need LangGraph,” they mixed up the socket and the state machine. Keep tool servers portable. Keep the control plane honest.
How do human-in-the-loop and checkpointing differ?
If HITL is a compliance requirement, treat checkpoint plus resume as a day-one acceptance test — not a slide.
| Concern | LangGraph | CrewAI | Custom |
|---|---|---|---|
| Pause for approval | First-class interrupt() + checkpointer (interrupts) | Task human_input; Flow @human_feedback 1.8.0+; enterprise webhooks | You implement queue + resume |
| Survive process death | Persistent checkpointer (Postgres for production; memory/SQLite are demo-grade in the reference table) | Crew checkpoint= JSON/SQLite; Flow state when HumanFeedbackPending is raised | Your job system owns it |
| Replay / time-travel | Checkpoint history is a design goal (time travel) | Crew.from_checkpoint(); rebuild the rest from logs | You build it or you do not |
| Evidence for auditors | Graph + StateSnapshot per super-step | Task outputs + your logs + optional crew checkpoints | Whatever you logged |
| Mid-node vs end-of-task | Interrupt inside a node | Default human review is end-of-task; Flow decorator is a step boundary | You choose |
Acceptance test we run before anyone says “we have HITL”:
- Start a run. Hit the pause. Kill the process.
- Wait longer than a request timeout — hours if that is the real wait.
- Resume on a new process with the same thread or checkpoint id.
- Confirm the write did not fire twice.
- Confirm an auditor can see the branch and the human payload.
If step 3 fails, you do not have HITL. You have a demo that waited on stdin.
What is an honest migration path?
A common, honest path in 2026:
- Week 0–1: Prove the job with CrewAI or a notebook custom loop — tools stubbed, evaluator on.
- Week 2: Freeze golden cases from real failures. Stop adding agents for sport.
- Week 3: If control, HITL, or durability requirements appear, re-express the same states as a LangGraph — or keep custom and add your checkpointer.
- Week 4: Cut over behind the same eval gate. Do not “rewrite and hope.”
| Move | Keep | Change |
|---|---|---|
| Crew → Flow-wrapped crew | Agents, tasks, tools | Outer order becomes @start / @listen |
| Crew → LangGraph | Named states, tool schemas, eval | Tasks become nodes; HITL becomes interrupt() |
| Custom → LangGraph | Policy, eval, budgets | Persistence becomes a checkpointer |
| Any → MCP tools | Control plane | Tool transport becomes MCP |
Migration checklist:
- Map each Crew task to a named state or node
- Move side-effect tools behind the same sandbox and idempotency keys
- Keep prompts versioned; do not rewrite copy and topology in one PR
- Re-run the golden set before enabling writes
- Re-run the HITL kill-and-resume test on the new control plane
The point of migration is a tighter control surface, not a new identity. If the golden set does not move with you, you did not migrate. You started over.
What failure mode should you expect?
What breaks: A hierarchical Crew burns three manager LLM calls, then a worker writes a CRM note that fails a soft criterion nobody scores online. The demo looked great because a human watched the happy path.
What it costs: Token spend without a pass; a sales lead that trusts the agent less; a week of “is it the model?” debugging when the real bug is missing evaluate/revise states.
What you do instead: Put the evaluator in the loop before you add agents. Trace tool calls. Prefer one agent with a hard gate over a crew that improvises order.
A second failure, LangGraph-shaped: you call interrupt() after a non-idempotent send. Official resume behavior restarts the node (Graph API). The human approves. The email goes out twice. The graph was correct. The side effect was not.
| Failure | Stack that invites it | Fix |
|---|---|---|
| Manager tokens, no pass | Hierarchical crew, no eval | One agent + scorer; add roles only after the scorer fails |
| Double write on resume | LangGraph node with pre-interrupt side effects | Interrupt first; write after; idempotency keys either way |
| Console HITL in prod | Default @human_feedback | Async provider + persisted pending state |
| MCP-as-orchestrator | Tool servers with no control plane | Keep MCP; pick a loop |
What is a sane SMB default in 2026?
For most Spurlock Studios SMB pilots:
| Starting point | When |
|---|---|
| Custom loop + pinned model + evaluator | Single job, few tools, writes gated |
| LangGraph | Long waits, multi-step approvals, must resume cleanly |
| CrewAI | Role collaboration is the product and you accept the metaphor |
Default bias: smallest control surface that clears the golden set. Multi-agent fashion is a separate decision — split only when trust, audience, or timing conflicts force it.
| Do this first | Skip until the golden set demands it |
|---|---|
| One agent, stubbed tools, written pass/fail | Hierarchical manager |
| Write sandbox + deny gate | Unrestricted CRM tools |
| Cost cap and kill switch | Unbounded revise loops |
| Pin versions | “Latest” on every deploy |
If you cannot name the evaluator, the sandbox, the stop condition, and the kill switch, you are not choosing a framework yet. You are choosing a costume. The operating manual is the parent for that bar. This page only picks the harness after the bar exists.
A useful tell: if the first slide in the design review is a logo, you started in the wrong place. The first slide should be the golden-set score and the resume test.
How does Spurlock choose on a pilot?
On a $1,500 · 5-day agentic pilot we do not start with a framework bake-off for sport. We:
- Lock the job, tools, and evaluator criteria
- Ship the thinnest loop that can fail safely
- Introduce LangGraph only when durability or HITL shows up in the real job
- Use CrewAI when the customer’s process is already a crew of humans and the mapping is honest
- Keep the golden set and cost band as the referee
Framework choice is a control decision, not a brand affiliation. Continue with the operating manual. If the job should not be an agent at all, stop at when not to build one and stay on /agentic only for the cases that earn a loop.
Paste this into the design doc before anyone opens a tutorial. Fill the left column from the job, not from a ranking thread. If the left column is blank, you are not ready to pick a harness.
| If you need… | Prefer… | Reject… |
|---|---|---|
| Explicit revise/eval cycles you can test | LangGraph or custom state machine | Prompt-only “try again” |
| Fast role-based prototype with real tools | CrewAI | Premature microservices of agents |
| One write path, one policy gate | Custom | Three frameworks “just in case” |
| Multi-client shared tools | MCP servers + any orchestrator | Rewriting tools per host |
| Hours-later human approval | LangGraph checkpointer + interrupt(), or a proven Crew/Flow resume | “The process will wait” |
| Fashion ranking from a blog table | Nothing | Shipping on vibes |
Worked example: lead enrichment agent
Job: Enrich a CRM lead, draft a note, stop for a human if confidence is low.
| Approach | Shape | Likely outcome |
|---|---|---|
| Custom | States: fetch → enrich → draft → evaluate → write or escalate | Fastest path for most SMBs |
| LangGraph | Same states as nodes; interrupt() before write; Postgres checkpointer | Right when humans approve asynchronously |
| CrewAI | Researcher + Writer + Reviewer crew, ideally inside a Flow | Attractive demo; watch manager-token overhead and write authority |
| State | Pass | Fail |
|---|---|---|
| fetch | Record id resolves; fields present | 404 or empty required fields → escalate |
| enrich | External facts cited; no invented title | Hallucinated company → block write |
| draft | Note ≤ N chars; no pricing claims | Soft prose, no facts → revise once, then escalate |
| evaluate | Scorer ≥ threshold | Below threshold → interrupt / human |
| write | Idempotent upsert | Duplicate key → no-op, log |
Failure we have seen in spirit across builds: three roles argue in prompts while none owns the write sandbox. Fix the authority boundary first. Then pick the harness that makes that boundary visible.
Bake-off on this job, not on a public leaderboard:
- Twenty frozen leads. Ten should write. Ten should escalate.
- Same stubbed enrich tool. Same model pin.
- Score write/escalate correctness, token spend, and minutes to debug a bad note.
- Ship the winner. Delete the other two repos.
Anti-patterns
Framework tourism. Rebuilding the same agent in three stacks without a frozen golden set.
Crew for a single tool call. Role theater around crm.update.
LangGraph without a checkpointer when you claim HITL — interrupts need persistence (checkpointers).
“We’ll add evals after the graph looks cool.” The graph is not the product. The pass criteria are.
MCP as the orchestrator. You standardized the screwdriver and forgot the assembly line (MCP architecture).
Console HITL in production. Default Flow feedback blocks on stdin (human feedback in Flows). That is a laptop demo.
- I can point at the golden set file
- I can point at the resume test
- I can point at the cost band
- I cannot point at a Hacker News thread as the reason we picked this
If the last box is the only one you can check, you are shopping. Stop shopping.
Control is the product requirement. Fashion is a feed. The golden set is the only ranking that ships.
FAQ
Is CrewAI only for demos?
No. CrewAI ships real production systems when the role/task metaphor matches the work and you use Flows — or equivalent rails — where order must be proven. Official docs tell you to start production apps with a Flow and drop a Crew in for autonomous steps. It becomes “demo-shaped” when teams skip evaluators, budgets, and write sandboxes. That failure is available in every framework.
When is direct API + thin harness best?
When you have one agent, a small tool set, short-lived runs, and you already own policy, eval, and durability elsewhere. If you are not using graph interrupts or role collaboration, a custom loop is often clearer and cheaper to operate. Ship custom when it clears the same golden set the frameworks would be scored on.
Does MCP replace LangGraph?
No. MCP is a protocol for exposing tools, resources, and prompts to AI hosts. It does not choose when to call a tool, how to branch, or how to stop. LangGraph is an orchestration runtime for stateful agent graphs. Use MCP for portable tool boundaries. Use LangGraph, CrewAI, or custom code for control flow.
How do human-in-the-loop and checkpointing differ across options?
LangGraph treats checkpointers and interrupt() as first-class and requires a thread_id to resume. CrewAI supports task-level human_input, Flow @human_feedback from 1.8.0, crew checkpoints, and enterprise webhooks — verify pause/resume durability for your deploy. Custom means you implement the queue, snapshot, and resume contract yourself, which is fine if you already have a job system.
What’s a sane SMB default in 2026?
Start with a custom loop, or a single LangGraph only if you need resumable HITL. Reach for CrewAI when collaboration-shaped work is real and you will wrap it in a Flow when order must be proven. Prove any choice on one golden set and a cost band before scaling writes.
How does Spurlock choose on a pilot?
We lock the job and evaluator first, ship the thinnest safe loop, and only adopt LangGraph or CrewAI when a concrete control or collaboration requirement appears. The $1,500 pilot on /agentic is built to make that call with evidence, not fashion.
CTA
Pick the control surface that clears your golden set — then harden it.
What questions does this article answer?
- Is CrewAI only for demos?
- No. CrewAI ships real production systems when the role/task metaphor matches the work and you use Flows — or equivalent rails — where order must be proven. Official docs tell you to start production apps with a Flow and drop a Crew in for autonomous steps. It becomes “demo-shaped” when teams skip evaluators, budgets, and write sandboxes. That failure is available in every framework.
- When is direct API + thin harness best?
- When you have one agent, a small tool set, short-lived runs, and you already own policy, eval, and durability elsewhere. If you are not using graph interrupts or role collaboration, a custom loop is often clearer and cheaper to operate. Ship custom when it clears the same golden set the frameworks would be scored on.
- Does MCP replace LangGraph?
- No. MCP is a protocol for exposing tools, resources, and prompts to AI hosts. It does not choose when to call a tool, how to branch, or how to stop. LangGraph is an orchestration runtime for stateful agent graphs. Use MCP for portable tool boundaries. Use LangGraph, CrewAI, or custom code for control flow.
- How do human-in-the-loop and checkpointing differ across options?
- LangGraph treats checkpointers and `interrupt()` as first-class and requires a `thread_id` to resume. CrewAI supports task-level `human_input`, Flow `@human_feedback` from 1.8.0, crew checkpoints, and enterprise webhooks — verify pause/resume durability for your deploy. Custom means you implement the queue, snapshot, and resume contract yourself, which is fine if you already have a job system.
- What's a sane SMB default in 2026?
- Start with a custom loop, or a single LangGraph only if you need resumable HITL. Reach for CrewAI when collaboration-shaped work is real and you will wrap it in a Flow when order must be proven. Prove any choice on one golden set and a cost band before scaling writes.
- How does Spurlock choose on a pilot?
- We lock the job and evaluator first, ship the thinnest safe loop, and only adopt LangGraph or CrewAI when a concrete control or collaboration requirement appears. The **$1,500** pilot on [/agentic](/agentic) is built to make that call with evidence, not fashion.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.