Spurlock Studios
Contact
Share LinkedIn X
A small stack of coins. Thesis: LANGGRAPH CREWAI CUSTOM LOOP CHOOSE.

LangGraph vs CrewAI vs writing the agent loop yourself is not a beauty contest. For production, pick the abstraction that matches how much control, durability, and auditability you need — then prove the choice on the same golden set and the same cost band.

This spoke belongs to the Agentic Systems Operating Manual. It assumes you already know when not to build an agent. If the path is known, you do not need this comparison. You need a workflow.

The short answer

  • LangGraph wins when control flow must be explicit: branches, cycles, checkpoints, and human interrupts you can evidence.
  • CrewAI wins when the work maps cleanly to roles and tasks and you need a working multi-agent shape fast — including production crews when the metaphor fits, with Flows when order must be fixed.
  • Custom loop wins when the job is a thin tool loop with your own policy, eval, and persistence — and you refuse framework tax you will not use.
  • Never rank by GitHub vibes. Rank by pass rate, cost per pass, escalate rate, and time-to-debug on your golden set.
  • MCP does not replace any of these. MCP is a tool protocol. These options are orchestration choices.

What problem does each abstraction solve?

LangChain’s own agent docs define the job without a brand: an agent is a model calling tools in a loop until a task is complete, and a harness is everything around that loop (LangChain agents). LangGraph, CrewAI, and a custom loop are three harnesses. They are not three religions.

OptionCore metaphorYou getYou pay
LangGraphExplicit state graphNodes, edges, typed state, checkpointers, interrupt() HITLYou design the graph; resume rules are strict
CrewAIRoles, tasks, crews (+ Flows)Fast multi-agent collaboration; optional deterministic FlowsHigher-level magic; extra manager/planner calls if you turn them on
Custom loopYour code owns the loopMinimal deps; exact policy, eval, and budget wiringYou build persistence, HITL, and resume yourself

As of August 2026, neither LangGraph nor CrewAI is “dead” or demo-only. Both ship actively maintained docs and production primitives. The failure mode is picking the wrong posture, not picking a corpse.

  • I can name the durability or HITL feature I need this month
  • I can stub tools and run offline evals without a vendor cloud
  • I can emit run, tool, and eval spans an operator can read
  • I can pin versions and re-run a golden set after upgrades

If every box stays unchecked, write the loop. Framework fashion is expensive.

How much control do you actually need?

Ask four questions before you open a tutorial:

  1. Must a regulator, auditor, or ops lead see the exact branch taken?
  2. Must a run pause for a human and resume hours later without losing state?
  3. Are there cycles (revise → evaluate → act) that are part of the product, not a hack?
  4. Will you outgrow a role/task metaphor within one quarter?
NeedLean toward
Yes to 1–3LangGraph, or custom with an equivalent checkpoint and HITL contract
Mostly collaboration, deadline pressure, role mapping is naturalCrewAI (Crews for open work; Flows when order must be fixed)
No to all four; one agent, few tools, short runsCustom loop

LangGraph’s Graph API is a state schema plus nodes plus edges (LangGraph Graph API). CrewAI’s intro tells you to start production apps with a Flow and drop a Crew in only when a step needs autonomous collaboration (CrewAI introduction). A custom loop is the same state names in your code.

Bravery is not a framework. Control is a product requirement.

When does framework tax exceed the benefit?

Framework tax shows up as extra LLM calls, opaque mid-run state, upgrade churn, and engineers debugging the library instead of the job.

CrewAI makes some of that tax explicit. A hierarchical process requires a manager_llm or manager_agent (CrewAI crews). Turn on planning and the library sends crew data to an AgentPlanner before each iteration and injects that plan into every task description — another model call you did not budget. Those are not bugs. They are the price of a manager metaphor.

LangGraph’s tax is different: you own the graph. If you compile without a checkpointer and then claim human-in-the-loop, you bought a drawing, not a resume contract (LangGraph checkpointers).

TaxWhere it hidesWhat you measure
Manager / planner tokensHierarchical crews, planning=TrueExtra calls per run vs a single-agent baseline
Resume fictionGraph with no checkpointer; Flow that blocks on console inputHours-later resume on a killed process
Upgrade driftUnpinned framework + unpinned modelGolden-set delta after pip
Debug theaterPretty crew transcript, no tool/eval spansMinutes from “CRM write wrong” to the offending node

Use this gate before adopting anything heavier than a thin harness:

  1. Name the control feature you cannot ship without (checkpoint, interrupt, role split).
  2. Price the extra tokens that feature will spend on the golden set.
  3. Confirm you can stub tools and fail closed.
  4. If you cannot do 1–3 this week, stay custom.

What does LangGraph give you in production?

LangGraph models the agent as a state machine you can draw. Official docs split the machine into three parts: shared State, Nodes that update it, and Edges that pick the next node (LangGraph Graph API). Production teams care about three primitives that are first-class in current docs:

  1. Checkpointers — snapshot state after each super-step; threads keyed by thread_id (checkpointers)
  2. interrupt() — pause inside a node, surface a JSON-serializable payload, resume with Command(resume=…) (interrupts)
  3. Persistence split — checkpointers for thread-scoped short-term state; stores for cross-thread facts (persistence)

That combination is why LangGraph shows up when runs must survive crashes, human waits, or audit questions. Official checkpointer docs list the jobs they exist for: human-in-the-loop, memory between turns, time travel, and fault tolerance. Time travel is not a slogan. You replay from a prior checkpoint or fork state to try an alternate path (time travel).

PrimitiveWhat the docs actually sayWhat breaks if you skip it
thread_idRequired in configurable when a checkpointer is onResume cannot load state
Super-step checkpointSnapshot after each graph tickYou can only resume at those boundaries
Pending writesSuccessful sibling nodes stay durable if another node failsYou re-run work that already succeeded
interrupt() + same threadResume value becomes the return of interrupt()HITL is a slide, not a contract

Resume is stricter than most demos admit. After an interrupt, the node restarts from the top. Code and side effects before the pause run again (Graph API resume notes). Put CRM writes after the interrupt, or make them idempotent. Otherwise a human “approve” double-sends the email.

LangChain’s create_agent factory is a configurable harness on top of this world (LangChain agents). It is a fast start, not a substitute for naming your states. Think in intake, act, evaluate, revise, terminal even if you stay custom. The graph library is optional. The state names are not.

A plain request/response agent often should not use LangGraph. You bought a graph runtime for a one-shot function.

What does CrewAI give you in production?

CrewAI’s posture is role-based collaboration: agents with roles and goals, tasks with expected outputs, crews that run sequentially or hierarchically (CrewAI crews). Tasks require a description and an expected output; they can carry human_input, guardrails, and explicit context from prior tasks (CrewAI tasks). For research → analyze → write → review shapes, the metaphor is productive and you get a working system quickly.

CrewAI’s own intro is blunt about production shape: start with a Flow. Use a Flow for structure, state, and logic. Use a Crew inside a Flow step when a task needs autonomy (CrewAI introduction). Flows are event-driven: @start marks an entry, @listen chains the next method, and state rides on the Flow object (CrewAI Flows).

CrewAI pieceOfficial jobWhen it is the wrong piece
Sequential crewTasks run in listed orderYou needed a manager you did not budget
Hierarchical crewManager assigns and validates; manager_llm or manager_agent requiredOne write path with one policy gate
FlowOuter state, branches, loopsYou only needed one agent and two tools
human_input on a taskHuman reviews the agent’s final answerYou needed mid-node approval, not end-of-task review
@human_feedback (Flows, 1.8.0+)Pause a Flow, collect feedback, route on outcomesYou left the default console block in production

HITL is documented, not imaginary. Task-level human_input can pause a crew in Pending Human Input and resume via webhook in enterprise deploys (CrewAI human-in-the-loop). Flow-level @human_feedback can emit outcomes like approved/rejected and route @listen methods. The same page says the default decorator blocks on console input; production needs an async HumanFeedbackProvider (human feedback in Flows).

Crews also grew a first-party checkpoint: checkpoint=True saves after events such as task_completed, defaulting to .checkpoints/ JSON, with Crew.from_checkpoint() to resume (CrewAI crews — checkpointing). That is real. It is still a different contract than LangGraph’s per-super-step thread_id snapshots and interrupt() inside a node. Verify pause/resume on your deploy before you tell an auditor you have HITL.

Teams that “hate CrewAI in production” often stayed in pure Crews when regulation required a fixed order — then blamed the library for a metaphor mismatch. CrewAI is not “only for demos.” Treat it as a velocity-first abstraction with a control ceiling. Hit the ceiling → migrate the control plane, not your company identity.

When does a custom loop win?

A custom loop is usually:

intake → plan (optional) → tool calls → evaluate → revise or terminal

plus your policy gate, budgets, and traces. Direct provider tool use lives here. LangChain would still call that a harness (LangChain agents). You just own the file.

SignalMeaning
One agent, ≤8 toolsFramework graph is optional
Runs finish in one request windowCheckpoint tax may not pay
You already own durable jobs (queues, Durable Objects, Temporal)Do not buy a second runtime
Policy and eval are non-negotiableKeep them in your code, not buried
You can name every state on a whiteboardYou do not need a role metaphor

Choose custom when the job is a thin tool loop and you already have persistence. Do not choose custom because you want to feel clever. Choose it because the golden set already passes and a graph library would add nodes you will not use.

Custom does not mean careless. It means you own the boring parts on purpose.

  1. Pin the model and the tool schemas.
  2. Put the evaluator in the loop before the write.
  3. Persist run id, tool args, and eval scores yourself.
  4. Add a queue + resume only when a human wait appears in the real job.

If step 4 never appears, you never needed LangGraph. If roles never appear, you never needed a crew.

How do you compare them on the same golden set?

Fashion rankings invent benchmarks. You should not. Run this bake-off on one frozen set.

StepWhat you lock
1Same job types and golden cases (pass/fail criteria frozen)
2Same tool stubs or sandboxed tools
3Same model pin and temperature policy
4Same max turns, budget, and kill switch
5Report pass rate, cost per pass, escalate rate, p95 latency, debug minutes per failure

Decision rule we use on pilots:

  1. If custom clears the bar, ship custom.
  2. If LangGraph clears the bar and you need HITL or checkpointing you do not want to rebuild, ship LangGraph.
  3. If CrewAI clears the bar faster and the job is collaboration-shaped, ship CrewAI — with Flows where order must be proven.
  4. If two options tie on quality, pick the one with lower cost per pass and faster incident debug.

Do not “average” three frameworks into a chimera. One control plane ships. The losers stay in a branch until the next golden-set miss forces a rematch.

MetricWhy it is the refereeHow it lies if you skip the lock
Pass rateQuality on your casesA vendor demo used different tools
Cost per passTokens including manager/plannerYou compared a crew to a one-shot
Escalate rateHow often a human must finishYou scored “looks done” as a pass
Debug minutesTime from bad write to the nodePretty logs, no tool ids

Intuition-only framework merges are how regressions ship. I have built 500+ automations and spent 20,000+ hours on agentic systems. The bake-offs that changed a pick were boring: same stubs, same cases, one spreadsheet. The bake-offs that wasted a week were three repos and no frozen scorer.

  • Golden cases frozen before the first framework import
  • Tool stubs identical across candidates
  • Model pin identical across candidates
  • Writes disabled until the eval gate is green
  • One person scores; no “it felt better”

Does MCP replace LangGraph or CrewAI?

No. MCP (Model Context Protocol) is an open standard for connecting AI applications to external systems — tools, data sources, and prompt-shaped workflows (MCP intro). The architecture overview is explicit: MCP focuses solely on the protocol for context exchange — it does not dictate how AI applications use LLMs or manage the provided context (MCP architecture).

LangChain’s MCP page matches that split. MCP standardizes how hosts discover and call tools. langchain-mcp-adapters turns those tools into LangChain tools you can hang on an agent (LangChain MCP). You can put the same MCP servers behind LangGraph, CrewAI, or a custom loop. Choosing MCP does not choose your orchestration layer.

LayerJobNot the job
MCPDiscover and call tools, resources, promptsBranch, persist, stop, approve
LangGraph / CrewAI / customWhen to call, how to branch, how to stopBeing a portable tool bus
EvaluatorPass/fail on the golden setPicking a vendor

If someone says “we standardized on MCP, so we do not need LangGraph,” they mixed up the socket and the state machine. Keep tool servers portable. Keep the control plane honest.

How do human-in-the-loop and checkpointing differ?

If HITL is a compliance requirement, treat checkpoint plus resume as a day-one acceptance test — not a slide.

ConcernLangGraphCrewAICustom
Pause for approvalFirst-class interrupt() + checkpointer (interrupts)Task human_input; Flow @human_feedback 1.8.0+; enterprise webhooksYou implement queue + resume
Survive process deathPersistent checkpointer (Postgres for production; memory/SQLite are demo-grade in the reference table)Crew checkpoint= JSON/SQLite; Flow state when HumanFeedbackPending is raisedYour job system owns it
Replay / time-travelCheckpoint history is a design goal (time travel)Crew.from_checkpoint(); rebuild the rest from logsYou build it or you do not
Evidence for auditorsGraph + StateSnapshot per super-stepTask outputs + your logs + optional crew checkpointsWhatever you logged
Mid-node vs end-of-taskInterrupt inside a nodeDefault human review is end-of-task; Flow decorator is a step boundaryYou choose

Acceptance test we run before anyone says “we have HITL”:

  1. Start a run. Hit the pause. Kill the process.
  2. Wait longer than a request timeout — hours if that is the real wait.
  3. Resume on a new process with the same thread or checkpoint id.
  4. Confirm the write did not fire twice.
  5. Confirm an auditor can see the branch and the human payload.

If step 3 fails, you do not have HITL. You have a demo that waited on stdin.

What is an honest migration path?

A common, honest path in 2026:

  1. Week 0–1: Prove the job with CrewAI or a notebook custom loop — tools stubbed, evaluator on.
  2. Week 2: Freeze golden cases from real failures. Stop adding agents for sport.
  3. Week 3: If control, HITL, or durability requirements appear, re-express the same states as a LangGraph — or keep custom and add your checkpointer.
  4. Week 4: Cut over behind the same eval gate. Do not “rewrite and hope.”
MoveKeepChange
Crew → Flow-wrapped crewAgents, tasks, toolsOuter order becomes @start / @listen
Crew → LangGraphNamed states, tool schemas, evalTasks become nodes; HITL becomes interrupt()
Custom → LangGraphPolicy, eval, budgetsPersistence becomes a checkpointer
Any → MCP toolsControl planeTool transport becomes MCP

Migration checklist:

  • Map each Crew task to a named state or node
  • Move side-effect tools behind the same sandbox and idempotency keys
  • Keep prompts versioned; do not rewrite copy and topology in one PR
  • Re-run the golden set before enabling writes
  • Re-run the HITL kill-and-resume test on the new control plane

The point of migration is a tighter control surface, not a new identity. If the golden set does not move with you, you did not migrate. You started over.

What failure mode should you expect?

What breaks: A hierarchical Crew burns three manager LLM calls, then a worker writes a CRM note that fails a soft criterion nobody scores online. The demo looked great because a human watched the happy path.

What it costs: Token spend without a pass; a sales lead that trusts the agent less; a week of “is it the model?” debugging when the real bug is missing evaluate/revise states.

What you do instead: Put the evaluator in the loop before you add agents. Trace tool calls. Prefer one agent with a hard gate over a crew that improvises order.

A second failure, LangGraph-shaped: you call interrupt() after a non-idempotent send. Official resume behavior restarts the node (Graph API). The human approves. The email goes out twice. The graph was correct. The side effect was not.

FailureStack that invites itFix
Manager tokens, no passHierarchical crew, no evalOne agent + scorer; add roles only after the scorer fails
Double write on resumeLangGraph node with pre-interrupt side effectsInterrupt first; write after; idempotency keys either way
Console HITL in prodDefault @human_feedbackAsync provider + persisted pending state
MCP-as-orchestratorTool servers with no control planeKeep MCP; pick a loop

What is a sane SMB default in 2026?

For most Spurlock Studios SMB pilots:

Starting pointWhen
Custom loop + pinned model + evaluatorSingle job, few tools, writes gated
LangGraphLong waits, multi-step approvals, must resume cleanly
CrewAIRole collaboration is the product and you accept the metaphor

Default bias: smallest control surface that clears the golden set. Multi-agent fashion is a separate decision — split only when trust, audience, or timing conflicts force it.

Do this firstSkip until the golden set demands it
One agent, stubbed tools, written pass/failHierarchical manager
Write sandbox + deny gateUnrestricted CRM tools
Cost cap and kill switchUnbounded revise loops
Pin versions“Latest” on every deploy

If you cannot name the evaluator, the sandbox, the stop condition, and the kill switch, you are not choosing a framework yet. You are choosing a costume. The operating manual is the parent for that bar. This page only picks the harness after the bar exists.

A useful tell: if the first slide in the design review is a logo, you started in the wrong place. The first slide should be the golden-set score and the resume test.

How does Spurlock choose on a pilot?

On a $1,500 · 5-day agentic pilot we do not start with a framework bake-off for sport. We:

  1. Lock the job, tools, and evaluator criteria
  2. Ship the thinnest loop that can fail safely
  3. Introduce LangGraph only when durability or HITL shows up in the real job
  4. Use CrewAI when the customer’s process is already a crew of humans and the mapping is honest
  5. Keep the golden set and cost band as the referee

Framework choice is a control decision, not a brand affiliation. Continue with the operating manual. If the job should not be an agent at all, stop at when not to build one and stay on /agentic only for the cases that earn a loop.

Paste this into the design doc before anyone opens a tutorial. Fill the left column from the job, not from a ranking thread. If the left column is blank, you are not ready to pick a harness.

If you need…Prefer…Reject…
Explicit revise/eval cycles you can testLangGraph or custom state machinePrompt-only “try again”
Fast role-based prototype with real toolsCrewAIPremature microservices of agents
One write path, one policy gateCustomThree frameworks “just in case”
Multi-client shared toolsMCP servers + any orchestratorRewriting tools per host
Hours-later human approvalLangGraph checkpointer + interrupt(), or a proven Crew/Flow resume“The process will wait”
Fashion ranking from a blog tableNothingShipping on vibes

Worked example: lead enrichment agent

Job: Enrich a CRM lead, draft a note, stop for a human if confidence is low.

ApproachShapeLikely outcome
CustomStates: fetch → enrich → draft → evaluate → write or escalateFastest path for most SMBs
LangGraphSame states as nodes; interrupt() before write; Postgres checkpointerRight when humans approve asynchronously
CrewAIResearcher + Writer + Reviewer crew, ideally inside a FlowAttractive demo; watch manager-token overhead and write authority
StatePassFail
fetchRecord id resolves; fields present404 or empty required fields → escalate
enrichExternal facts cited; no invented titleHallucinated company → block write
draftNote ≤ N chars; no pricing claimsSoft prose, no facts → revise once, then escalate
evaluateScorer ≥ thresholdBelow threshold → interrupt / human
writeIdempotent upsertDuplicate key → no-op, log

Failure we have seen in spirit across builds: three roles argue in prompts while none owns the write sandbox. Fix the authority boundary first. Then pick the harness that makes that boundary visible.

Bake-off on this job, not on a public leaderboard:

  1. Twenty frozen leads. Ten should write. Ten should escalate.
  2. Same stubbed enrich tool. Same model pin.
  3. Score write/escalate correctness, token spend, and minutes to debug a bad note.
  4. Ship the winner. Delete the other two repos.

Anti-patterns

Framework tourism. Rebuilding the same agent in three stacks without a frozen golden set.

Crew for a single tool call. Role theater around crm.update.

LangGraph without a checkpointer when you claim HITL — interrupts need persistence (checkpointers).

“We’ll add evals after the graph looks cool.” The graph is not the product. The pass criteria are.

MCP as the orchestrator. You standardized the screwdriver and forgot the assembly line (MCP architecture).

Console HITL in production. Default Flow feedback blocks on stdin (human feedback in Flows). That is a laptop demo.

  • I can point at the golden set file
  • I can point at the resume test
  • I can point at the cost band
  • I cannot point at a Hacker News thread as the reason we picked this

If the last box is the only one you can check, you are shopping. Stop shopping.

Control is the product requirement. Fashion is a feed. The golden set is the only ranking that ships.

FAQ

Is CrewAI only for demos?

No. CrewAI ships real production systems when the role/task metaphor matches the work and you use Flows — or equivalent rails — where order must be proven. Official docs tell you to start production apps with a Flow and drop a Crew in for autonomous steps. It becomes “demo-shaped” when teams skip evaluators, budgets, and write sandboxes. That failure is available in every framework.

When is direct API + thin harness best?

When you have one agent, a small tool set, short-lived runs, and you already own policy, eval, and durability elsewhere. If you are not using graph interrupts or role collaboration, a custom loop is often clearer and cheaper to operate. Ship custom when it clears the same golden set the frameworks would be scored on.

Does MCP replace LangGraph?

No. MCP is a protocol for exposing tools, resources, and prompts to AI hosts. It does not choose when to call a tool, how to branch, or how to stop. LangGraph is an orchestration runtime for stateful agent graphs. Use MCP for portable tool boundaries. Use LangGraph, CrewAI, or custom code for control flow.

How do human-in-the-loop and checkpointing differ across options?

LangGraph treats checkpointers and interrupt() as first-class and requires a thread_id to resume. CrewAI supports task-level human_input, Flow @human_feedback from 1.8.0, crew checkpoints, and enterprise webhooks — verify pause/resume durability for your deploy. Custom means you implement the queue, snapshot, and resume contract yourself, which is fine if you already have a job system.

What’s a sane SMB default in 2026?

Start with a custom loop, or a single LangGraph only if you need resumable HITL. Reach for CrewAI when collaboration-shaped work is real and you will wrap it in a Flow when order must be proven. Prove any choice on one golden set and a cost band before scaling writes.

How does Spurlock choose on a pilot?

We lock the job and evaluator first, ship the thinnest safe loop, and only adopt LangGraph or CrewAI when a concrete control or collaboration requirement appears. The $1,500 pilot on /agentic is built to make that call with evidence, not fashion.

CTA

Pick the control surface that clears your golden set — then harden it.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

Is CrewAI only for demos?
No. CrewAI ships real production systems when the role/task metaphor matches the work and you use Flows — or equivalent rails — where order must be proven. Official docs tell you to start production apps with a Flow and drop a Crew in for autonomous steps. It becomes “demo-shaped” when teams skip evaluators, budgets, and write sandboxes. That failure is available in every framework.
When is direct API + thin harness best?
When you have one agent, a small tool set, short-lived runs, and you already own policy, eval, and durability elsewhere. If you are not using graph interrupts or role collaboration, a custom loop is often clearer and cheaper to operate. Ship custom when it clears the same golden set the frameworks would be scored on.
Does MCP replace LangGraph?
No. MCP is a protocol for exposing tools, resources, and prompts to AI hosts. It does not choose when to call a tool, how to branch, or how to stop. LangGraph is an orchestration runtime for stateful agent graphs. Use MCP for portable tool boundaries. Use LangGraph, CrewAI, or custom code for control flow.
How do human-in-the-loop and checkpointing differ across options?
LangGraph treats checkpointers and `interrupt()` as first-class and requires a `thread_id` to resume. CrewAI supports task-level `human_input`, Flow `@human_feedback` from 1.8.0, crew checkpoints, and enterprise webhooks — verify pause/resume durability for your deploy. Custom means you implement the queue, snapshot, and resume contract yourself, which is fine if you already have a job system.
What's a sane SMB default in 2026?
Start with a custom loop, or a single LangGraph only if you need resumable HITL. Reach for CrewAI when collaboration-shaped work is real and you will wrap it in a Flow when order must be proven. Prove any choice on one golden set and a cost band before scaling writes.
How does Spurlock choose on a pilot?
We lock the job and evaluator first, ship the thinnest safe loop, and only adopt LangGraph or CrewAI when a concrete control or collaboration requirement appears. The **$1,500** pilot on [/agentic](/agentic) is built to make that call with evidence, not fashion.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot