Spurlock Studios
Contact
Share LinkedIn X
Two clipped paper packets. Thesis: AUTOMATE INVOICING ADMIN TASKS AI.

You automate invoicing and admin tasks with AI by putting the model on extract, a named human on approve, and a deterministic API on send. The graph turns a messy PDF, email, or timesheet into a schema-valid candidate. A person confirms customer, lines, and total. Then n8n creates, finalizes, and notifies with idempotency keys. The model does not own Stripe. A confidence score is not a send button.

This spoke sits under the Production n8n handbook. The ledger spine — draft hold, auto_advance=false, payment reconcile — lives in invoice and ops pipelines. This page owns the AI hop in front of that spine: what the extractor is allowed to fill, what it must leave null, and why a guessed collection rate is not a metric.

Across 600+ automations built and 500+ live, the finance graphs that survive treat the model as a clerk with a form. The graphs that blow up let the clerk mail the bill.

The short answer

  • Extract. Unstructured input → JSON that matches a written schema. Unknown fields are null, not prose.
  • Approve. A human sees source, lines, header total, and a draft link. Approve / Reject / Request edit. Stale cards do not silent-send.
  • Send. Finalize and notify as separate API calls, each with an idempotency key. Replay the failed step, not the whole graph.
  • Measure the loop, not AR folklore. Reject reasons, duplicate invoice count, exception age, hours of retyping on this path.
  • Do not invent a collection rate. If you cannot export the observation window, you do not have a promotion case.
HopWho owns itAllowed to guess?
ExtractModel + schemaNo. null or exception queue
ValidateDeterministic codeNo. Totals either match or they do not
ApproveNamed finance / ops humanJudgment only — not data entry from memory
SendBilling API via n8nNo. Keys, not vibes

The model drafts the form. The ledger writes the money.

How do I automate invoicing and admin tasks with AI?

You do not “turn on AI billing.” You build one loop that every invoice and every admin write has to survive: extract → approve → send. AI belongs in the first hop. Money and customer-visible mail belong in the last hop, after a person (or a written class they signed) says so.

n8n’s own human-in-the-loop for tools list is purchases, external communications, and deletes. Invoice send is all three in slow motion: cash, a customer email, and a record you cannot quietly unsend.

JobAI mayHuman mustGraph must never
Invoice from a signed proposal PDFPull legal name, lines, currency, PO into JSONConfirm entity + total before notifyLet the extractor call /send
Invoice from a usage exportMap meter rows into line items for a closed periodConfirm period boundsBill an open period because the model “filled the dates”
Invoice from approved timeCopy approved hours onlyConfirm the approval stamp is realPull unapproved time because the email sounded urgent
Vendor bill from a PDFCandidate vendor, amount, due date, invoice numberConfirm vendor master + amountPay from the extract
Expense from a receiptMerchant, amount, date, category suggestionConfirm GL / projectReimburse on confidence
Credit-memo candidate from a ticketAmount requested + original invoice id if presentFinance role authorizesVoid because the model “read frustration”
W-9 / vendor onboardingTax fields into a staging rowA person commits vendor masterFile the form from a guessed EIN
Collections reminderDraft copy into needs_reviewHuman send on first cohortsAuto-dunning with invented aging math

Procedure — the only v1 that is honest:

  1. Pick one input class (one mailbox, one Drive folder, one CRM “won” reason).
  2. Write the output schema before you pick a model.
  3. Extract into that schema.
  4. Validate with code: sum(lines) === header_total, currency present, billing email present, source id present.
  5. Create a draft in the billing system with auto_advance=false.
  6. Pause for a human.
  7. On approve: finalize, then send — two calls.
  8. Persist invoice id + extract id on the CRM row.
  9. Failures go to an exception queue with those ids. Replay the failed step.

If step 5 and step 7 are the same node, you built a demo. Stripe will happily email and retry collection when auto_advance stays on. That is Stripe being helpful. It is not your gate.

What is extract → approve → send — and why is the model not the sender?

Extract is structured output. Approve is a logged decision. Send is a side effect with a key. Collapse any two hops and you get the failure mode this post exists to prevent: a guessed line that left the building.

HopInputOutputFailure if skipped
ExtractPDF text, email body, CSV, ticketJSON matching the schema, plus unknown[]Hallucinated SKU, wrong entity, invented rush fee
ApproveCandidate + source link + draft URLapproved / rejected / edit, actor, timestampRubber-stamp of a total nobody checked
SendApproved candidate + invoice idCustomer-visible notify + CRM writeDuplicate invoice, or a send that cannot be replayed cleanly

The extractor is a clerk. Clerks do not mail invoices because the stack of paper “looked complete.”

Decision list — where the model stops:

  1. It may fill fields that exist in the source document.
  2. It may mark a field null and push a reason into unknown[].
  3. It may not invent a line item, a tax rate, a days_until_due, or a collections schedule.
  4. It may not choose collection_method. That is policy.
  5. It may not call create, finalize, or send. Those are n8n HTTP / Stripe nodes after the gate.

n8n’s Information Extractor is built for hop one: unstructured text in, schema out. It is not an agent with Stripe tools. Keep it that way. If you hand the same model a create_invoice tool, you have left this design and entered a write-capable agent. That is a different blast radius and a different post.

Checklist — you still have three hops:

  • Extract node cannot reach live Stripe credentials
  • Approve payload shows the source (PDF page, email id), not only the JSON
  • Send nodes sit after the Wait / Slack approve
  • Reject writes a reason code the extractor can be scored on
  • Edit returns to extract or to a human-fixed JSON — it does not “send the old total and fix later”

A one-node “AI invoice” is a send waiting to happen.

Which invoice and admin jobs belong in the first loop?

Start where volume is high, the source is boring, and a wrong send is still reversible because you have not sent yet. Draft-first is the reversibility. Auto-send is not v1.

Candidatev1?WhyAI’s actual job
Domestic standard SKU from a signed PDF / proposalYesHigh retype; hold state existsLines + entity + currency into JSON
Closed-period usage file with one meter ownerYesMechanical mappingRows → line items; period bounds stay code
Approved time entries already stamped in the time toolYesCopy, do not judge hoursFilter to status=approved only
Payment webhook → mark paidYes, but no AIReconciliation is matching, not extractionSkip the model
Weekly digest of drafts > 48h and open exceptionsYes, but no AIVisibilitySkip the model
First invoice to a new customerExtract + always-human sendEntity riskExtract still helps; send never auto
Custom SOW / one-off SKUExtract + always-humanLines are the productModel proposes; human owns every line
Multi-entity / multi-currency / tax-ID requiredDefer or extra gateWrong entity is a legal messDo not “figure it out”
Auto-send every extract that scores “high”NoConfidence is not a control—
Auto-pay vendors from OCRNoCash out is dual-control—
Auto collections sequence from agingNoYou will invent a rate and then defend it—
AI-written dunning that cites a DSO targetNoNo sourced number, and it is customer contact—

Checklist — v1 scope posted where sales can read it:

  • One customer class, one currency, one billing rail
  • One input class (mailbox or folder or CRM trigger — not all three on day one)
  • Extract + validate + draft + approve + send as distinct nodes
  • Admin cousins (expenses, vendor bills, W-9) get their own schema, same three hops
  • Explicit non-goals: auto-pay, auto-dunning, auto-void

If sales hears “the AI bills whoever signed,” you already have a runbook gap. Say extract-to-draft out loud.

Admin tasks pay rent when they share the loop, not when they share a Slack channel. An expense receipt and an invoice PDF are different schemas. They are the same discipline: structured candidate, human gate, idempotent write. Dumping both into #finance as walls of JSON is how the channel dies — the same mute physics as lead routing. One card, one owner, one reason. Not a firehose.

How do I implement this in n8n?

The rail is n8n. The design is three hops with a claim before every write. Do not start by wiring Stripe. Start by naming the source id you will refuse to process twice.

Reference graph:

  1. Trigger. Email (attachment), Drive, form, or CRM “won.” Capture source_id (message id, file hash, deal id + milestone).
  2. Claim. Insert source_id + schema_version into a store with a uniqueness constraint. If the insert loses, exit. This is the graph’s idempotency, before Stripe’s.
  3. Binary → text. Extract From File (PDF, XLSX, CSV, or text). The Information Extractor wants text; the docs even use {{ $json.text }} after a PDF extract.
  4. Extract. Information Extractor with a schema you wrote, not “generate from a happy JSON example” as your only contract.
  5. Validate. Code / IF: required fields, sum(lines) vs header, currency, email format, period closed if usage. Fail → exception queue. Do not create.
  6. Draft. Stripe create as draft, auto_advance=false, Idempotency-Key = extract_id:draft.
  7. Approve. Wait on webhook ($execution.resumeUrl) or Slack send-and-wait. Show source + totals + draft URL. Auth on the resume URL.
  8. Promote. Key invoiceId:finalize, then invoiceId:send. Two POSTs.
  9. Write-back. CRM: invoice_id, extract_id, approved_by, sent_at.
  10. Errors. Error workflow with the Error Trigger. Page money paths; do not rely on a 2xx that created nothing.
NodeJobNot its job
Extract From FileBytes → text / rowsGuessing what a blurry total “probably” is
Information ExtractorText → JSONCalling Stripe
Code / IF validateChecksums and requiredsRewriting lines to make the hash match
StripeLedger objectBeing your approval log
WaitPause + resumeA two-minute timer you hope a human saw
Error TriggerAutomatic failuresEditor clicks — n8n will not fire those as production alerts
CredentialUsed byMust not reach
Model / extractor keyInformation Extractor onlyStripe live, QBO/Xero pay, mailbox send
Stripe testStaging graphProduction customers
Stripe liveSend hops after approveExtractor, prompts, execution logs
Resume-URL authWait webhookA public Slack paste

If the extractor credential can create an invoice, you do not have three hops. You have one hop with extra prose.

Wait specifics that bite billing graphs:

  • The node offloads execution to the database until resume. That is the point. A 60-second “wait for the founder” is not a gate.
  • $execution.resumeUrl is unique per execution. Bind it to extract_id. A forwarded naked URL is a credential.
  • Turn on Limit Wait Time. Past the limit: hold or reject. Never auto-send because nobody clicked. The overnight cousin of this failure is in when automation fails overnight: silent success, then a surprise in the customer inbox.
  • Put Header / JWT auth on the resume webhook. Wait documents Basic, Header, JWT, or None. None is for demos.

Staging proof — do this before live keys:

  1. Drop one real (redacted) PDF into the trigger.
  2. Confirm the candidate JSON. Confirm unknown[] is honest.
  3. Confirm the customer inbox is empty after draft create.
  4. Approve. Confirm one invoice id, one send, one CRM write.
  5. Replay the trigger. Confirm you still have one invoice.
  6. Reject a second sample. Confirm no send, reason stored.
  7. Force a validate fail (break the total). Confirm no Stripe create.

If step 5 creates a second object, stop. You do not have a loop. You have a retry tax.

How do I write an extraction schema that fails closed?

Write the schema as a contract the model cannot charm its way around. Happy-path JSON examples are a trap: n8n’s Information Extractor treats every field as mandatory when you generate the schema from a JSON example. Mandatory-plus-missing is how you get a guessed email.

Prefer JSON Schema (or attribute descriptions with explicit optional fields) and make “I do not know” representable.

FieldTypeRequired?If missing
source_idstringYes — from the trigger, not the modelDo not extract. You already have it.
customer_legal_namestringYesException. Do not bill a nickname.
billing_emailstringYesException. Do not invent accounts@
currencyenum (USD, …)YesException. Do not assume home currency
po_numberstring or nullNonull + unknown reason
line_items[]array of {description, qty, unit_amount}Yes, min 1Exception
header_totalinteger minor unitsYesException
period_start / period_enddate or nullUsage onlyException if usage and either is null
tax_idstring or nullPolicynull unless you collect it
unknown[]array of {field, reason, quote}Yes (may be empty)If the model is unsure, it goes here
days_until_due—ForbiddenPolicy table, not extract
collection_rate / DSO / ”% likely to pay”—ForbiddenNot a field. Delete it from the prompt

Validation that is not the model:

  • sum(qty * unit_amount) === header_total in integer minor units. No floats.
  • Every unit_amount is an integer. “Twelve hundred” in prose that became 12000 is a checksum fail if the PDF also printed 1,200.00.
  • Currency is in the enum copied from the billing rail, not a free-text “dollars.”
  • JSON Schema additionalProperties: false on the object you will persist. Extra keys are how a “rush fee” sneaks in.
  • source_id on the candidate equals the claimed trigger id. The model does not get to rename the email.
  • Quotes in unknown[] must be substrings of the source text. If the quote is not in the PDF, the extract is fiction.

System prompt rules that belong in the extractor template (n8n appends format instructions after your template — do not fight that; add constraints):

  1. Copy amounts from the document. Do not convert, round, or “fix” formatting.
  2. If a required field is not in the source, return null and explain in unknown[]. Do not substitute.
  3. Do not add line items that are not printed or listed in the source.
  4. Do not output tax, due-date policy, or collection strategy unless those strings appear as facts in the source — and even then, due-date policy is stripped by validate.
  5. Never output card numbers. If a PAN appears, redact and exception-queue. Stripe’s security guide is blunt: card data on your systems explodes PCI scope. Invoice automation carries invoice ids and amounts. Not PANs. Not CVCs.

A guessed required field is a send waiting to happen.

What must the approval card show before anyone clicks send?

If the approver needs five tabs, you designed research. The card is the product. Totals are computed in the workflow. The model does not get to narrate the money in Slack prose.

Card fieldWhy it exists
Customer legal name + CRM linkStops billing the signer instead of the entity
Currency + header total + sum(lines)Two numbers. If they differ, the card should not have an Approve
Line table (description, qty, unit amount)Stops a $0 draft that grew a mystery line
unknown[] quotesMakes the guess visible instead of burying it in JSON
Source link (email / PDF)Approver inspects the document, not the clerk’s summary
Draft invoice URLApprover inspects the ledger object
Extract id + schema versionSupport can replay the same candidate
Mode (first invoice / standard SKU / over $X)Policy visible in the click
Approve / Reject / Request editRejects need a reason code

Checklist — people will actually click this:

  • One screen, two primary actions, a clock (SLA + named backup)
  • Approve disabled in the UI if checksum failed — do not rely on “they will notice”
  • Token / auth on $execution.resumeUrl so a forwarded link cannot finalize a stranger’s invoice
  • Double-click cannot double-send — claim invoiceId:finalize before the POST
  • PTO backup named in the runbook
  • Finance channel, not #general. Volume is how routing bots get muted; billing cards are the same animal

SLA rule, because stale money is still money:

  1. Publish the clock (same day, next business morning — write it).
  2. Past SLA: ping the backup, not the whole company.
  3. Past limit: hold or reject. Wait’s Limit Wait Time is for this. Auto-approve-on-timeout is how a vacation becomes a customer incident.
  4. Overnight extract failures can wait for morning. Overnight send failures after an approve are a page — see automation fails overnight. Auth drift and silent 2xx belong in that triage, not in a weekly retro.

The click is the control. Homework is not a control.

How do idempotency keys stop a second invoice?

Retries are normal. Email triggers replay. Humans double-click. Stripe webhooks are at-least-once and live destinations retry for up to three days. If your only defense is “we probably will not hit that,” you will bill twice.

You need two layers: a business key you store, and Stripe’s idempotency keys on every POST.

LayerKeyStoreReplay behavior
Graph claimsource_id + schema_versionYour DB / Airtable / Data table, uniqueSecond trigger exits before extract
Draft createextract_id:draftStripe header, ≤255 charsSame payload returns the first invoice
FinalizeinvoiceId:finalizeStripe headerDoes not create a sibling invoice
SendinvoiceId:sendStripe headerDoes not mail twice as a new object
Side effectsinvoiceId:crm / invoiceId:slackYour storeOne CRM write, one notify
Payment (later)Stripe event.idUnique insert before CRM “paid”Covered in the ops pipeline spoke — do not skip it

Stripe’s rules, as of their current docs: keys up to 255 characters; they suggest V4 UUIDs or enough entropy; do not put emails or personal identifiers in the key; they may prune keys after at least 24 hours; all POSTs accept keys; GET/DELETE ignore them. Use stable business keys (extract_id:draft), not a random UUID generated inside the retry.

Procedure — keys in the workflow:

  1. Compute extract_id = source_id + schema_version (and attachment hash if the email can carry two PDFs).
  2. Claim it uniquely before the extractor spends a token, or at latest before Stripe create. Extract is cheap. Duplicate drafts are not.
  3. Send Idempotency-Key on create, finalize, and send.
  4. On Stripe 409 / mismatch (same key, different body), go to the exception queue. Do not “make a new key and try again.” That is how you create the second object on purpose.
  5. After 24 hours, Stripe may have forgotten the key. Your stored invoice id is the long-term lock. Resume send with invoiceId:send, never a second create.
Anti-patternWhat happensFix
UUID generated in the Function node on every runEvery retry is a new key → new invoiceDerive from extract_id
Email address in the keyPII in a header, and aliases fork the keyUse ids
One key for create+finalize+sendParameter mismatch on the second hop, or a create replay that includes send fieldsOne key per POST
Claim after sendYou already mailedClaim before create
Replay whole graph after invoice.paidKickoff tasks multiplyFan-out behind paymentId:… keys

Keys are cheaper than refunds. If step “replay trigger” in staging still creates invoice two, you are not ready for live.

Which admin tasks reuse the same loop without becoming a second Slack process?

Admin is not a junk drawer. Each task gets a schema and the same three hops. The destination changes. The discipline does not.

Admin taskExtract fromApprove showsSend / writev1 send?
Expense reportReceipt image / PDFMerchant, amount, date, projectDraft expense in booksNo auto-reimburse
Vendor billVendor PDFVendor, amount, due date, invoice #Draft billNo auto-pay
Time → invoice linesNotes / calendar / time toolHours, project, approved stampDraft invoice linesOnly if time was already approved
Ticket → credit candidateHelp desk threadOriginal invoice id, amount, reasonCredit-note draftFinance role only
W-9 / vendor masterPDFEIN/name match, addressStaging vendor rowHuman commits master data
Renewal reminderCRM + invoice statusCustomer, amount, dateEmail draftHuman send on first cohorts
Close-week packGraph statsCounts, aging, failed runsInternal digestAuto OK — no customer

Shared rules:

  • New schema, new extract_id namespace (expense:, bill:, w9:). Do not reuse invoice keys.
  • Same fail-closed validate. OCR that is unsure on the total does not create a bill. That is the AP cousin of a guessed invoice line — the ops pipeline spoke covers pay-side dual control; this page covers not letting the model skip it.
  • Same mute test as routing: if #finance gets every receipt, they will ignore the invoice that matters.
  • No “AI collections.” A reminder is customer contact. It is hop three. It waits.

Decision list — when an admin task is not a cousin:

  1. If success is taste (brand copy, legal strategy), keep it manual.
  2. If the write hits a bank account, you need dual control on pay, not a smarter extract.
  3. If nobody will own the exception queue, you built a pile.
  4. If the only metric anyone wants is “faster collections,” stop. That number is not in this design.

Retype the form. Do not retype a culture of chasing.

What breaks this in production?

Concrete failure mode:

  1. A proposal PDF prints $1,200.00. OCR / the model emits 12000 (lost decimal, or a comma treated as grouping in the wrong place).
  2. The schema stored amount as a string. There is no integer checksum against sum(lines) and no unknown[] quote of the printed total.
  3. Confidence looks “high.” The graph treats complete JSON as approved, or a timeout auto-sends because Limit Wait Time was wired to continue.
  4. Stripe create runs with auto_advance left at the default. The customer is notified without a human. Or you do have a draft, then a retried email trigger creates a second draft because source_id was never claimed.
  5. Finance spends the week on a credit note and a founder apology. The workflow is turned off. Close goes back to copy-paste. The next three finance automations never get approved.

We do not attach a dollar figure or a collection-rate “hit” to a named client. You can price your cleanup: two invoices or one wrong total, one credit path, one week of manual billing, and a freeze on every money graph. That is the fail cost. The handbook exists so you do not learn it on a live customer.

SkipWhat breaksWhat you do instead
Model can call sendWrong entity / wrong total leavesExtract credentials ≠ Stripe live key
JSON-example schema, all fields mandatoryGuessed email, guessed POOptional + null + unknown[]
No checksum12000 vs $1,200.00Integer minor units, sum === header
auto_advance=trueStripe emails and retries for youfalse; you own finalize/send
No source claimDuplicate drafts on retryUnique source_id before create
Random UUID keysNew key → new invoiceextract_id:draft
Timeout = approveVacation sendLimit Wait Time → hold / reject
Slack as the ledgerMonth-end is archaeologyInvoice id + extract id in CRM
Personal Stripe loginGraph dies when they leaveService seat
Card image in the prompt “for context”PCI scopeRedact; exception queue
Metric = “we’ll collect faster”You will invent a rateSee the next section

Anti-patterns that only show up on this AI loop (not the generic pipeline):

  • Prompt: “If the total looks low, add a rush line.”
  • Prompt: “Estimate likelihood of payment.” Delete that sentence.
  • Using chat history from a sales thread as line items without a signed source.
  • Letting the model set days_until_due from “Net 30” folklore in the email signature.
  • Retrying extract and create because the Wait node was cancelled, without the claim row.

Bravery is not a restore strategy. A checksum and a gate are.

How do I measure whether the loop is working without inventing collection rates?

Measure whether extract is honest, whether humans still catch the misses, and whether retries stay at one invoice. Do not measure “AI collections,” DSO shaved, ”% recovered,” or “invoices paid faster.” Those numbers depend on customers, terms, product, and dunning policy. This graph does not own them. If someone wants a collection rate in the dashboard, they can run a finance report from the ledger after a dated window. They may not attribute it to the extractor.

MetricHow you get itTarget shapeNot a target
Schema-pass rateValidate node pass / all extractsRising as intake gets cleaner100% by letting the model fill gaps
Reject reason mixReason codes on the card“Wrong PO” → add a form field; “invented line” → tighten schemaA single “thumbs down”
Checksum fail countsum !== headerFalling, then rareZero because you stopped checking
Time-to-approvedecidedAt - draftedAt vs SLAInside the published clockHeroics at 11pm
Duplicate invoice countCount of Stripe invoices per extract_idZero“Only a few”
Exception ageQueue timestampHours, not weeksA graveyard you screenshot
Retype hours on this pathSame clock as baseline weekDown on draft assemblyA company-wide productivity %
Auto-send class sizeExport of invoices with mode=autoEmpty in v1; later, one written classAll invoices because extract “is good now”
Number leadership will ask forWhat you say
“What’s our collection rate with AI?”We do not have one. We have a send-control rate.
“How much faster do we get paid?”Not this project’s KPI. Terms and customers drive that.
“What’s the model accuracy?”Reject reasons + checksum fails, dated. Not a single % from a vendor slide.
“ROI?”Hours of retyping removed on this path, minus incident cost. No AR multiplier.

Observation window — dated, exportable:

  • Every extract: extract_id, pass/fail, unknown[] size
  • Every approve: actor, reason, mode
  • Every send: invoice id, idempotency key
  • Duplicate query: COUNT(invoice_id) GROUP BY extract_id having count > 1
  • No field named predicted_pay_date or collection_score

If you cannot export that list, you cannot promote a class to auto-send. A clean demo week is not a window.

Promotion still lives on the invoice pipeline rules: known customer, standard SKU, under a written $X, after this window is boring. The AI loop does not earn a shortcut.

When should I hire vs DIY this automation?

DIY the extract-to-draft loop when the write is reversible and the blast radius is one person with a pause switch. Hire (or book the $500 Automation Audit) when a retry can bill a customer twice, tax/entity is in play, or nobody named will own the queue at 2am.

SituationDIYAudit / build
One currency, one legal entity, standard SKUYes — extract → draft → you click—
You will watch every card for a monthYes—
Staging proof of “replay trigger = still one invoice” already greenYes to go live on drafts—
Live Stripe send on the first weekNoYes
Multi-entity, tax IDs, mixed FXNoYes
Vendor pay or customer refunds in the same graphNoYes — different gates
Personal Stripe / QBO login todayFix seats firstYes if you want it production
Overnight send after approve, no on-callNoYes — page money paths
Leadership wants a collections KPI on the extractorStop and rewrite the scoreboardYes, to keep that number off the graph
Volume is a handful of invoices a month and intake is already cleanMaybe not worth a canvasOnly if the cost is learning the spine

Decision list:

  1. If you cannot pause the workflow in two minutes, you are not in DIY-live. You are in demo.
  2. If finance will not name an approver and a backup, you are not ready to send. Extract-to-draft can still save retyping.
  3. If the only “AI” request is dunning copy that cites a recovery percentage, decline the metric. Build extract-approve-send or build nothing.
  4. If you already have the ledger spine from the ops pipeline spoke and you only need hop one, DIY the extractor against that spine. Do not rebuild Stripe.

DIY is a smaller loop, not a sloppier one. Keys and checksums still ship.

What should I skip if I only have a week?

Skip autonomy theater. Ship one input class through a gate you can replay.

Do this week:

  1. Pick one mailbox or one folder. Write source_id.
  2. Write the schema table (required vs null). Ban collection fields.
  3. Information Extractor + checksum + exception queue.
  4. Draft create, auto_advance=false, extract_id:draft.
  5. Approval card with source + totals. Limit Wait Time → hold.
  6. Forced duplicate test. Forced checksum-fail test.
  7. Runbook: who pauses, who approves, where exceptions live.

Skip this week:

  • Auto-send, even on “high confidence”
  • AP auto-pay
  • AI dunning / “smart collections”
  • Three billing rails writing the same invoice
  • Agent-with-Stripe-tools
  • A dashboard tile named collection rate
  • Every admin cousin on the same canvas

A week of one honest loop beats a month of a model that mails.

When is this not worth doing yet?

Skip the canvas when you cannot name the source of truth for lines, when nobody will click the card, or when volume is already smaller than the cost of a wrong send. AI does not fix a billing process that does not exist.

BlockerWhy the loop failsDo this first
Lines live in three spreadsheetsExtract will average the fictionOne assembly rule (proposal or usage or approved time)
No finance ownerCards rot; timeout becomes policyName approver + backup
Personal billing loginGraph dies on PTOService seat
Tax/entity is tribal knowledgeModel will guess the letterheadWrite the entity table
Volume is tiny and already cleanYou are buying incident risk for minutesStay manual; steal the schema for later
Success = “collect faster”You will invent a ratePick retyping hours or walk away
Cannot pause in two minutesNot productionError workflow + named pause

Go-live pause test — if any box is empty, keep send off:

  • Named human can disable the workflow in two minutes without a deploy
  • Extractor credential cannot see live Stripe
  • livemode / tenant is visible on the approval card
  • Limit Wait Time holds; it does not approve
  • Duplicate query on extract_id is in the Friday pack

If the process changes every sprint, freeze a v1 class (domestic standard SKU) or wait. The extractor will happily encode last week’s exception as this week’s line item.

Worth doing the moment retyping is weekly, the PDF/email already contains the facts, and a human will still own send. That is the whole product.

FAQ

How do I automate invoicing and admin tasks with AI?

Put the model on extract, a named human on approve, and the billing API on send. Schema-validate the candidate (checksum, currency, email, source id), create a draft with auto_advance=false, then finalize and send as separate idempotent POSTs. Admin tasks — expenses, vendor bills, W-9 staging — reuse the same three hops with their own schemas. The model never owns Stripe send.

How do I measure whether automating invoicing and admin tasks with AI is working?

Track schema-pass rate, reject reason codes, checksum fails, time-to-approve versus SLA, duplicate invoices per extract_id (target zero), exception age, and retyping hours on this path. Do not invent a collection rate, DSO change, or ”% recovered.” Those belong to customers and terms, not to the extractor.

What usually fails first when teams try this?

A guessed total or invented line that still looks like valid JSON, then a send without a gate — or a retry that creates a second invoice because source_id was never claimed. Close second: JSON-example schemas that make every field mandatory, so the model fills an email that was never in the PDF. Confidence is not a control.

How long does this take to show results?

You should see retyping drop once extract-to-draft is live and humans are clicking a card that shows source and totals — often inside a focused stretch after credentials and one input class exist. I will not invent a days-to-DSO or studio-wide payback figure. Early proof is one invoice per extract and a reject log you can read.

What should I skip if I only have a week?

Skip auto-send, AP autopay, AI dunning, and a collections dashboard. Do one input class, a fail-closed schema, draft create with an idempotency key, an approval card with Limit Wait Time, and a forced duplicate test. A week of that loop beats a week of a model with Stripe tools.

When is this not worth doing yet?

When line items have no single source of truth, nobody will own the approval SLA, logins are personal, or the KPI you want is a collection rate the graph cannot honestly produce. Fix assembly rules and ownership first. Extract-approve-send is for facts that already exist in a document, not for inventing a billing process.

CTA

Extract the form. Approve the money. Then send once.

If you want extract → approve → send built to production standard, start with the handbook, then use automation or book the $500 Automation Audit.

FAQ

What questions does this article answer?

How do I automate invoicing and admin tasks with AI?
Put the model on extract, a named human on approve, and the billing API on send. Schema-validate the candidate (checksum, currency, email, source id), create a draft with `auto_advance=false`, then finalize and send as separate idempotent POSTs. Admin tasks — expenses, vendor bills, W-9 staging — reuse the same three hops with their own schemas. The model never owns Stripe send.
How do I measure whether automating invoicing and admin tasks with AI is working?
Track schema-pass rate, reject reason codes, checksum fails, time-to-approve versus SLA, duplicate invoices per `extract_id` (target zero), exception age, and retyping hours on this path. Do not invent a collection rate, DSO change, or "% recovered." Those belong to customers and terms, not to the extractor.
What usually fails first when teams try this?
A guessed total or invented line that still looks like valid JSON, then a send without a gate — or a retry that creates a second invoice because `source_id` was never claimed. Close second: JSON-example schemas that make every field mandatory, so the model fills an email that was never in the PDF. Confidence is not a control.
How long does this take to show results?
You should see retyping drop once extract-to-draft is live and humans are clicking a card that shows source and totals — often inside a focused stretch after credentials and one input class exist. I will not invent a days-to-DSO or studio-wide payback figure. Early proof is one invoice per extract and a reject log you can read.
What should I skip if I only have a week?
Skip auto-send, AP autopay, AI dunning, and a collections dashboard. Do one input class, a fail-closed schema, draft create with an idempotency key, an approval card with Limit Wait Time, and a forced duplicate test. A week of that loop beats a week of a model with Stripe tools.
When is this not worth doing yet?
When line items have no single source of truth, nobody will own the approval SLA, logins are personal, or the KPI you want is a collection rate the graph cannot honestly produce. Fix assembly rules and ownership first. Extract-approve-send is for facts that already exist in a document, not for inventing a billing process.
Sources

Last reviewed

More from this lane

Automation

All →
Book the audit