How do I automate invoicing and admin tasks with AI
AI extracts invoice and admin fields into a schema. A human approves the candidate. Then you send with idempotency keys — the model never owns Stripe.
William Spurlock Founder — Spurlock Studios 29 MIN
You automate invoicing and admin tasks with AI by putting the model on extract, a named human on approve, and a deterministic API on send. The graph turns a messy PDF, email, or timesheet into a schema-valid candidate. A person confirms customer, lines, and total. Then n8n creates, finalizes, and notifies with idempotency keys. The model does not own Stripe. A confidence score is not a send button.
This spoke sits under the Production n8n handbook. The ledger spine — draft hold, auto_advance=false, payment reconcile — lives in invoice and ops pipelines. This page owns the AI hop in front of that spine: what the extractor is allowed to fill, what it must leave null, and why a guessed collection rate is not a metric.
Across 600+ automations built and 500+ live, the finance graphs that survive treat the model as a clerk with a form. The graphs that blow up let the clerk mail the bill.
The short answer
- Extract. Unstructured input → JSON that matches a written schema. Unknown fields are
null, not prose. - Approve. A human sees source, lines, header total, and a draft link. Approve / Reject / Request edit. Stale cards do not silent-send.
- Send. Finalize and notify as separate API calls, each with an idempotency key. Replay the failed step, not the whole graph.
- Measure the loop, not AR folklore. Reject reasons, duplicate invoice count, exception age, hours of retyping on this path.
- Do not invent a collection rate. If you cannot export the observation window, you do not have a promotion case.
| Hop | Who owns it | Allowed to guess? |
|---|---|---|
| Extract | Model + schema | No. null or exception queue |
| Validate | Deterministic code | No. Totals either match or they do not |
| Approve | Named finance / ops human | Judgment only — not data entry from memory |
| Send | Billing API via n8n | No. Keys, not vibes |
The model drafts the form. The ledger writes the money.
How do I automate invoicing and admin tasks with AI?
You do not “turn on AI billing.” You build one loop that every invoice and every admin write has to survive: extract → approve → send. AI belongs in the first hop. Money and customer-visible mail belong in the last hop, after a person (or a written class they signed) says so.
n8n’s own human-in-the-loop for tools list is purchases, external communications, and deletes. Invoice send is all three in slow motion: cash, a customer email, and a record you cannot quietly unsend.
| Job | AI may | Human must | Graph must never |
|---|---|---|---|
| Invoice from a signed proposal PDF | Pull legal name, lines, currency, PO into JSON | Confirm entity + total before notify | Let the extractor call /send |
| Invoice from a usage export | Map meter rows into line items for a closed period | Confirm period bounds | Bill an open period because the model “filled the dates” |
| Invoice from approved time | Copy approved hours only | Confirm the approval stamp is real | Pull unapproved time because the email sounded urgent |
| Vendor bill from a PDF | Candidate vendor, amount, due date, invoice number | Confirm vendor master + amount | Pay from the extract |
| Expense from a receipt | Merchant, amount, date, category suggestion | Confirm GL / project | Reimburse on confidence |
| Credit-memo candidate from a ticket | Amount requested + original invoice id if present | Finance role authorizes | Void because the model “read frustration” |
| W-9 / vendor onboarding | Tax fields into a staging row | A person commits vendor master | File the form from a guessed EIN |
| Collections reminder | Draft copy into needs_review | Human send on first cohorts | Auto-dunning with invented aging math |
Procedure — the only v1 that is honest:
- Pick one input class (one mailbox, one Drive folder, one CRM “won” reason).
- Write the output schema before you pick a model.
- Extract into that schema.
- Validate with code:
sum(lines) === header_total, currency present, billing email present, source id present. - Create a draft in the billing system with
auto_advance=false. - Pause for a human.
- On approve: finalize, then send — two calls.
- Persist invoice id + extract id on the CRM row.
- Failures go to an exception queue with those ids. Replay the failed step.
If step 5 and step 7 are the same node, you built a demo. Stripe will happily email and retry collection when auto_advance stays on. That is Stripe being helpful. It is not your gate.
What is extract → approve → send — and why is the model not the sender?
Extract is structured output. Approve is a logged decision. Send is a side effect with a key. Collapse any two hops and you get the failure mode this post exists to prevent: a guessed line that left the building.
| Hop | Input | Output | Failure if skipped |
|---|---|---|---|
| Extract | PDF text, email body, CSV, ticket | JSON matching the schema, plus unknown[] | Hallucinated SKU, wrong entity, invented rush fee |
| Approve | Candidate + source link + draft URL | approved / rejected / edit, actor, timestamp | Rubber-stamp of a total nobody checked |
| Send | Approved candidate + invoice id | Customer-visible notify + CRM write | Duplicate invoice, or a send that cannot be replayed cleanly |
The extractor is a clerk. Clerks do not mail invoices because the stack of paper “looked complete.”
Decision list — where the model stops:
- It may fill fields that exist in the source document.
- It may mark a field
nulland push a reason intounknown[]. - It may not invent a line item, a tax rate, a
days_until_due, or a collections schedule. - It may not choose
collection_method. That is policy. - It may not call create, finalize, or send. Those are n8n HTTP / Stripe nodes after the gate.
n8n’s Information Extractor is built for hop one: unstructured text in, schema out. It is not an agent with Stripe tools. Keep it that way. If you hand the same model a create_invoice tool, you have left this design and entered a write-capable agent. That is a different blast radius and a different post.
Checklist — you still have three hops:
- Extract node cannot reach live Stripe credentials
- Approve payload shows the source (PDF page, email id), not only the JSON
- Send nodes sit after the Wait / Slack approve
- Reject writes a reason code the extractor can be scored on
- Edit returns to extract or to a human-fixed JSON — it does not “send the old total and fix later”
A one-node “AI invoice” is a send waiting to happen.
Which invoice and admin jobs belong in the first loop?
Start where volume is high, the source is boring, and a wrong send is still reversible because you have not sent yet. Draft-first is the reversibility. Auto-send is not v1.
| Candidate | v1? | Why | AI’s actual job |
|---|---|---|---|
| Domestic standard SKU from a signed PDF / proposal | Yes | High retype; hold state exists | Lines + entity + currency into JSON |
| Closed-period usage file with one meter owner | Yes | Mechanical mapping | Rows → line items; period bounds stay code |
| Approved time entries already stamped in the time tool | Yes | Copy, do not judge hours | Filter to status=approved only |
| Payment webhook → mark paid | Yes, but no AI | Reconciliation is matching, not extraction | Skip the model |
| Weekly digest of drafts > 48h and open exceptions | Yes, but no AI | Visibility | Skip the model |
| First invoice to a new customer | Extract + always-human send | Entity risk | Extract still helps; send never auto |
| Custom SOW / one-off SKU | Extract + always-human | Lines are the product | Model proposes; human owns every line |
| Multi-entity / multi-currency / tax-ID required | Defer or extra gate | Wrong entity is a legal mess | Do not “figure it out” |
| Auto-send every extract that scores “high” | No | Confidence is not a control | — |
| Auto-pay vendors from OCR | No | Cash out is dual-control | — |
| Auto collections sequence from aging | No | You will invent a rate and then defend it | — |
| AI-written dunning that cites a DSO target | No | No sourced number, and it is customer contact | — |
Checklist — v1 scope posted where sales can read it:
- One customer class, one currency, one billing rail
- One input class (mailbox or folder or CRM trigger — not all three on day one)
- Extract + validate + draft + approve + send as distinct nodes
- Admin cousins (expenses, vendor bills, W-9) get their own schema, same three hops
- Explicit non-goals: auto-pay, auto-dunning, auto-void
If sales hears “the AI bills whoever signed,” you already have a runbook gap. Say extract-to-draft out loud.
Admin tasks pay rent when they share the loop, not when they share a Slack channel. An expense receipt and an invoice PDF are different schemas. They are the same discipline: structured candidate, human gate, idempotent write. Dumping both into #finance as walls of JSON is how the channel dies — the same mute physics as lead routing. One card, one owner, one reason. Not a firehose.
How do I implement this in n8n?
The rail is n8n. The design is three hops with a claim before every write. Do not start by wiring Stripe. Start by naming the source id you will refuse to process twice.
Reference graph:
- Trigger. Email (attachment), Drive, form, or CRM “won.” Capture
source_id(message id, file hash, deal id + milestone). - Claim. Insert
source_id+schema_versioninto a store with a uniqueness constraint. If the insert loses, exit. This is the graph’s idempotency, before Stripe’s. - Binary → text. Extract From File (
PDF,XLSX,CSV, or text). The Information Extractor wants text; the docs even use{{ $json.text }}after a PDF extract. - Extract. Information Extractor with a schema you wrote, not “generate from a happy JSON example” as your only contract.
- Validate. Code / IF: required fields,
sum(lines)vs header, currency, email format, period closed if usage. Fail → exception queue. Do not create. - Draft. Stripe create as
draft,auto_advance=false, Idempotency-Key =extract_id:draft. - Approve. Wait on webhook (
$execution.resumeUrl) or Slack send-and-wait. Show source + totals + draft URL. Auth on the resume URL. - Promote. Key
invoiceId:finalize, theninvoiceId:send. Two POSTs. - Write-back. CRM:
invoice_id,extract_id,approved_by,sent_at. - Errors. Error workflow with the Error Trigger. Page money paths; do not rely on a 2xx that created nothing.
| Node | Job | Not its job |
|---|---|---|
| Extract From File | Bytes → text / rows | Guessing what a blurry total “probably” is |
| Information Extractor | Text → JSON | Calling Stripe |
| Code / IF validate | Checksums and requireds | Rewriting lines to make the hash match |
| Stripe | Ledger object | Being your approval log |
| Wait | Pause + resume | A two-minute timer you hope a human saw |
| Error Trigger | Automatic failures | Editor clicks — n8n will not fire those as production alerts |
| Credential | Used by | Must not reach |
|---|---|---|
| Model / extractor key | Information Extractor only | Stripe live, QBO/Xero pay, mailbox send |
| Stripe test | Staging graph | Production customers |
| Stripe live | Send hops after approve | Extractor, prompts, execution logs |
| Resume-URL auth | Wait webhook | A public Slack paste |
If the extractor credential can create an invoice, you do not have three hops. You have one hop with extra prose.
Wait specifics that bite billing graphs:
- The node offloads execution to the database until resume. That is the point. A 60-second “wait for the founder” is not a gate.
$execution.resumeUrlis unique per execution. Bind it toextract_id. A forwarded naked URL is a credential.- Turn on Limit Wait Time. Past the limit: hold or reject. Never auto-send because nobody clicked. The overnight cousin of this failure is in when automation fails overnight: silent success, then a surprise in the customer inbox.
- Put Header / JWT auth on the resume webhook. Wait documents Basic, Header, JWT, or None. None is for demos.
Staging proof — do this before live keys:
- Drop one real (redacted) PDF into the trigger.
- Confirm the candidate JSON. Confirm
unknown[]is honest. - Confirm the customer inbox is empty after draft create.
- Approve. Confirm one invoice id, one send, one CRM write.
- Replay the trigger. Confirm you still have one invoice.
- Reject a second sample. Confirm no send, reason stored.
- Force a validate fail (break the total). Confirm no Stripe create.
If step 5 creates a second object, stop. You do not have a loop. You have a retry tax.
How do I write an extraction schema that fails closed?
Write the schema as a contract the model cannot charm its way around. Happy-path JSON examples are a trap: n8n’s Information Extractor treats every field as mandatory when you generate the schema from a JSON example. Mandatory-plus-missing is how you get a guessed email.
Prefer JSON Schema (or attribute descriptions with explicit optional fields) and make “I do not know” representable.
| Field | Type | Required? | If missing |
|---|---|---|---|
source_id | string | Yes — from the trigger, not the model | Do not extract. You already have it. |
customer_legal_name | string | Yes | Exception. Do not bill a nickname. |
billing_email | string | Yes | Exception. Do not invent accounts@ |
currency | enum (USD, …) | Yes | Exception. Do not assume home currency |
po_number | string or null | No | null + unknown reason |
line_items[] | array of {description, qty, unit_amount} | Yes, min 1 | Exception |
header_total | integer minor units | Yes | Exception |
period_start / period_end | date or null | Usage only | Exception if usage and either is null |
tax_id | string or null | Policy | null unless you collect it |
unknown[] | array of {field, reason, quote} | Yes (may be empty) | If the model is unsure, it goes here |
days_until_due | — | Forbidden | Policy table, not extract |
collection_rate / DSO / ”% likely to pay” | — | Forbidden | Not a field. Delete it from the prompt |
Validation that is not the model:
-
sum(qty * unit_amount) === header_totalin integer minor units. No floats. - Every
unit_amountis an integer. “Twelve hundred” in prose that became12000is a checksum fail if the PDF also printed1,200.00. - Currency is in the enum copied from the billing rail, not a free-text “dollars.”
- JSON Schema
additionalProperties: falseon the object you will persist. Extra keys are how a “rush fee” sneaks in. -
source_idon the candidate equals the claimed trigger id. The model does not get to rename the email. - Quotes in
unknown[]must be substrings of the source text. If the quote is not in the PDF, the extract is fiction.
System prompt rules that belong in the extractor template (n8n appends format instructions after your template — do not fight that; add constraints):
- Copy amounts from the document. Do not convert, round, or “fix” formatting.
- If a required field is not in the source, return
nulland explain inunknown[]. Do not substitute. - Do not add line items that are not printed or listed in the source.
- Do not output tax, due-date policy, or collection strategy unless those strings appear as facts in the source — and even then, due-date policy is stripped by validate.
- Never output card numbers. If a PAN appears, redact and exception-queue. Stripe’s security guide is blunt: card data on your systems explodes PCI scope. Invoice automation carries invoice ids and amounts. Not PANs. Not CVCs.
A guessed required field is a send waiting to happen.
What must the approval card show before anyone clicks send?
If the approver needs five tabs, you designed research. The card is the product. Totals are computed in the workflow. The model does not get to narrate the money in Slack prose.
| Card field | Why it exists |
|---|---|
| Customer legal name + CRM link | Stops billing the signer instead of the entity |
Currency + header total + sum(lines) | Two numbers. If they differ, the card should not have an Approve |
| Line table (description, qty, unit amount) | Stops a $0 draft that grew a mystery line |
unknown[] quotes | Makes the guess visible instead of burying it in JSON |
| Source link (email / PDF) | Approver inspects the document, not the clerk’s summary |
| Draft invoice URL | Approver inspects the ledger object |
| Extract id + schema version | Support can replay the same candidate |
Mode (first invoice / standard SKU / over $X) | Policy visible in the click |
| Approve / Reject / Request edit | Rejects need a reason code |
Checklist — people will actually click this:
- One screen, two primary actions, a clock (SLA + named backup)
- Approve disabled in the UI if checksum failed — do not rely on “they will notice”
- Token / auth on
$execution.resumeUrlso a forwarded link cannot finalize a stranger’s invoice - Double-click cannot double-send — claim
invoiceId:finalizebefore the POST - PTO backup named in the runbook
- Finance channel, not
#general. Volume is how routing bots get muted; billing cards are the same animal
SLA rule, because stale money is still money:
- Publish the clock (same day, next business morning — write it).
- Past SLA: ping the backup, not the whole company.
- Past limit: hold or reject. Wait’s Limit Wait Time is for this. Auto-approve-on-timeout is how a vacation becomes a customer incident.
- Overnight extract failures can wait for morning. Overnight send failures after an approve are a page — see automation fails overnight. Auth drift and silent 2xx belong in that triage, not in a weekly retro.
The click is the control. Homework is not a control.
How do idempotency keys stop a second invoice?
Retries are normal. Email triggers replay. Humans double-click. Stripe webhooks are at-least-once and live destinations retry for up to three days. If your only defense is “we probably will not hit that,” you will bill twice.
You need two layers: a business key you store, and Stripe’s idempotency keys on every POST.
| Layer | Key | Store | Replay behavior |
|---|---|---|---|
| Graph claim | source_id + schema_version | Your DB / Airtable / Data table, unique | Second trigger exits before extract |
| Draft create | extract_id:draft | Stripe header, ≤255 chars | Same payload returns the first invoice |
| Finalize | invoiceId:finalize | Stripe header | Does not create a sibling invoice |
| Send | invoiceId:send | Stripe header | Does not mail twice as a new object |
| Side effects | invoiceId:crm / invoiceId:slack | Your store | One CRM write, one notify |
| Payment (later) | Stripe event.id | Unique insert before CRM “paid” | Covered in the ops pipeline spoke — do not skip it |
Stripe’s rules, as of their current docs: keys up to 255 characters; they suggest V4 UUIDs or enough entropy; do not put emails or personal identifiers in the key; they may prune keys after at least 24 hours; all POSTs accept keys; GET/DELETE ignore them. Use stable business keys (extract_id:draft), not a random UUID generated inside the retry.
Procedure — keys in the workflow:
- Compute
extract_id=source_id+schema_version(and attachment hash if the email can carry two PDFs). - Claim it uniquely before the extractor spends a token, or at latest before Stripe create. Extract is cheap. Duplicate drafts are not.
- Send
Idempotency-Keyon create, finalize, and send. - On Stripe 409 / mismatch (same key, different body), go to the exception queue. Do not “make a new key and try again.” That is how you create the second object on purpose.
- After 24 hours, Stripe may have forgotten the key. Your stored invoice id is the long-term lock. Resume send with
invoiceId:send, never a second create.
| Anti-pattern | What happens | Fix |
|---|---|---|
| UUID generated in the Function node on every run | Every retry is a new key → new invoice | Derive from extract_id |
| Email address in the key | PII in a header, and aliases fork the key | Use ids |
| One key for create+finalize+send | Parameter mismatch on the second hop, or a create replay that includes send fields | One key per POST |
| Claim after send | You already mailed | Claim before create |
Replay whole graph after invoice.paid | Kickoff tasks multiply | Fan-out behind paymentId:… keys |
Keys are cheaper than refunds. If step “replay trigger” in staging still creates invoice two, you are not ready for live.
Which admin tasks reuse the same loop without becoming a second Slack process?
Admin is not a junk drawer. Each task gets a schema and the same three hops. The destination changes. The discipline does not.
| Admin task | Extract from | Approve shows | Send / write | v1 send? |
|---|---|---|---|---|
| Expense report | Receipt image / PDF | Merchant, amount, date, project | Draft expense in books | No auto-reimburse |
| Vendor bill | Vendor PDF | Vendor, amount, due date, invoice # | Draft bill | No auto-pay |
| Time → invoice lines | Notes / calendar / time tool | Hours, project, approved stamp | Draft invoice lines | Only if time was already approved |
| Ticket → credit candidate | Help desk thread | Original invoice id, amount, reason | Credit-note draft | Finance role only |
| W-9 / vendor master | EIN/name match, address | Staging vendor row | Human commits master data | |
| Renewal reminder | CRM + invoice status | Customer, amount, date | Email draft | Human send on first cohorts |
| Close-week pack | Graph stats | Counts, aging, failed runs | Internal digest | Auto OK — no customer |
Shared rules:
- New schema, new
extract_idnamespace (expense:,bill:,w9:). Do not reuse invoice keys. - Same fail-closed validate. OCR that is unsure on the total does not create a bill. That is the AP cousin of a guessed invoice line — the ops pipeline spoke covers pay-side dual control; this page covers not letting the model skip it.
- Same mute test as routing: if
#financegets every receipt, they will ignore the invoice that matters. - No “AI collections.” A reminder is customer contact. It is hop three. It waits.
Decision list — when an admin task is not a cousin:
- If success is taste (brand copy, legal strategy), keep it manual.
- If the write hits a bank account, you need dual control on pay, not a smarter extract.
- If nobody will own the exception queue, you built a pile.
- If the only metric anyone wants is “faster collections,” stop. That number is not in this design.
Retype the form. Do not retype a culture of chasing.
What breaks this in production?
Concrete failure mode:
- A proposal PDF prints $1,200.00. OCR / the model emits
12000(lost decimal, or a comma treated as grouping in the wrong place). - The schema stored
amountas a string. There is no integer checksum againstsum(lines)and nounknown[]quote of the printed total. - Confidence looks “high.” The graph treats complete JSON as approved, or a timeout auto-sends because Limit Wait Time was wired to continue.
- Stripe create runs with
auto_advanceleft at the default. The customer is notified without a human. Or you do have a draft, then a retried email trigger creates a second draft becausesource_idwas never claimed. - Finance spends the week on a credit note and a founder apology. The workflow is turned off. Close goes back to copy-paste. The next three finance automations never get approved.
We do not attach a dollar figure or a collection-rate “hit” to a named client. You can price your cleanup: two invoices or one wrong total, one credit path, one week of manual billing, and a freeze on every money graph. That is the fail cost. The handbook exists so you do not learn it on a live customer.
| Skip | What breaks | What you do instead |
|---|---|---|
| Model can call send | Wrong entity / wrong total leaves | Extract credentials ≠ Stripe live key |
| JSON-example schema, all fields mandatory | Guessed email, guessed PO | Optional + null + unknown[] |
| No checksum | 12000 vs $1,200.00 | Integer minor units, sum === header |
auto_advance=true | Stripe emails and retries for you | false; you own finalize/send |
| No source claim | Duplicate drafts on retry | Unique source_id before create |
| Random UUID keys | New key → new invoice | extract_id:draft |
| Timeout = approve | Vacation send | Limit Wait Time → hold / reject |
| Slack as the ledger | Month-end is archaeology | Invoice id + extract id in CRM |
| Personal Stripe login | Graph dies when they leave | Service seat |
| Card image in the prompt “for context” | PCI scope | Redact; exception queue |
| Metric = “we’ll collect faster” | You will invent a rate | See the next section |
Anti-patterns that only show up on this AI loop (not the generic pipeline):
- Prompt: “If the total looks low, add a rush line.”
- Prompt: “Estimate likelihood of payment.” Delete that sentence.
- Using chat history from a sales thread as line items without a signed source.
- Letting the model set
days_until_duefrom “Net 30” folklore in the email signature. - Retrying extract and create because the Wait node was cancelled, without the claim row.
Bravery is not a restore strategy. A checksum and a gate are.
How do I measure whether the loop is working without inventing collection rates?
Measure whether extract is honest, whether humans still catch the misses, and whether retries stay at one invoice. Do not measure “AI collections,” DSO shaved, ”% recovered,” or “invoices paid faster.” Those numbers depend on customers, terms, product, and dunning policy. This graph does not own them. If someone wants a collection rate in the dashboard, they can run a finance report from the ledger after a dated window. They may not attribute it to the extractor.
| Metric | How you get it | Target shape | Not a target |
|---|---|---|---|
| Schema-pass rate | Validate node pass / all extracts | Rising as intake gets cleaner | 100% by letting the model fill gaps |
| Reject reason mix | Reason codes on the card | “Wrong PO” → add a form field; “invented line” → tighten schema | A single “thumbs down” |
| Checksum fail count | sum !== header | Falling, then rare | Zero because you stopped checking |
| Time-to-approve | decidedAt - draftedAt vs SLA | Inside the published clock | Heroics at 11pm |
| Duplicate invoice count | Count of Stripe invoices per extract_id | Zero | “Only a few” |
| Exception age | Queue timestamp | Hours, not weeks | A graveyard you screenshot |
| Retype hours on this path | Same clock as baseline week | Down on draft assembly | A company-wide productivity % |
| Auto-send class size | Export of invoices with mode=auto | Empty in v1; later, one written class | All invoices because extract “is good now” |
| Number leadership will ask for | What you say |
|---|---|
| “What’s our collection rate with AI?” | We do not have one. We have a send-control rate. |
| “How much faster do we get paid?” | Not this project’s KPI. Terms and customers drive that. |
| “What’s the model accuracy?” | Reject reasons + checksum fails, dated. Not a single % from a vendor slide. |
| “ROI?” | Hours of retyping removed on this path, minus incident cost. No AR multiplier. |
Observation window — dated, exportable:
- Every extract:
extract_id, pass/fail,unknown[]size - Every approve: actor, reason, mode
- Every send: invoice id, idempotency key
- Duplicate query:
COUNT(invoice_id) GROUP BY extract_idhaving count > 1 - No field named
predicted_pay_dateorcollection_score
If you cannot export that list, you cannot promote a class to auto-send. A clean demo week is not a window.
Promotion still lives on the invoice pipeline rules: known customer, standard SKU, under a written $X, after this window is boring. The AI loop does not earn a shortcut.
When should I hire vs DIY this automation?
DIY the extract-to-draft loop when the write is reversible and the blast radius is one person with a pause switch. Hire (or book the $500 Automation Audit) when a retry can bill a customer twice, tax/entity is in play, or nobody named will own the queue at 2am.
| Situation | DIY | Audit / build |
|---|---|---|
| One currency, one legal entity, standard SKU | Yes — extract → draft → you click | — |
| You will watch every card for a month | Yes | — |
| Staging proof of “replay trigger = still one invoice” already green | Yes to go live on drafts | — |
| Live Stripe send on the first week | No | Yes |
| Multi-entity, tax IDs, mixed FX | No | Yes |
| Vendor pay or customer refunds in the same graph | No | Yes — different gates |
| Personal Stripe / QBO login today | Fix seats first | Yes if you want it production |
| Overnight send after approve, no on-call | No | Yes — page money paths |
| Leadership wants a collections KPI on the extractor | Stop and rewrite the scoreboard | Yes, to keep that number off the graph |
| Volume is a handful of invoices a month and intake is already clean | Maybe not worth a canvas | Only if the cost is learning the spine |
Decision list:
- If you cannot pause the workflow in two minutes, you are not in DIY-live. You are in demo.
- If finance will not name an approver and a backup, you are not ready to send. Extract-to-draft can still save retyping.
- If the only “AI” request is dunning copy that cites a recovery percentage, decline the metric. Build extract-approve-send or build nothing.
- If you already have the ledger spine from the ops pipeline spoke and you only need hop one, DIY the extractor against that spine. Do not rebuild Stripe.
DIY is a smaller loop, not a sloppier one. Keys and checksums still ship.
What should I skip if I only have a week?
Skip autonomy theater. Ship one input class through a gate you can replay.
Do this week:
- Pick one mailbox or one folder. Write
source_id. - Write the schema table (required vs
null). Ban collection fields. - Information Extractor + checksum + exception queue.
- Draft create,
auto_advance=false,extract_id:draft. - Approval card with source + totals. Limit Wait Time → hold.
- Forced duplicate test. Forced checksum-fail test.
- Runbook: who pauses, who approves, where exceptions live.
Skip this week:
- Auto-send, even on “high confidence”
- AP auto-pay
- AI dunning / “smart collections”
- Three billing rails writing the same invoice
- Agent-with-Stripe-tools
- A dashboard tile named collection rate
- Every admin cousin on the same canvas
A week of one honest loop beats a month of a model that mails.
When is this not worth doing yet?
Skip the canvas when you cannot name the source of truth for lines, when nobody will click the card, or when volume is already smaller than the cost of a wrong send. AI does not fix a billing process that does not exist.
| Blocker | Why the loop fails | Do this first |
|---|---|---|
| Lines live in three spreadsheets | Extract will average the fiction | One assembly rule (proposal or usage or approved time) |
| No finance owner | Cards rot; timeout becomes policy | Name approver + backup |
| Personal billing login | Graph dies on PTO | Service seat |
| Tax/entity is tribal knowledge | Model will guess the letterhead | Write the entity table |
| Volume is tiny and already clean | You are buying incident risk for minutes | Stay manual; steal the schema for later |
| Success = “collect faster” | You will invent a rate | Pick retyping hours or walk away |
| Cannot pause in two minutes | Not production | Error workflow + named pause |
Go-live pause test — if any box is empty, keep send off:
- Named human can disable the workflow in two minutes without a deploy
- Extractor credential cannot see live Stripe
-
livemode/ tenant is visible on the approval card - Limit Wait Time holds; it does not approve
- Duplicate query on
extract_idis in the Friday pack
If the process changes every sprint, freeze a v1 class (domestic standard SKU) or wait. The extractor will happily encode last week’s exception as this week’s line item.
Worth doing the moment retyping is weekly, the PDF/email already contains the facts, and a human will still own send. That is the whole product.
FAQ
How do I automate invoicing and admin tasks with AI?
Put the model on extract, a named human on approve, and the billing API on send. Schema-validate the candidate (checksum, currency, email, source id), create a draft with auto_advance=false, then finalize and send as separate idempotent POSTs. Admin tasks — expenses, vendor bills, W-9 staging — reuse the same three hops with their own schemas. The model never owns Stripe send.
How do I measure whether automating invoicing and admin tasks with AI is working?
Track schema-pass rate, reject reason codes, checksum fails, time-to-approve versus SLA, duplicate invoices per extract_id (target zero), exception age, and retyping hours on this path. Do not invent a collection rate, DSO change, or ”% recovered.” Those belong to customers and terms, not to the extractor.
What usually fails first when teams try this?
A guessed total or invented line that still looks like valid JSON, then a send without a gate — or a retry that creates a second invoice because source_id was never claimed. Close second: JSON-example schemas that make every field mandatory, so the model fills an email that was never in the PDF. Confidence is not a control.
How long does this take to show results?
You should see retyping drop once extract-to-draft is live and humans are clicking a card that shows source and totals — often inside a focused stretch after credentials and one input class exist. I will not invent a days-to-DSO or studio-wide payback figure. Early proof is one invoice per extract and a reject log you can read.
What should I skip if I only have a week?
Skip auto-send, AP autopay, AI dunning, and a collections dashboard. Do one input class, a fail-closed schema, draft create with an idempotency key, an approval card with Limit Wait Time, and a forced duplicate test. A week of that loop beats a week of a model with Stripe tools.
When is this not worth doing yet?
When line items have no single source of truth, nobody will own the approval SLA, logins are personal, or the KPI you want is a collection rate the graph cannot honestly produce. Fix assembly rules and ownership first. Extract-approve-send is for facts that already exist in a document, not for inventing a billing process.
CTA
Extract the form. Approve the money. Then send once.
If you want extract → approve → send built to production standard, start with the handbook, then use automation or book the $500 Automation Audit.
What questions does this article answer?
- How do I automate invoicing and admin tasks with AI?
- Put the model on extract, a named human on approve, and the billing API on send. Schema-validate the candidate (checksum, currency, email, source id), create a draft with `auto_advance=false`, then finalize and send as separate idempotent POSTs. Admin tasks — expenses, vendor bills, W-9 staging — reuse the same three hops with their own schemas. The model never owns Stripe send.
- How do I measure whether automating invoicing and admin tasks with AI is working?
- Track schema-pass rate, reject reason codes, checksum fails, time-to-approve versus SLA, duplicate invoices per `extract_id` (target zero), exception age, and retyping hours on this path. Do not invent a collection rate, DSO change, or "% recovered." Those belong to customers and terms, not to the extractor.
- What usually fails first when teams try this?
- A guessed total or invented line that still looks like valid JSON, then a send without a gate — or a retry that creates a second invoice because `source_id` was never claimed. Close second: JSON-example schemas that make every field mandatory, so the model fills an email that was never in the PDF. Confidence is not a control.
- How long does this take to show results?
- You should see retyping drop once extract-to-draft is live and humans are clicking a card that shows source and totals — often inside a focused stretch after credentials and one input class exist. I will not invent a days-to-DSO or studio-wide payback figure. Early proof is one invoice per extract and a reject log you can read.
- What should I skip if I only have a week?
- Skip auto-send, AP autopay, AI dunning, and a collections dashboard. Do one input class, a fail-closed schema, draft create with an idempotency key, an approval card with Limit Wait Time, and a forced duplicate test. A week of that loop beats a week of a model with Stripe tools.
- When is this not worth doing yet?
- When line items have no single source of truth, nobody will own the approval SLA, logins are personal, or the KPI you want is a collection rate the graph cannot honestly produce. Fix assembly rules and ownership first. Extract-approve-send is for facts that already exist in a document, not for inventing a billing process.
Last reviewed
Automation
Automation After the show is not you at 1 a.m.
Post-show onboarding — thank-you, join path, merch nudge — belongs in a human-gated n8n rail, not your thumb at load-out.
Automation Paperwork that is not the plant
Invoice and PO matching, intake, and support triage in n8n with Metrc fences — the paperwork operators hate, not a menu widget.
Automation Saturday still books — the missed-call rail for trades
A missed-call text-back that routes zip and books a slot beats voicemail and Saturday desk coverage you cannot keep staffed. If a kid is cheaper, say so.
Automation Why doesn’t worker concurrency cap my n8n sub-workflows
Worker concurrency does not cap n8n sub-workflows. Each Execute Workflow child is a new execution the production limit skips, usually on the parent worker.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.