Paper-craft illustration: open navy ledger, cream calculator and coin stack, a cream paper hand with an orange wax seal.

Guide

AI for Finance Teams: What’s Safe, What Needs a Human, What’s Off-Limits

Enthusiasm, then a pause: which of these numbers is a business actually willing to let a model touch? Three zones, written as a policy.

Hand AI the high-volume, checkable work — invoice extraction, reconciliation prep, expense coding — and keep a qualified human on sign-off, audit judgement, tax positions, and anything that goes in front of a board. Co-pilot the forecast and the variance narrative. Never paste real financials into a free or personal consumer chat. Enterprise tier, source-system check, a materiality line agreed in advance, and an audit trail: those four keep the safe zone actually safe.

Every finance demo ends the same way. The model extracts the invoice, matches the bank line, drafts the variance note, and someone in the room says, quietly, “we could never let it touch that.” The pause is the useful part. Enthusiasm is cheap. Which of these numbers is the business actually willing to let a model touch?

This page is the answer, written as a policy: three zones — safe, co-pilot, no-go — plus the four guardrails that keep them from collapsing into each other. It is the finance version of the sorting rule in our Create, Understand, Act framework, and it sits next to the AI guardrail playbook. The tasks change. The line does not.

$12.88–$19.83 fully loaded cost to process an invoice by hand — the number AI is measured against
39% vs <0.1% manual invoice error rate versus AI extraction, on the same industry benchmarks
~70% cycle-time cut the same programmes cite, with about three-quarters off the cost
98% reconciliation prep vendors claim they can auto-match — treat as a ceiling, then measure

Three zones, written as a policy

Do not launch “AI for finance.” Write which tray each task sits in, who owns it, and what happens when the model is wrong. Most teams already know the split. They have just never put it on one page.

Zone What it covers Who owns the output
Safe AP/AR invoice extraction, reconciliation prep, expense categorisation The model runs. A person samples. Exceptions queue.
Co-pilot Forecasting support, variance narratives, board-pack drafts, month-end close support The model drafts. A named human reads, edits, and signs.
No-go Statement sign-off, controls and audit judgement, tax positions, consumer AI tools A person. The model is not in the room.
Paper-craft illustration: three navy trays on cream handmade paper, joined by a dashed orange line — a cream tick in the first, a cream pencil with an orange stripe in the second, a cream padlock in the third.
Safe, co-pilot, no-go — three trays, not a slider. Put the task in one of them before anyone opens a tool.

Safe: extraction, recs, categorisation

Safe means high volume, low judgement, grounded in a source system you already run. The model reads. It does not decide what the number means.

AP/AR invoice extraction

This is the cleanest win in the function, and the one with the oldest receipts. Manual processing still costs $12.88 to $19.83 per invoice, fully loaded, and about 39% of those invoices carry at least one error. AI-automated extraction is cited below 0.1% error, with cycle times cut by around 70% and cost cut by about three-quarters — typically down to the $2.36–$2.78 band.[1]

Those are industry benchmarks, not a promise on your supplier mix. A clean PDF from a recurring vendor is not the same job as a photographed receipt in two languages. Run the model on a month of your own invoices, count the exceptions, and only then talk about headcount.

Reconciliation prep

Matching bank lines to the ledger is the other high-volume tray. Vendors claim they can auto-match up to 98% of the prep work and leave a person with the exception report.[2] Treat 98% as a ceiling someone else measured, then measure yours. The safe job is the match and the flag. The posting, and anything material, stays with a person.

Expense categorisation

Coding a card feed against a chart of accounts you actually maintain is the same shape as invoice extraction: repetitive, checkable, reversible. Policy exceptions — a meal over the cap, a vendor you have never seen, a category that would move a filing — go to the queue. If the model cannot cite the policy line, it does not code it.

Co-pilot: forecasts, narratives, the close

Co-pilot means the model is in the room and a named human still owns the number. Four jobs sit here, and none of them ship unread.

  • Forecasting support. Vendors quote 92–97% accuracy. Treat that as a vendor claim until you have scored it on your own pack, over a hold-out that includes a messy month, not a quiet Tuesday.[3] AI can draft the driver tree and the first-pass range. A person still picks the case that goes to the board.
  • Variance-analysis narratives. “Payroll is $180k over because headcount landed three weeks early in APAC” is a useful first draft. It is also the sentence that will be wrong in a confident voice. Every name, number and cause gets checked against the source system before it leaves the tab.
  • Board-pack drafting. Structure, charts, the first pass of the commentary. Not the pack that is sent. In 2025 Deloitte had to refund part of an AU$440,000 Australian government report after a researcher found fabricated citations in a document built with AI assistance.[4] That is a co-pilot failure, not a no-go one: the model drafted, and nobody ran the citation check.
  • Month-end close support. Checklists, flux drafts, the “what still isn’t reconciled” list. The close itself is still a person’s calendar.

In February 2026 COSO published Achieving Effective Internal Control Over Generative AI. The line that matters for a close: set-and-forget assurance is inadequate for probabilistic models.[5] A rule-based bot that always posts the same way is not the same control as a model that drafts a different variance note on Tuesday than it did on Monday. If you cannot say who reviewed this week’s output, you do not have a control. You have a habit.

A forecast the model drafted is a draft. The number that goes to the board has a name next to it.

No-go: sign-off, judgement, tax, consumer tools

Four things the model does not do. No copilot draft goes out unread; no agent resolves them end-to-end.

  • Final sign-off on the financial statements. The signature is a person. PCAOB amendments to AS 2201 — an audit of internal control over financial reporting, integrated with the financial-statement audit — apply for fiscal years beginning on or after 15 December 2026.[6] Whatever tooling you add this year has to survive that audit, not just a demo.
  • Controls and audit judgement. Whether a control is designed well, whether a deficiency is material, whether a sample is enough — those are judgements. The model can fetch the evidence. It cannot own the conclusion.
  • Tax positions. IRS Circular 230 applies in full to AI-assisted tax work. The Office of Professional Responsibility’s June 2026 guidance did not write a new rule; it said the old ones already cover the tool. Due diligence, competence, confidentiality, reasonable fees — none of them get lighter because a model drafted the memo.[7] A practitioner who cannot verify the citation does not file it.
  • Free and personal consumer AI. Never paste confidentials — invoices, a P&L, a tax pack, an employee file — into a free or personal ChatGPT, Claude or Gemini account. Enterprise and business tiers only: ChatGPT Enterprise, Claude for Financial Services, or the equivalent contracted stack your information-security lead has actually signed.[8][9] A personal login on a company close is a data incident with a friendly UI.
Paper-craft illustration: a navy padlock resting on a cream ledger of figures, with a burnt-orange wax seal stamped ENTERPRISE beside it.
Confidentials stay on the contracted tier. The seal is the point — not the model.

Guardrails that make the zones hold

The zones fail the same way every other AI policy fails: they look sharp in the launch meeting and nobody can find them in month three. Four rules, specific enough to follow:

  1. 01

    Enterprise tier only

    Company work on a contracted account — ChatGPT Enterprise, Claude for Financial Services, or the equivalent. No personal logins. No free tiers. No “just this once.”

  2. 02

    Verify against source systems

    Every number, name, date and citation in a co-pilot draft is checked against the ERP, the bank, the tax file — not against the model’s confidence.

  3. 03

    Materiality, in advance

    Write the threshold before the first close. Below it, the safe tray may post with a sample. At or above it, a named human signs. Do not invent the number on day 28.

  4. 04

    An audit trail

    Who ran it, on what, who checked it, what changed. If you cannot reconstruct last month’s flux narrative, you do not have a control COSO would recognise.

That is the same discipline as the one-page guardrail template — stance, approved tools, data rules, review points — pointed at a ledger instead of a blog post. Print it. Pin it. Put it on page one of onboarding. A zone nobody has read is not a zone.

Write it, train it, check it

A policy that lives in a shared drive is not a policy. Treat the rollout the way you would treat any other change that has to survive month-end: a named owner, training on the actual page, and a look-back at 30, 60 and 90 days to see whether the trays are still being used or quietly skipped.

That is the same adoption curve we have run at scale — more than 5,000 people trained across the Philippines and Australia, 98% satisfaction on the programmes we measure, and in one 4,000-person rollout, 72% weekly use inside a month and 46% daily at six months — applied to a finance team instead of a whole company. The tooling is the easy part. Whether the controller trusts the split is the work. If you want that mapped onto the pack you already export, that is an AI consulting conversation, not a chatbot subscription. The training half lives in corporate AI training.

Sources

  1. Manual invoice processing commonly $12.88–$19.83 fully loaded, with around 39% of invoices carrying at least one error (IOFM / Ardent Partners); AI-automated extraction cited below 0.1% error, costs typically $2.36–$2.78, cycle times from the mid-teens in days down to around 3.1 for best-in-class — the ~70% cycle cut and ~three-quarters cost cut used above. Factura.ai, “20 Statistics Proving How AI Technology Transforms Invoice Processing Accuracy in 2026,” May 13, 2026. factura.ai (benchmarks compiled from Ardent Partners 2025 and related AP studies; see also Parseur, “AI Invoice Processing Benchmarks 2026.” parseur.com)
  2. Reconciliation prep “up to 98%” auto-matched is a vendor-reported ceiling, not an independent benchmark — HighRadius cites 98% auto-tagging on bank-statement processing; ChatFin cites customers eliminating 98% of manual rec work. Measure on your own feed. transformance.ai (HighRadius) and chatfin.ai
  3. Forecast-accuracy figures in the 92–97% band are vendor claims (HighRadius 95%; Anaplan-class tools commonly quoted 92–96%). Treat as marketing until scored on your own pack. Transformance, “Cash Forecasting Tools Compared: 8 Options for 2026,” May 6, 2026. transformance.ai
  4. “Deloitte to refund government after using AI in $440k report,” Accounting Times, Oct. 9, 2025 — Department of Employment and Workplace Relations report; fabricated academic references and misquoted Federal Court orders; partial refund. accountingtimes.com.au (also CFO Dive)
  5. Committee of Sponsoring Organizations of the Treadway Commission (COSO), Achieving Effective Internal Control Over Generative AI, February 2026. coso.org — Deloitte on WSJ: generative AI moves assurance from deterministic, rule-based systems to probabilistic models, and “set-and-forget” monitoring does not work. deloitte.wsj.com
  6. PCAOB AS 2201, An Audit of Internal Control Over Financial Reporting That Is Integrated with An Audit of Financial Statements — amendments to paragraph .09 and new paragraph .99, effective 15 December 2026 (fiscal years beginning on or after that date). pcaobus.org
  7. IRS Office of Professional Responsibility, OPR Alert 2026-19, “Introductory Guidelines for Responsible AI Use in Federal Tax Practice,” June 24, 2026 — Circular 230 duties (due diligence, competence, confidentiality, fees) apply in full to AI-assisted tax work. Journal of Accountancy, June 26, 2026. journalofaccountancy.com (also Forbes, June 25, 2026. forbes.com)
  8. OpenAI, ChatGPT Enterprise — contracted, no-training-on-business-data tier for company work. openai.com/enterprise
  9. Anthropic, Claude for Financial Services — enterprise/business tier with source-attributed outputs and FSI compliance posture (SOC 2, FedRAMP). claude.com/solutions/financial-services and support.claude.com

Every citation above was checked against its source before this piece was published — the same check the article asks a controller to run on a flux narrative.

FAQ

Common questions

Which finance tasks are actually safe to hand to AI?

High-volume, low-judgement work grounded in a source system you already run: AP/AR invoice extraction, reconciliation prep, and expense categorisation. The model reads and codes. A person still samples, and still posts anything material.

Can AI sign off financial statements or take a tax position?

No. Final sign-off on the statements, controls and audit judgement, and any tax position are no-go. IRS Circular 230 applies in full to AI-assisted tax work — due diligence, competence and confidentiality do not get lighter because a model drafted the memo.

Is it all right to paste invoices or a P&L into free ChatGPT?

No. Never paste confidentials into a free or personal consumer tool. Use an enterprise or business tier with a contract — ChatGPT Enterprise, Claude for Financial Services — and still verify every number against the source system. See the AI guardrail playbook for the staff-facing data rule.

Should we believe vendor claims of 92–97% forecast accuracy?

No. Treat it as a vendor claim until you have measured it on your own pack, over a hold-out that includes a messy month. AI can draft the forecast and the variance narrative. A named human still owns the number that goes to the board.

How do we turn this into a policy the team will actually follow?

One page, three zones, four guardrails, a named owner. Pin it, train it, check it at 30/60/90. That is the same adoption discipline as corporate AI training — and the conversation we run in AI consulting.

Done-for-you

Book a readiness conversation

We map the three zones onto the pack you already run — what the model may extract, what a named human still signs, and what never leaves the ledger. Then we train the team on the page, not a deck.