Practice · AI

Where AI belongs in a finance function — and where it does not

The useful question is not whether AI can do accounting. It is which specific tasks it does reliably enough that a Chartered Accountant would sign the output — and that line is much sharper than the marketing suggests.

Finance is unusually well suited to automation and unusually unforgiving of it. Well suited, because most of the work is high-volume, rule-governed and heavily patterned. Unforgiving, because the output is consumed by auditors, lenders and revenue authorities, none of whom accept "the system produced it" as an explanation, and because a single confident error costs more trust than a year of correct months earns back.

So the question worth answering is narrow and practical: for each task, is the model reliable enough that a named Chartered Accountant would put their signature on the result?

The test that decides it

Four properties determine whether a finance task can be safely automated. A task that has all four is a strong candidate. A task missing any one of them needs a human in the loop, and a task missing two should not be automated at all.

PropertyQuestionIf absent
VerifiableCan the output be checked against something objective — a bank statement, a contract, a return?Errors are invisible until they compound. Do not automate.
BoundedIs the space of correct answers small and enumerable?The model will be plausibly wrong rather than obviously wrong.
ReversibleIf it is wrong, can it be corrected before anyone external relies on it?Requires human approval before the action, always.
PatternedAre there hundreds of prior examples with known-correct answers?No basis for a confidence threshold. Treat as judgement.
The distinction that matters

The risk is not that a model gets things wrong. Junior accountants get things wrong too. The risk is that a model gets things wrong confidently and consistently, at volume, in a way that looks exactly like being right. A junior who is unsure asks. A model that is unsure produces a well-formatted answer.

Everything in a serious control design exists to reintroduce the asking.

What works reliably today

These pass all four tests. In our own engagements they run with high autonomy and a human reviewing exceptions rather than every item.

  • Transaction coding. Against an established chart with a year of history, accuracy on recurring patterns exceeds what a rotating junior achieves — because the model does not get bored on transaction four hundred. Novel suppliers still queue for review.
  • Bank reconciliation. Rule and pattern matching, with genuine exceptions surfaced. This is close to a solved problem.
  • Document extraction. Pulling supplier, date, amount, GST and line detail from invoices and receipts. Reliable, and trivially verifiable against the document.
  • Obligation tracking. Knowing what is due, for which entity, in which jurisdiction, and when. Calendar logic dressed up as intelligence, and none the worse for it.
  • Reconciliation preparation. Assembling supporting schedules and flagging the items that do not tie.
  • Anomaly detection. Duplicate payments, unusual amounts against a supplier's history, out-of-hours activity, round-number invoices. Machines are simply better than people at this.
  • First-draft commentary. "Revenue is $84k below budget, driven by two delayed enterprise contracts." Factual variance description, generated from the ledger, then edited by a person who knows why.

What works under supervision

These are useful, and materially faster than doing it from scratch, but the output is a draft and must be reviewed by someone qualified before it goes anywhere.

  • Return preparation. A GST or BAS return prepared from the ledger and reconciled back to it. The arithmetic is reliable; the treatment of edge cases is not. A CA reviews and files.
  • Month-end journals. Accruals, prepayments, depreciation, FX revaluation — all mechanical given correct inputs. The judgement is whether the inputs are complete, and that is a human question.
  • Board and investor packs. Assembly, formatting, charting and factual commentary all draft well. What the board should be worried about does not, and that is the part that matters.
  • Forecast mechanics. Rolling the model, loading actuals, recalculating, running scenarios. The mechanics are reliable. The assumptions are a judgement about the future, and a model has no basis for those beyond extrapolation.
  • Diligence responses. Drafting answers from source documents is fast and largely accurate. What to disclose, how to frame it, and when to say nothing is entirely human.
  • Policy application. Applying a written accounting policy to a transaction works well. Deciding whether the policy is right, or whether this transaction is the exception, does not.

What does not, and will not soon

Not because the technology is inadequate, but because the task is not the kind of thing that automation resolves.

Judgement under genuine ambiguity

Is this contract one performance obligation or three? Is this receivable impaired? Is the tax position arguable enough to take? These have no verifiable answer at the time you must decide. They require a professional to form a view and be accountable for it — and accountability is not a property a model can hold.

Anything where being wrong is not reversible

Payment initiation. Filing a return. Signing accounts. Committing to a lender. The control principle is simple: a machine may prepare, a human must authorise. Not because the machine is more likely to err, but because when it does, someone must have been in a position to stop it.

Advising a person

Telling a founder that the raise is not going to close at the price they want. Telling a board that the CEO's plan does not survive its own arithmetic. This is what a CFO is for, and it depends on relationship, timing and knowing how the person will hear it.

Deciding what matters

A model can surface every variance above a threshold. It cannot tell you that the small one in a minor cost line is the early signal of something structural, because that requires knowing what the company is trying to do. Materiality is a judgement about consequence, not size.

A working heuristic

If the answer can be checked against a document, automate it and review the exceptions. If the answer depends on a view about the future, an interpretation of a standard, or how a person will react — a professional decides, and the machine's role is to prepare the evidence.

The controls that make it safe

An AI-assisted finance function is not a normal one with a feature bolted on. It needs a control design of its own, and the design is the product.

  • Confidence floors. Agents act only above an explicit threshold; below it the item queues for a person. A short review queue is a healthy sign, not a failure.
  • Narrow agents, not a general assistant. A model asked to do one thing against written policy is testable and defensible. One asked to "do the accounts" is neither.
  • Full traceability. Every figure links to its transactions and every transaction to its source document. If you cannot click a variance and land on the invoice, you cannot defend the number.
  • Immutable audit log. Every action, override and approval, timestamped and attributed. An auditor must be able to reconstruct how any figure came to exist.
  • Regression testing on procedures. When a rule changes, run it against a held-out set of cases with known answers before releasing it. Finance teams have never done this. Software teams always have. Adopt the software habit.
  • Earned autonomy. New clients start with everything in draft-and-review, and autonomy is granted agent by agent as accuracy is demonstrated on their own data. Never on a vendor's benchmark.
  • Data discipline. Segregated by client, encrypted, never used to train models, with zero-retention terms from model providers where available.

The part nobody says out loud

This changes finance jobs, and it removes some of them. Roles built mainly on processing volume — coding, matching, chasing, keying — shrink substantially. Pretending otherwise is not kindness.

What grows is everything the processing work was crowding out: analysis, forecasting, business partnering, controls, and the judgement calls that were being made hurriedly because nobody had time. A finance team of six spending most of its week on transactions can become a team of four doing work that changes decisions — and that is a better job for the four, and an honest conversation to have with all six.

The right sequence is to be straight about it early, help people move to the work that remains where the capability is there, and treat the transition as a real project rather than a consequence discovered at go-live.

Agents do the work. A named Chartered Accountant owns the opinion. Any arrangement where nobody in particular is accountable for the output is not an operating model — it is an unallocated risk.

XLCFOWe run eight supervised finance agents behind a Chartered Accountant review gate. Nothing reaches a board, a lender or a regulator without a named human signature.

Book a platform walkthrough How the agents work