AI bookkeeping is sold as work removed from the week. The effect on an accountant's time points somewhere else. In a study by academics at Stanford and MIT, reported by the Journal of Accountancy, the average accountant using generative AI on foundational accounting work reallocated about 8.5% of their time out of routine data entry and toward high value tasks such as business communication and quality assurance.
Quality assurance is review. That is the trade a firm owner is actually buying, and it decides which steps of the bookkeeping cycle a model can take and which one it hands straight back.
What AI Bookkeeping Means Inside the Ledger
AI bookkeeping is software that reads a client's bank and card lines, and the documents behind them, and proposes or posts the account each line belongs to. Everything else the category sells is built on that one step.
It is worth separating from the automation your firm already runs. A bank rule is a condition somebody wrote by hand: if the description contains this string, code it to that account. A model is not given the rule. It infers a coding from patterns in the client's history and in other ledgers, and it returns a confidence level along with the answer. The profession's own account of these tools says they "continuously improve with user feedback", "provide clear indications on confidence levels", and "offer predefined automation workflows" (Journal of Accountancy, January 2026).
The confidence level is the setting that decides how the software behaves inside your firm. Somebody picks the point at which a coding posts on its own and the point at which it stops for a person. Ask your firm who picked it, and when it was last looked at.
Where AI Bookkeeping Lands in the Cycle, Step by Step
Take the cycle as your firm already runs it: feed intake, coding, document capture, reconciliation, close. A model does not take those steps evenly. It takes one of them almost entirely, helps with two, and hands the last one back.
Bank and card feed intake. A feed is a connection, and a model reads whatever arrives through it. It does not know that nothing arrived. Confirming that each feed delivered, and that a quiet week is a quiet client rather than a broken connection, stays a human item in the work between closes.
Transaction coding. Coding is where a line gets assigned to an account in the chart of accounts, and it is the highest volume decision in the cycle. This is the step the category is really about, and it is the one a model takes.
Document and receipt capture. Extraction reads a document that is in front of it. Matching that document to the right transaction is a second decision, and noticing the document that never arrived is not something the model is looking for at all.
Reconciliation. Matching two populations against each other is arithmetic, and software has done it for years. The residual is the job, and what a finished reconciliation has to contain does not change with how the matching was done.
Close. Here it hands back. Accruals, estimates and the judgments that decide what the statements say are not coding decisions, so the close sequence runs the way it always did, on a ledger that arrives in better shape.
The study reports one thing about the close, and it is timing. Accountants using generative AI closed the month-end books 7½ days sooner than those who did not (Journal of Accountancy, August 2025). Read it as an association among firms that adopted, not as a promise about yours.
Hold a demo against the split below, row by row, and ask which column the software is being sold on.
| Cycle step | What a model takes | What comes back to a person |
|---|---|---|
| Feed intake | Reading and importing what the feed delivers | Noticing a feed that stopped, or a quiet week that is a broken connection |
| Transaction coding | A proposed account for each line, with a confidence level | Every line under the threshold, and every rule the model inferred wrongly |
| Document capture | Extracting fields from an invoice or receipt | Matching the document to its transaction, and chasing the one that never came |
| Reconciliation | Matching two populations and listing the residual | Explaining the residual, and calling it an error or a timing difference |
| Close | Assembling schedules from what is already posted | Accruals, estimates and every significant judgment |
The Exception Queue Is Where the Work Moves
AI coding does not remove bookkeeping work from a firm. It converts it into review work, and the conversion is not one for one.
The Stanford and MIT study puts a size on the shift. It reports that 8.5% as a pickup of about 3.5 hours in a 40-hour workweek, and the high value tasks it names are business communication and quality assurance (Journal of Accountancy, August 2025). The hours did not leave the week. They changed job.
Now look at what the coding step produces. AI adoption in the study was linked to a 12% increase in general ledger granularity, measured by the number of unique accounts used to categorize transactions (Journal of Accountancy, August 2025). Finer granularity is usually better reporting. It is also more distinct coding decisions per period, and every one of them is a place a reviewer can disagree.
Then read what these tools are said to free a team up to do, as a job description rather than as a benefit. They let teams "focus on identifying, analyzing, and managing unexpected anomalies in workflows and reviewing new or unusual transactions" (Journal of Accountancy, January 2026). That is the exception queue, and somebody in your firm works it every week.
The size of that queue is a setting before it is a result. Raise the confidence threshold and more lines stop for a person, so the queue grows and the ledger gets safer. Lower it and the queue shrinks, because more lines post on their own, including the ones the model got wrong and nobody now looks at. An automation rate, the share of lines the software posts without stopping for a person, describes where that dial was left, not how often the software was right.
One honest note on the evidence. The authors surveyed 277 accountants and partnered with a company providing AI-based accounting software to analyze transactions at 79 small and midsize firms (Journal of Accountancy, August 2025). So the time figures are self-reported, from a survey that asked accountants to document three consecutive workweeks of their own work, and the transaction data came from a vendor in the category. The findings describe firms that chose to adopt. They tell you the shape of the change rather than its size in your practice.
What Your Firm Owes on an AI-Coded Ledger
AR-C section 70 reaches an accountant in public practice engaged to prepare a client's financial statements and not engaged to audit, review or compile them, at .01, and merely assisting management in preparing them is a bookkeeping service the section does not cover, at .02. Where it does apply, it does not ask you to verify what you were given. It says such an engagement does not require the accountant to verify the accuracy or completeness of the information provided by management, or otherwise gather evidence to express an opinion or a conclusion on the financial statements, at .04. It then says the accountant should prepare the financial statements using the records, documents, explanations, and other information provided by management, at .13 (AICPA, Statements on Standards for Accounting and Review Services).
Hold that against a ledger the client's own tool coded. The AI output is the records provided by management. It arrives with the standing a shoebox of receipts has, and no more.
The duty that follows is the one to read slowly. If the accountant becomes aware that the records, documents, explanations, or other information, including significant judgments, used in the preparation of the financial statements are incomplete, inaccurate, or otherwise unsatisfactory, the accountant should bring that to management's attention and request additional or corrected information, and where management fails to provide it, should disclose the material misstatement in the financial statements or withdraw from the engagement and tell management the reasons, at .17 (AICPA, Statements on Standards for Accounting and Review Services).
Awareness is the trigger, and awareness is a function of how much you look. A firm that reads nothing becomes aware of nothing. The section does not send you looking, which is exactly why the depth of your review is your firm's own decision, and it is the difference between a defensible file and a lucky one.
A third requirement is worth reading against the AI case. Where the accountant assists management with significant judgments about amounts or disclosures reflected in the statements, the accountant should discuss those judgments with management so management understands them and accepts responsibility for them, at .16 (AICPA, Statements on Standards for Accounting and Review Services). Some coding decisions are exactly that, because they change what a line means rather than where it sits.
That requirement turns on the accountant assisting, not on the software acting. Where your firm's tool produced one of those judgments during the preparation, the discussion is still owed, and it has not happened until somebody has it. Where the client's tool produced it and your firm only accepted the output, the requirement never attaches, and the judgment sits in the statements with nobody having discussed it.
Which standard an engagement runs under is a prior question, settled at acceptance rather than at the close. The ladder from preparation to compilation to review is laid out in audit, review and compilation, and the scope change coming to preparation engagements is covered in client accounting services.
Write the Exception Standard Into the Scope
Before you accept a ledger a model coded, whether the tool is yours or the client's, write down what an exception is and who owns it. Five things belong in that note.
- What counts as an exception. A coding below the confidence threshold, a vendor with no history, an account that has never carried this kind of line, and any single line above a size you set.
- Which accounts get read in full regardless of confidence. Revenue, payroll, related party accounts, and anything feeding a covenant or a return.
- Who resolves an exception, and what a resolved one leaves behind. A resolution that exists only as a changed account is a decision nobody can grade later.
- Who may change a coding rule, and where the change is recorded. A model that learns a wrong rule applies it across many lines at once, which is why the weekly re-read of a sample already sits in the between-closes cycle.
- What happens when the vendor changes the model. Your exception rate is not comparable across a version change unless somebody re-baselines it on known work.
This is a standard about reviewing model output. It is a different question from keying accuracy, where the single-pass default sets the bar. It is also not a controls document: whose login the software acts under and which actions it may commit without a person are decided per action, before any of this.
When a Client Arrives With Books Their Own Tool Coded
This is the case the buying decision usually skips, because it is not a purchase at all. The client is not asking you to do the bookkeeping. They are asking you to accept it.
Four things decide whether you can. Ask which periods the tool coded and under whose rules, because a ledger coded on the client's defaults is not the same artifact as one coded on yours. Ask whether a person ever reviewed the opening balances you would be inheriting. Ask whether the client can still produce the source documents behind the coded lines, since a coding without its document is an assertion. And ask whether remediating prior coding sits inside the recurring scope or stands as separate work with its own start and end, because that is the boundary firms discover in March rather than agree in advance.
The engagement letter is where those answers survive a disagreement, and its required contents are already set out in the client onboarding checklist. What an AI-coded ledger adds is a named source for the records and a stated boundary around remediation.
The advice runs the other way too, at your own use of these tools. Firms are told to consider having disclosure language added to all engagement letters to inform clients that AI tools may be used in the provision of professional services. The same guidance is firmer on review. It says AI should not be used to make, finalize, or support decisions related to client services, such as tax filings, the issuance of opinions on financial statements, or advisory recommendations, without thorough human review and approval (Journal of Accountancy, July 2026).
Questions Firms Ask
Can AI Do My Bookkeeping?
It can do the coding step and hand back an exception queue. Whether that is a good trade depends on two things you can check before buying anything: whether your firm has the review capacity to work the queue every week, and whether the client's feeds and documents are in good enough shape for a model to have something to learn from. Where the cycle is undocumented, a model produces a ledger that looks right and cannot be checked.
Can ChatGPT Do My Bookkeeping?
A general assistant can help you reason about a transaction. It is not connected to the ledger, so it cannot post one, and what you may safely paste into it is a separate decision with its own rules. Software that does post is a different purchase, and what decides it is permissions rather than intelligence, taken per action.
Buy the Review, Not the Automation Rate
An automation rate is a claim about where a dial was left. Ask the harder question instead. On a month of one real client's transactions, what did the exception queue look like, who worked it, and how much of what posted without stopping held up when somebody checked it?
Capacity is usually the reason a firm reaches for any of this, and the test is the same whether the coding was done by software or by a person: grade real output on files where you already know the answer. Don't trust us. Test us. Accountably places trained offshore accountants and tax preparers inside US CPA and EA firms, and the Free 40-Hour Proof Pilot puts a fixed block of your own representative work through multi-layer review so your reviewer grades it first. See how the pilot works.
