Whether AI has reached a piece of accounting work has very little to do with what the task is called. It turns on what the work receives at its front door: data that arrived in named fields, a document somebody has to read, or a set of facts nobody has written down yet.
The IRS handbook for authorized e-file providers makes the point in a single instruction. It tells firms to check that their software does not auto-populate one particular identifier, because the number in that field has to come off the document in front of the preparer.
Run your own task list through that one question and AI in accounting stops being a category argument. Three shapes, and each leaves a different job for a person.
The Test Is the Shape of the Input, Not the Name of the Task
Data means values that arrived in named fields, written by another system. Nothing was interpreted, so there is nothing to check for a misreading.
A picture of data, the document shape, means a document whose values a person can read and a program cannot, until something reads them for it. That reading is new information which did not exist before, and it belongs to whoever accepted it.
Judgment written down means facts that were never in fields at all: a contract, an incomplete set of records, a client explaining what happened. Reading those is the work rather than the input to it.
The same service line arrives in all three shapes across a client list. One payroll client sends an export, the next photographs a wage statement on a kitchen table, and a third sends a spreadsheet somebody rebuilt by hand. So the sort runs per client and per month, not once per service.
Where the Input Is Already Data
This is the end of the list where the tools work, and it is the least interesting part of the AI question because most of it was solved before models arrived. Coding a bank line to an account is the clearest case, and which steps of the cycle a model takes there is worked through in AI bookkeeping. An invoice that arrives as structured data can clear without a person once the order and receipt it has to match already exist, while the same invoice as a PDF cannot, a split set out in invoice processing automation. Return data a third party already filed can be pulled instead of chased, covered in tax preparation automation.
What a person still owns here is everything the data cannot contain. Completeness is the standing example, and so is deciding what a correctly coded line actually means for this client.
Where the Input Is a Picture of Data
This is where most of the current argument about AI in accounting lives, and it is the one shape that creates work as well as removing it.
What Scan and Populate Actually Produces
Three things happen in sequence, and firms buy them as one product.
Optical character recognition, or OCR, turns an image of text into characters a program can handle. Extraction then decides which of those characters belong to which field, a judgment about layout rather than about accounting. Auto-population writes the result into the working file, where it sits looking exactly like something a person typed.
The reading is a claim, and nobody has checked it yet. Its characteristic failure is not a typo. It is a correct value taken from the wrong place: the state wage box instead of the federal one, a prior-year copy sitting in the same PDF, the second page of a consolidated statement, the spouse's form in a joint bundle. Each of those values is plausible on its own, which is why a reviewer skimming for something that looks wrong will not find it.
The Federal E-File Rules Already Name Auto-Population
One field in the tax return record is fenced off from automatic filling by rule, and the wording is unusually direct about software.
Taxpayers who file using an individual taxpayer identification number, known as an ITIN, may hold a wage statement showing a different number. For those returns, software must require the manual key entry of the taxpayer identification number, or TIN, as it appears on the Form W-2 reporting the wages. Electronic return originators, known as EROs, should also ascertain that the software they use does not auto-populate the ITIN in the Form W-2 (IRS Publication 1345).
Read what that instruction assumes. The system already knows the number that identifies the taxpayer, and the field in front of it asks for the number printed on a document. Filling it from what the system knows, rather than from the paper, is the failure the rule is written against, and it shares its shape with a misread box: a value that reached the field from somewhere other than the document in front of the preparer.
The entry itself has a standard. Providers should take care to ensure that they transcribe all TINs correctly, and the TIN entered in the Form W-2 in the electronic return record must be identical to the TIN on the version provided by the taxpayer (IRS Publication 1345).
Getting it wrong is not a quiet problem. Incorrect TINs, using the same TIN on more than one return, and associating the wrong name with a TIN are named as some of the most common causes of rejected returns. The name control that has to agree with it is the first four significant letters of an individual taxpayer's last name, or of a business name, as recorded by the Social Security Administration or the IRS (IRS Publication 1345).
The Documents Extraction Reads Worst Carry Extra Duties
Documents a person has altered or filled in by hand are the ones extraction handles least reliably, and the handbook attaches its own obligations to a list that includes them.
EROs must always enter the non-standard form code in the electronic record of individual income tax returns for Forms W-2, W-2G or 1099-R that are altered, handwritten or typed, and an alteration includes any pen-and-ink change (IRS Publication 1345).
Two further duties attach to the same documents. Providers must never alter the information after the taxpayer has given the forms to them, and the IRS has identified questionable Forms W-2 as a key indicator of potentially abusive and fraudulent returns, telling providers to be on the lookout for suspicious or altered Forms W-2, W-2G and 1099-R and for forged or fabricated documents (IRS Publication 1345).
Setting a flag and forming a suspicion are not readings. No confidence score performs either one, and a scanner that produces clean fields from a document somebody edited by hand has done its job and left yours untouched.
The Verification Pass Runs Field to Document
Auto-population does not remove the checking. It changes what is being checked, from whether somebody typed the number correctly to whether the machine read the right box on the right document.
That gives the review a direction. Compare the populated field against the image it came from, not against another output that was built from the same reading, and concentrate on the fields where a wrong value survives a plausibility check: identifiers, account numbers, dates, and anything that decides which entity or which year a figure belongs to.
One rule shows how a tidy-looking result can be the wrong one. Addresses on Forms W-2, W-2G or 1099-R, on Schedule C, or on other tax forms supplied by the taxpayer that differ from the taxpayer's current address must be input into the electronic record of the return, even if the addresses are old or the taxpayer has moved (IRS Publication 1345). Software that helpfully replaces an old address with the current one produces a cleaner record and a wrong one, and a reviewer checking the return against the client file rather than against the document will agree with it.
Three nearby decisions are already settled elsewhere. Where the confidence threshold sits, and who works the queue it creates, is the dial covered in AI bookkeeping. What the stored image owes you as a record of its own is in invoice processing automation. If the keying leaves the building instead of going to software, the accuracy standard and its denominator belong in the scope, as set out in accounting data entry services.
Where the Input Is Judgment, Written Down
The input in this shape is a lease, a partnership agreement, a shoebox with three months missing, a client email that tells half the story, or a question about how a rule applies to facts nobody has finished gathering. A model produces fluent output on all of it, and fluency is the difficulty rather than the benefit, because checking the answer can cost close to what producing it would have cost.
Four questions in this shape are worked through where they belong. What a prompt has to carry before it is worth running, and the failure mode that looks most like competence, is in ChatGPT prompts for accountants. What your firm owes on a position a model helped reach is in AI tax preparation. When model output is treated as audit evidence, the reliability questions do not change, and they are in AI in auditing. When the software stops answering and starts acting under somebody's login, the controls are in agentic AI in accounting.
Run the Sort on Your Own List
Take last season's task list and write one of three words beside each line: data, document, judgment. The label is about what arrives, not about how hard the work feels.
The data lines are where a tool pays for itself quickly, and where the remaining human job is completeness rather than accuracy. The document lines are where a tool changes the shape of the work: fewer keystrokes, and a new verification pass that has to be designed, staffed and written down. The judgment lines are where buying software adds a draft to review and nothing else.
Two things tend to fall out of that exercise. A service line sold as one product splits across all three labels, so the demo that impressed you may have been running on the data lines. And a task that looks mechanical often turns out to be a document line, which means the saving is real and smaller than the proposal implied.
Two further decisions sit downstream of the sort and are settled elsewhere. Whether a given line is better answered by software or by another pair of hands is compared task by task in AI versus offshoring in accounting. Whether a tool you already bought ever reached the whole job type is the adoption question in accounting technology adoption.
Questions Firms Ask About AI in Accounting
How Does AI Work in Accounting?
In three different ways, depending on the input. On structured data it classifies and matches. On documents it reads and populates fields, producing a claim your firm then owns. On judgment work it drafts, and the draft has to be verified by someone who could have written it.
Is There an AI That Can Do Accounting?
There are tools that take defined steps well, and the honest test is what a wrong answer costs and how quickly it can be spotted. Coding a recurring bank line is cheap to check. A misread identifier on a return is not, which is why the W-2 identifier on a return filed with an ITIN carries a manual key entry rule.
Is AI Going to Replace Accounting?
The federal labor projections point in two directions at once for two different occupations, which is the actual answer and is set out in the future of the accounting profession. For a firm deciding this season, the practical version is the three-way sort, plus the routing between software and people in AI versus offshoring in accounting.
Start With the Front Door, Not the Software
Pick the job type that cost you the most evenings last season and answer two questions before you look at a single product. What does this work receive, and who checks the reading? A demo cannot answer either one, and both answers are already sitting in your own files.
If the honest answer is that the work arrives as judgment and the queue in front of it is qualified attention, the constraint is capacity rather than software. Accountably places trained offshore accountants and tax preparers inside US CPA, EA and accounting firms, on your software and your SOPs, in about 3 to 4 weeks, with the signature, the opinion and the final judgment staying with your firm. Since 2022 that has meant 20+ US firms and 30+ placements. Don't trust us, test us: the Free 40-Hour Proof Pilot puts a fixed block of your own representative work through full multi-layer review, so your reviewer grades real output before a client file depends on it, and if someone is not the right fit in the first 30 days the 30-Day Fit Guarantee replaces them free.
