Blog

Agentic AI in Accounting: What Changes When the Software Acts

Agentic AI in accounting acts inside your systems. Decide whose login it uses, which actions it may commit unattended, and what your audit trail will show.

Accountably Editorial Team 11 min read Updated 2026-08-14

The familiar kind of AI in an accounting firm drafts. It writes a memo, proposes a coding, summarizes a document, and then it stops, and a person decides what happens next. Agentic AI in accounting is the category where the software does not stop. It takes the next step itself, inside your systems, under some account, and leaves behind a record with a name on it. That is a different purchase from a smarter assistant, and it needs a different set of questions before anyone turns it on.

What Agentic AI in Accounting Actually Means

The line is autonomy, not intelligence. Agentic AI refers to artificial intelligence systems that function as autonomous agents capable of independently making decisions, learning from interactions, and adapting to changing environments. Unlike traditional AI models, it exhibits independent reasoning, goal-driven behavior, and the ability to interact dynamically with users, systems, and real-world scenarios (NIST, agentic AI).

The profession's own description is blunter. Agentic AI "is an AI that has agency," with the ability to act autonomously and make decisions, where generative tools like ChatGPT "require user-provided, step-by-step prompts" (Journal of Accountancy, June 2025).

For a firm, though, the definition is settled by permissions rather than by the vendor's language. Ask what the system is allowed to write to. If the output lands in a chat window, a draft, or a suggestion someone accepts, it drafts. If it posts to the ledger, sends a client an email, files something, moves data between two applications, or changes a record without waiting, it acts. The same product can do both, in different modules, on the same afternoon.

Every Action Commits Under Somebody's Identity

The moment software acts, the action carries a name, and the name is whatever credential it used.

Security architecture already has a plain term for an actor that is not a person. A zero trust design, one that grants nothing on the basis of network location alone, has to know who its subjects are, and subjects could encompass both human and possible non-person entities, such as service accounts that interact with resources, at section 7.3.1 (NIST Special Publication 800-207, Zero Trust Architecture). A service account is the login issued to a non-human actor, so its work is attributable to it and not to a colleague.

NIST's own treatment of this sits in a section about agents that administer the security controls themselves, and it is candid that the ground is unfinished. Artificial intelligence and other software-based agents "are being deployed to manage security issues on enterprise networks," need to interact with the management components of the architecture "sometimes in lieu of a human administrator," and how they authenticate "is an open issue," at section 5.7 (NIST Special Publication 800-207, Zero Trust Architecture). The same passage notes that a software agent "may have a lower bar for authentication (e.g., API key versus MFA) to perform administrative or security-related tasks compared with a human user," and that an attacker "could gain access to a software agent's credentials and impersonate the agent when performing tasks." The identity problem it names does not stay inside security tooling.

Read that against the rule most US firms are already covered by, and the wording is worth reading closely. The FTC Safeguards Rule requires multi-factor authentication "for any individual accessing any information system," unless the Qualified Individual you designate to oversee your information security program "has approved in writing the use of reasonably equivalent or more secure access controls," at section 314.4(c)(5), and it requires policies, procedures and controls designed to monitor and log "the activity of authorized users," at section 314.4(c)(8) (eCFR, Safeguards Rule elements). Those are the two clauses you already work through when a new person is provisioned.

Both are written around people. An agent is not an individual, and the rule's defined term does not obviously reach it either: an authorized user "means any employee, contractor, agent, customer, or other person that is authorized to access any of your information systems or data," at section 314.2(a) (eCFR, Safeguards Rule definitions). The word agent sits in that list, but the sentence closes on "or other person," so the category it names is a person acting for you, not a piece of software.

Which means the rule does not answer the identity question for you, and one of its own clauses makes the answer matter: access controls have to "authenticate and permit access only to authorized users," at section 314.4(c)(1)(i) (eCFR, Safeguards Rule elements). An actor you cannot place in that category is one you cannot cleanly permit. Give the agent credentials of its own and the authentication and logging clauses have something to attach to. Run it on a staff member's login and you have settled it by default, in her name.

That failure is quiet, and it surfaces late. Six months on, a peer reviewer or your client's auditor reads a change history that attributes a run of postings to someone who was not at her desk, and nothing in the file says otherwise. Ask any vendor two things in writing: does the agent get an account of its own, and does its own name appear in the record of every action it takes.

The Tax Standard Already Names Artificial Intelligence, and It Names It a Tool

The AICPA revised its Statements on Standards for Tax Services effective January 1, 2024, and the revision added a standard on relying on tools. The SSTSs apply to AICPA members providing tax services, which is narrower than every firm that prepares returns. A tool is defined there as a resource used in the provision of tax services, and the list of examples runs from tax preparation software and research publications through data analytics and statistical models to artificial intelligence, at paragraph 1.4.2 (AICPA, Statements on Standards for Tax Services).

The standard permits reliance and then closes the exit. A member may reasonably rely on tools used in providing tax services to a taxpayer, and "use of a tool does not absolve the member of professional obligations under AICPA or other applicable ethical standards," at paragraph 1.4.4 (AICPA, Statements on Standards for Tax Services).

The diligence duty is stated as a positive obligation rather than as advice. A member who employs tools remains responsible for the completed work product, and "accordingly, members should take reasonable steps to determine that the tools used are appropriate for the intended purpose," at paragraph 1.4.7 (AICPA, Statements on Standards for Tax Services).

Then comes the sentence worth sitting with. Tools "should be used to enhance or improve the member's understanding of a tax issue, not to supplant the member's professional judgment," and the example given is Form 1040, where a member must still attest under penalties of perjury that, to the best of the preparer's knowledge and belief, the return and accompanying schedules are true, correct and complete, because "that responsibility cannot be transferred entirely to reliance on a tool," at paragraph 1.4.8 (AICPA, Statements on Standards for Tax Services).

That language describes a resource you consult. A system that files, posts or sends is not improving anyone's understanding of anything, so applying the standard to an actor is a reading rather than a ruling, and reasonable people will land in different places on it. What is not a reading is who carries the duty: the standard puts the reasonable steps on the member who employs the tool, not on the vendor who sells it. The same logic sits under Treasury practice rules, where the reliance presumption is written around a person you engage, supervise, train and evaluate rather than around a product.

Decide the Approval Gate Per Action, Not Per Tool

Firms tend to decide this at the wrong unit. The question gets framed as whether to adopt a product, when the decision that actually protects you is which individual actions that product may commit without a person in the loop.

NIST's AI Risk Management Framework is voluntary, and it still puts the process itself on the list. Processes for human oversight are to be "defined, assessed, and documented in accordance with organizational policies," at MAP 3.5 (NIST, Artificial Intelligence Risk Management Framework, AI RMF 1.0). Defined and documented is a higher bar than a shared understanding that somebody checks the important ones.

The same sorting question turns up in the profession's own reporting, with a shape attached. It is important to determine where a human needs to be involved, and that "includes determining critical decision points and thinking about processes that might be difficult to undo or might cause immediate issues if done incorrectly" (Journal of Accountancy, June 2025).

Reversibility and visibility are the two axes that fall out of that, and they sort an action list into three tiers.

Commit unattended. Actions that stay inside your walls and that a person sees before anyone outside does. Assembling a workpaper, pulling documents into the right client folder, preparing a reconciliation for review, populating a checklist from source documents. A mistake here is caught by the review that was always going to happen.

Stop for a human before it commits. Anything that leaves the building or changes a number someone relies on. A client email, a transmission, a posting that moves a reported balance, a change to a recurring schedule. The cost of a wrong action here is not the correction, it is the conversation.

Never delegated at all. Some of these your firm has already settled. Payment release and changes to the vendor master file belong inside the client's own controls, and approval authority for journal entries is written by entry type before the first close. An agent does not change either answer. It only makes the question arrive faster.

One design note decides whether the middle tier is real. A gate works only if the reviewer can see what is about to happen in a form they can check in seconds. A prompt that shows the proposed entry, the account, the amount and the source document is a control. A queue that asks for one click to approve a batch of pending actions is a rubber stamp with an audit trail attached.

Diligence Questions for Something That Acts

Software diligence usually asks about data: where it lives, who can see it, how it is encrypted. Those questions still apply. An actor adds five more.

What can it write to, and what is the smallest permission set that still does the job? Start from the list of systems and end at the narrowest scope, because a permission granted for a pilot is rarely taken back.

Whose identity does it act under? If the honest answer is a shared login or a staff member's credentials, you have decided the audit trail question by accident.

What does the record show, and can you read it without the vendor? You want the sequence of actions, the inputs, the timestamp and the actor, exportable, because the person who needs it will be reconstructing a decision long after the run.

How do you stop it mid-run, and what happens to half-finished work? NIST's framework names "the ability to shut down, modify, or have human intervention into systems that deviate from intended or expected functionality" among the practical approaches to AI safety (NIST, Artificial Intelligence Risk Management Framework, AI RMF 1.0). Ask what a stop leaves behind.

How does it leave? The same framework asks for mechanisms to inventory AI systems, at GOVERN 1.6, for processes to decommission and phase them out safely, at GOVERN 1.7, and for internal risk controls over the components of an AI system, third-party technologies included, to be identified and documented, at MAP 4.2 (NIST, Artificial Intelligence Risk Management Framework, AI RMF 1.0). Decommissioning is the one firms skip, and the credentials an agent held outlive the subscription unless somebody revokes them.

Independent assurance still has its place alongside those questions, and which report to ask for depends on the work the system touches. A report is evidence about a vendor's controls. It is not an answer about your own permission boundary, which nobody else can draw for you.

When the Honest Answer Is Read-Only

If you cannot say whose identity an action will commit under, do not grant write access yet. Run the same system in a mode where its output is a draft that a person acts on, which is exactly the non-agentic mode, and lose nothing except the step you were not ready to supervise.

The same goes for work that is not written down. An agent that chooses its own steps through an undocumented process will produce a plausible sequence nobody can grade, and it will do it quickly. Where automation genuinely attaches in a return pipeline is a question about the form of the input, and it is answerable before any agent is involved.

Questions Firms Ask

How Does Agentic AI Reach an Accounting Firm?

Expect it to arrive rather than be bought. Entry into everyday use "won't necessarily come with signs popping up to announce, 'You're now using agentic AI!'", and for accounting and finance work it "will likely be phased into existing software" as more capable features inside tools a firm already owns (Journal of Accountancy, June 2025).

That has a governance consequence worth planning for. The adoption decision may reach your firm as a release note rather than as a purchase order, which puts the gate with whoever reviews changes to the systems that hold client data, alongside the written program the Safeguards Rule already asks for.

How Is This Different From Automation a Firm Already Runs?

Scripted automation follows a sequence somebody wrote. An agent picks the sequence. One practitioner quoted in the profession's own reporting expects agentic AI to accomplish what robotic process automation was promised to do and often missed, "because of the inflexibility and complexity that popped up when it was used" (Journal of Accountancy, June 2025).

The control consequence is the part that matters for a firm. With a script, you can read it and enumerate every action it will ever take. With an agent, you cannot list the action set in advance, so the control moves off the script and onto the permission boundary and the approval gate. That is why the tiers are written per action rather than per tool.

Write the Gate Before You Grant the Access

Take one process you would hand an agent tomorrow and list the actions it would perform, not the outcome you want. Mark each action reversible or not, and visible to a person before it leaves the firm or not. Assign each to a tier, then grant only the permissions the top tier needs and hold the rest back until the log has shown you something real.

Capacity is usually the reason a firm reaches for any of this, and the test that protects your name is the same whether the work is done by software or by a person: grade real output on files where you already know the answer, before anything is committed. Don't trust us. Test us. Accountably places trained offshore accountants and tax preparers inside US CPA and EA firms, and the Free 40-Hour Proof Pilot puts a fixed block of your own representative work through multi-layer review so your reviewer grades it first. See how the pilot works.

See the work before your name is on it

Run a Free 40-Hour Proof Pilot on your own representative work, through full multi-layer review, before a single client file moves.