Skip to content

Invest1 publisher3 min readPublished

Your SOX Agent Vendor Has 90 Days To Prove It Is Not A Chatbot

Average programs run $2.3 million and 15,000 hours while 57% of internal audit functions hold headcount flat. The diligence is now about evidence layers, not demos.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Every vendor now promises agents that read the evidence, test the samples, and hand a finished workpaper to a reviewer.
  • KPMG's 2025 SOX Survey puts the average SOX program at $2.3 million a year.
  • KPMG's 2025 SOX Survey puts the average SOX program at more than 15,000 hours a year.
  • The Internal Audit Foundation found 57% of internal audit functions holding headcount flat while their mandate grows.
  • A year ago, AI tools for SOX testing could barely handle a clean three-way match.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

Every vendor in the space is now selling agents that read the evidence, test the samples and hand a finished workpaper to a reviewer [1]. They are pitching against a function that, per KPMG's 2025 SOX Survey, costs an average of $2.3 million and more than 15,000 hours a year [2][3], staffed by teams where the Internal Audit Foundation found 57% of internal audit functions holding headcount flat while the mandate grows [4]. That combination makes the buy decision hard to defer; the open question is what you are actually buying. The capability shift is recent. A year ago, according to a column in CPA Practice Advisor by the author of The Audit Leader's Guide to AI for SOX Testing, these tools could barely handle a clean three-way match; since the end-of-2025 model releases they have been working through complex reconciliations and hundred-megabyte spreadsheets [5][6]. The firms that will judge the output are building their own: KPMG has published a framework for generative AI in SOX, Grant Thornton a seven-step path, and Deloitte and EY each have multibillion-dollar AI audit platforms in progress [7]. They believe the technology works, and they will look at how you apply it [8]. The arithmetic explains the urgency. Roughly 40% of those 15,000 hours goes to testing [9], which is about 6,000 hours a year [10], and the program cost spread across the hours implies about $153 of spend per hour [11]. Question one is whether the thing is an agent or a chatbot with a better interface. A chatbot answers a question; an agent is given a goal, the rules for reaching it and a set of tools, and returns structured work a reviewer can inspect [12]. A SOX test is a chain of steps, and a long prompt can describe the chain but cannot reliably run it [13]. Teams that tried prompt libraries got stuck: nobody owned the prompts, nobody could measure the output, and every control needed its own [14]. The failure mode that matters is accuracy you cannot locate. A model right 95% of the time is useless if you cannot tell which 5% to check, because then you check everything and redo the work the tool was meant to save [15]. Question two is evidence. SOX evidence is screenshots, PDFs, spreadsheets, emails and system exports with no naming convention, with about one in five PBC requests coming back wrong the first time [16]. That mess is what broke robotic process automation, which Deloitte found only 3% of companies ever scaled [17]. The recommendation is to bring your own ugliest files to the demo and refuse the vendor's samples [18]. Question three is the audit trail, and the column specifies five layers: an approved testing plan signed off before testing starts, source evidence stored exactly as received, an extraction layer that can say what was extracted, from which file, where in that file and how, an assessment layer that turns those facts into a result with plain-English reasoning, and human review, which is where most of the auditor's time now goes [19]. The design assumption is that the same test run twice can produce slightly different answers [20]. There is no PCAOB standard for AI in audits, and the staff's July 2024 observations never became guidance, which leaves AS 1215 as the bar: a file an experienced auditor who never saw the engagement can follow [21]. Question five is the pilot: five to 10 controls, one owner, and a baseline recorded before you commit to a vendor, covering hours per control, PBC follow-up rate, rework hours and time to a final workpaper [22]. Run old and new on the same controls and compare time, documentation quality, reviewer comfort signing off, and whether both runs caught the same exceptions [23]. The stated payback is four to six months [24]. Watch two things.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories