Invest1 distinct publisher3 min readUpdated
Average programs run $2.3 million and 15,000 hours while 57% of internal audit functions hold headcount flat. The diligence is now about evidence layers, not demos.
The Investor · Invest desk
Compiled by The InvestorSomething wrong?How this is made
Every vendor in the space is now selling agents that read the evidence, test the samples and hand a finished workpaper to a reviewer [1]. They are pitching against a function that, per KPMG's 2025 SOX Survey, costs an average of $2.3 million and more than 15,000 hours a year [2][3], staffed by teams where the Internal Audit Foundation found 57% of internal audit functions holding headcount flat while the mandate grows [4]. That combination makes the buy decision hard to defer; the open question is what you are actually buying. The capability shift is recent. A year ago, according to a column in CPA Practice Advisor by the author of The Audit Leader's Guide to AI for SOX Testing, these tools could barely handle a clean three-way match; since the end-of-2025 model releases they have been working through complex reconciliations and hundred-megabyte spreadsheets [5][6]. The firms that will judge the output are building their own: KPMG has published a framework for generative AI in SOX, Grant Thornton a seven-step path, and Deloitte and EY each have multibillion-dollar AI audit platforms in progress [7]. They believe the technology works, and they will look at how you apply it [8]. The arithmetic explains the urgency. Roughly 40% of those 15,000 hours goes to testing [9], which is about 6,000 hours a year [10], and the program cost spread across the hours implies about $153 of spend per hour [11]. Question one is whether the thing is an agent or a chatbot with a better interface. A chatbot answers a question; an agent is given a goal, the rules for reaching it and a set of tools, and returns structured work a reviewer can inspect [12]. A SOX test is a chain of steps, and a long prompt can describe the chain but cannot reliably run it [13]. Teams that tried prompt libraries got stuck: nobody owned the prompts, nobody could measure the output, and every control needed its own [14]. The failure mode that matters is accuracy you cannot locate. A model right 95% of the time is useless if you cannot tell which 5% to check, because then you check everything and redo the work the tool was meant to save [15]. Question two is evidence. SOX evidence is screenshots, PDFs, spreadsheets, emails and system exports with no naming convention, with about one in five PBC requests coming back wrong the first time [16]. That mess is what broke robotic process automation, which Deloitte found only 3% of companies ever scaled [17]. The recommendation is to bring your own ugliest files to the demo and refuse the vendor's samples [18]. Question three is the audit trail, and the column specifies five layers: an approved testing plan signed off before testing starts, source evidence stored exactly as received, an extraction layer that can say what was extracted, from which file, where in that file and how, an assessment layer that turns those facts into a result with plain-English reasoning, and human review, which is where most of the auditor's time now goes [19]. The design assumption is that the same test run twice can produce slightly different answers [20]. There is no PCAOB standard for AI in audits, and the staff's July 2024 observations never became guidance, which leaves AS 1215 as the bar: a file an experienced auditor who never saw the engagement can follow [21]. Question five is the pilot: five to 10 controls, one owner, and a baseline recorded before you commit to a vendor, covering hours per control, PBC follow-up rate, rework hours and time to a final workpaper [22]. Run old and new on the same controls and compare time, documentation quality, reviewer comfort signing off, and whether both runs caught the same exceptions [23]. The stated payback is four to six months [24]. Watch two things.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
KPMG has published a framework for generative AI in SOX, Grant Thornton a seven-step path, and Deloitte and EY each have multibillion-dollar AI audit platforms in progress.
Every major firm is building AI testing tools of its own: they believe the technology works, but they will look at how you apply it.
SOX evidence includes screenshots, PDFs, spreadsheets, emails and system exports with no naming convention, and about one in five PBC requests come back wrong the first time.
Messy evidence is what broke robotic process automation, which Deloitte found only 3% of companies ever scaled.
KPMG's 2025 SOX Survey puts the average SOX program at $2.3 million a year.
KPMG's 2025 SOX Survey puts the average SOX program at more than 15,000 hours a year.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One vendor-authored source, third-party figures uncited
The cluster contains a single article. Its checklist content is internally coherent and specific, but every external number (KPMG 2025 SOX Survey, Internal Audit Foundation 57%, Deloitte 3% RPA) is cited without a link or reference, the one-in-five PBC failure rate has no attribution at all, and the central capability and payback claims have no benchmark, pilot result, or customer behind them. Nothing here is corroborated by a second publisher.
Firm-level tooling programs, no documented buyer deployments
Adoption signals are all supply-side: published frameworks from KPMG and Grant Thornton, asserted Deloitte and EY platform programs, and a generic claim that every vendor promises finished workpapers. No named company has been shown running SOX testing agents in production, no pilot outcomes are disclosed, and the article's own precedent is that only 3% of companies scaled RPA. The 90-day pilot framing itself implies the category is still pre-deployment for most buyers.
Self-limiting checklist carrying two unsupported upside claims
Modestly overstated overall. Much of the piece is deflationary: demos are dismissed, RPA's 3% scaling rate is raised, nondeterminism is conceded, and most auditor time is said to still go to human review, all of which pulls the narrative toward evidence. The gap comes from two vendor-favorable assertions with no support, a sudden post-end-of-2025 jump to complex reconciliations and hundred-megabyte spreadsheets, and a four-to-six-month payback, plus the sweeping 'every vendor' and 'multibillion-dollar platform' framing that makes the category look more mature than any disclosed deployment shows.
Advice written by a vendor CEO promoting his own guide
Disclosed but material. The author is CEO and co-founder of Bead AI, a company in the SOX AI testing category, and the article is built around the evaluation framework from his own book, which is offered as 'vendor-neutral'. The five criteria and the five-layer architecture double as a specification a buyer would take into competitive demos, and the unsourced payback figure is the kind of claim a seller benefits from. The piece is hosted by a trade publisher whose page gates the resource behind sign-in, adding a lead-capture incentive.
Low, single interested source
Confidence is limited by one publisher, one interested author, and no corroboration for the numbers or the capability and ROI claims. It is not lower because several checkable elements hold up on their face (the AS 1215 documentation bar and the absence of a PCAOB AI standard, the named KPMG and Grant Thornton artifacts) and because the prescriptive pilot and audit-trail content is specific enough to be tested by a reader without trusting the author.
invest
Pleasant, and lonelier: a 12,365-person trial cuts against the AI companion pitch1 distinct publisher
product
Rillet's $1bn bet: rebuild the general ledger around agents, not bolt a copilot onto it2 distinct publishers
invest
FASB settles its 2027 succession early, and the read is continuity not reset1 distinct publisher
science
Two governments got AI slop in writing, and the review chain caught none of it1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026