Skip to content

Invest1 publisher2 min readPublished

Advisors building their own tools face the testing burden of maintaining them

A Microsoft Macabacus survey puts AI-generated errors in 62% of teams and comprehensive guardrails in 24%, and the advisors quoted alongside it describe prompt-built apps that need retesting to return the same answer twice.

The Investor · Invest desk

Illustration accompanying Advisors building their own tools face the testing burden of maintaining them

What happened

  • Advisors are using vibe coding to turn natural-language prompts into client-facing tools, work that previously took hours or days of Excel and PowerPoint time.
  • According to American Banker, PwC joined EY and KPMG this past summer in a public "walk of shame" after AI-related inaccuracies turned up in their reporting.
  • A report by Microsoft's Macabacus found that 62% of its teams believed they had shipped a model or a presentation containing an AI-generated error.
  • The same report found 24% of Macabacus users have comprehensive guardrails, while 36% use AI daily or weekly with none at all.
  • Aditi Kapadia of Wealth IQ has vibe coded three applications for her Denver practice: a cashflow management tool, a money mindset assessment and a goals visualizer.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • cost Building in-house moves upkeep onto the advisor's own hours, a cost that falls outside any licence fee or vendor support contract.
  • exposure Client email addresses, names, account numbers and tax IDs can end up inside a tool whose data path the advisor who prompted it cannot describe.
  • constraint Compliance teams have to review software they never commissioned, and on Macabacus's own numbers most respondents are unsure whether their firm has the controls to do it.
  • contradiction The documented AI failures are in reports, models and presentations, while the tools under discussion are applications holding client data, so the liability case for vibe coding is an inference from adjacent evidence.

In this account the cost comes after the build. Aditi Kapadia of Wealth IQ found the building fun. But she said "maintaining and updating these applications is time consuming," and that the cashflow app "requires a lot of testing and updating to make it generate the same outcome repeatedly" [7][8]. An application that returns different numbers from the same inputs fails a test every advisor already applies to a spreadsheet. Getting it to pass is a recurring hour cost, and whoever wrote the prompt carries it.

Kapadia also narrowed her scope. She started out thinking advisors "could vibe code their way into most features they needed from their tech stack." She has since concluded that advisors "may want to vibe code their way to an app (or two) to plug any existing tech stack gaps" [9][10]. Mike Wilson, chief executive and co-founder of Hamachi.ai, drew the same boundary in build terms. "It's pretty darn easy now to build a prototype," he said. "It's still very difficult to take that prototype and make it something that you should feel comfortable putting client information in" [11][12].

The survey numbers rest on a narrower base than the percentages suggest. They describe Macabacus's own teams and users, and American Banker did not report a sample size or a field date [20]. Within that base, 45% said the correct guardrails existed at their firm and 55% said no or were not sure [5]. The share who believe they have shipped an AI-generated error runs 17 points above the share who believe their firm has the controls [19].

Sean Sandys, chief technology officer of Syntax Data, said advisors risk treating these tools "like they're a senior executive." He offered the control: "You wouldn't take something that your junior analyst generated and then put it in front of an external partner, a client or customer without reviewing it" [13][14]. Todd Wardzinski, writing on Red Hat Developer, said "the code itself becomes the only source of truth for what the software does" [15].

In my view the first bill here is hours. The review Sandys describes is a per-artifact cost that scales with the number of prompt-built tools a practice keeps in service. The counter-thesis is that guardrail platforms, Wilson's own company among them, cut review cost far enough that a prompt-built app really does displace a licence [12]. What would settle it is an error rate measured on vibe-coded applications; if that number lands at or below the rate for bought software, the upkeep argument fails.

What to watch

  • An error rate measured on prompt-built advisor applications, reported with its sample size, would test whether the 62% deck-and-model figure holds up there.
  • Macabacus publishing the sample size and field dates behind the 62%, 24% and 45% splits.
  • The first client complaint or examination finding that names a vibe-coded advisor tool.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories