Build1 publisher3 min readPublished
Fail closed, not fluent: separating the jobs an LLM should never have had
A developer's design note on OpenThesis argues that unreliable AI company research is an architecture problem, not a prompting one, and splits it into failure modes with different fixes.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author of a dev.to post titled "I stopped letting LLMs guess financial facts" describes building OpenThesis after finding that LLM reasoning about companies could be useful while the financial facts underneath it were much harder to trust.
- OpenThesis is described as an Apache-2.0 desktop system for evidence-first, AI-assisted company research.
- The author states that the observed failures are different failure modes, and that treating all of them as one giant prompting problem did not feel like a reliable architecture.
- The named failure modes are: mixing fiscal periods, accounting scopes or currencies; missing values quietly becoming zeros; a deterministic calculation being performed probabilistically; and a citation pointing to a real filing without actually supporting the claim.
- In the common company-question-to-LLM-to-answer workflow, a single model call is implicitly responsible for remembering reported values, selecting the right fiscal period, recognizing the accounting scope, finding sources, performing calculations, comparing scenarios, identifying risks, and writing a conclusion.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
The developer behind OpenThesis has published a design note on dev.to explaining why company-research workflows built around one large model call keep producing numbers that read well and do not hold, and has released the system as an Apache-2.0 desktop tool for evidence-first research [1][2]. The useful contribution is not the tool, it is the taxonomy: the post insists that several distinct failure modes are being treated as a single prompting problem [3].
The modes it names are genuinely separate. A model can mix fiscal periods, accounting scopes, or currencies; missing values can quietly become zeros; a calculation that is deterministic can be performed probabilistically; and a citation can point at a real filing that does not actually support the claim attached to it [4]. None of those is fixed by better instructions, because each one has a different mechanism.
The diagnosis is structural. In the naive pipeline, a single model call is implicitly responsible for remembering reported values, selecting the right fiscal period, recognising the accounting scope, finding sources, performing calculations, comparing scenarios, identifying risks, and writing the conclusion [5]. Per the author, qualitative reasoning, connecting evidence, forming scenarios, and challenging assumptions are reasonable uses of a language model, while recalling an exact reported value, deciding whether a value is missing, and computing a margin or valuation are poor places to accept probabilistic behaviour [6]. The stated rule: deterministic work stays deterministic, and the model is not the database or the calculator underneath the reasoning [7].
Implementation follows the rule. Ingestion starts from official filings rather than model memory or general web search, currently SEC EDGAR for the US, CNInfo for SSE, SZSE and BSE companies, and HKEXnews for Main Board and GEM issuers [8]. The pipeline is staged from filings through evidence extraction, validated financial facts, deterministic finance, specialist agents, synthesis, verification, and a traceable thesis [9]. Extracted evidence is assigned IDs, agents receive a bounded evidence set and are expected to attach those IDs to factual claims, so unknown references and unsupported claims can be flagged rather than accepted because the prose sounds plausible [10]. A final verification step can mark a run partial when required evidence coverage is absent [11].
The arithmetic gets the same treatment. Financial summaries and reverse-DCF calculations are written in ordinary code, missing values are not silently converted to zero, and the pipeline fails closed when core coverage is insufficient instead of handing an incomplete numeric picture to a model [12]. The example given is revenue growth with one period missing: a fluent model will still emit a percentage because completing patterns is what it does, while a deterministic function returns "not available", preserves the reason, and keeps the gap from contaminating downstream work [13].
There are eight declared roles, covering financial analysis, business analysis, accounting risk, growth, skepticism, forecasting, synthesis and verification [14][15]. The author is explicit that agent count is not the claim, since eight agents holding eight inconsistent versions of the facts would produce more confidence without more reliability; what matters is that the roles operate over the same research pack and evidence model [16][17]. The post also concedes that none of this removes extraction errors or resolves ambiguous filings, only that it makes the failure boundary inspectable [18], and it presents no measured error rate or benchmark [19]. It is not a stock picker or a trading bot [20].
Worth watching: whether partial runs stay labelled partial once someone wants a finished report, and whether extraction accuracy holds equally across three filing regimes with very different formats. Coverage and error statistics per jurisdiction would tell you more than the architecture diagram does.