Skip to content

Build1 publisher3 min readPublished

OpenAI indexes its financial data providers in-house to make every number clickable

The engineering in OpenAI's financial-services launch sits in the retrieval path, and the flagship model underneath it is the one the API already serves. That split decides what a buyer can build for itself.

The Engineer · Build desk

Illustration accompanying OpenAI indexes its financial data providers in-house to make every number clickable

What happened

  • OpenAI launched ChatGPT for Financial Services on September 10, 2026, and the model underneath is GPT-6 Astra, its current flagship, with no finance-tuned variant built for the product.
  • Around that model OpenAI added built-in financial data, firm-specific templates and a governance setup a compliance team can sign off on.
  • The built-in data at launch comes from Daloopa, PitchBook, LSEG News and Crunchbase, with Quartr included according to reporting around the release.
  • Morgan Stanley and Evercore worked as design partners, and OpenAI says that relationship pointed the first release at investment banking and equity research.
  • Astra is the first OpenAI model to reach the Critical tier of the company's internal cybersecurity capability framework, and OpenAI is gating some access because of it.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision A build-or-buy call here turns on the licensed corpus, not the model: templates and access rules are ordinary internal work, while an in-house copy of five providers' data starts with five commercial agreements.
  • constraint Because retrieval runs against OpenAI's own copy of the provider data, index freshness becomes an OpenAI operations question that the firm has to take on trust.
  • exposure A bank deploying agents on Astra is working with a model OpenAI itself places in its Critical cybersecurity tier, so access approval lands inside procurement.
  • capability With close to a million tokens in one session, cross-document work over a 10-K, years of transcripts and a comp set no longer needs a chunking pipeline the firm maintains.

The buildable half of this product is the half a large bank has already half-built. Templates carrying house style are an internal project. So is a rule about which desk may open which deal file. What a firm cannot stand up in a quarter is a licensed, indexed copy of five data providers' output sitting on the vendor's own hardware. Four of the five appear on the launch list; the fifth, Quartr, rests on reporting around the release [11][15].

OpenAI says hosting and indexing that data itself, instead of calling out to each provider on every request, improves retrieval speed, latency and how reliably it can trace a generated claim back to its source [12]. The retrieval reason is stable identifiers. An index you control gives every chunk an id that still resolves when a reviewer clicks it a month later, while five live vendor APIs put five availability windows and five rate limits in the answer path. The trade-off is freshness, and the write-up does not say how often the index refreshes.

That choice buys citations. The product surfaces granular citations so a user can check a number against where it came from while they work [13]. According to the dev.to account, the failure being designed against is a general chat model with no persistent access to properly entitled data and no awareness of a firm's house style: the output reads well and fails compliance review [14].

OpenAI publishes two benchmark numbers with the model, and both are claims about someone else's documents. OpenAI reports Astra about ten points above the previous model on OfficeQA Pro, which it describes as a rough proxy for office and financial document work [6]. For that to predict anything on an equity research desk, the desk's files and its questions would have to look like the benchmark's. On agentic, terminal-style tests, the same account says Astra beat its predecessor and Anthropic's comparable model at the time, at a lower cost per task [7]. That result describes OpenAI's bench that week, since the competitor is identified only as the comparable model of that moment.

The API exposes a reasoning effort setting with levels running from low up to max [5]. The analyst in the chat window has no equivalent dial; the setting sits with the team building on the API. That is where per-call spend gets decided, because summarizing a memo and building a comp table with sourcing should not cost the same.

The dev.to post paraphrases Nick Turley, OpenAI's VP of product, as describing the work as teaching ChatGPT to research like an analyst and back up its conclusions the way an analyst would [10].

What to watch

  • A published refresh cadence for the hosted index would turn staleness from a matter of trust into a measurable one.
  • Whether entitlement checks run against a firm's own systems or are configured inside OpenAI's governance layer.
  • Whether OpenAI adds data providers beyond the five named around the launch.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories