Skip to content

Invest1 publisher3 min readPublished

Banks' data nutrition label asks credit bureaus to sign for the goodness of what they sell

The Financial Services Sector Coordinating Council workgroup's label grades data quality by use case and asks vendors to attest to it, which makes it a procurement instrument rather than the model-governance extension it resembles.

The Investor · Invest desk

What happened

  • A workgroup of the Financial Services Sector Coordinating Council built a data "nutrition label" meant to check that the data fed to AI models meets minimal requirements, with PNC's Ned Carroll among those who worked on it.
  • PNC wrote a separate generative AI risk policy rather than stretching its model risk management, on the grounds that deterministic models had a ground truth in structured data and probabilistic ones do not.
  • Carroll says OpenAI and Anthropic would like PNC to open its data for training and will not get it, because the bank's IP is embedded in its policies, procedures and unstructured sources.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint With no regulator quoted as adopting or requiring it, the label's enforcement mechanism is the purchase order, so it binds the vendors a bank is large enough to make sign and nobody else.
  • contradiction Read through SR 11-7 the label looks like model governance stretched over generative systems, but the practitioner who helped build it says generative AI risk was kept out of model risk management on purpose.
  • decision Refusing to let frontier labs train on bank data commits the bank to buying inference and owning the context layer itself, along with the quality burden for everything in it.
  • exposure An attestation names an accountable party outside the bank, which turns a stale or wrong vendor field from an operational surprise into somebody's signature.

The label asks less of the data behind a marketing offer than of the data behind a loan approval, because the loan is where the bank exposes risk, as Carroll puts it [5]. That taxonomy grades by credit exposure, and the cautionary case American Banker sets alongside it ran the other way. Air Canada's virtual assistant gave a customer the wrong answer about a bereavement discount, the customer sued, and a Canadian court made the answer stick [10]. A customer-facing sentence sits low on the nutritional scale and high on the enforceable-promise scale, and the label as described does not close that gap.

Regulators published SR 11-7 on managing, monitoring and validating models in 2011 [8]; in 2012 Jamie Dimon's pay was cut in half over a $6 billion trading loss that stemmed partly from a model's use of faulty risk-management data [9]. The guidance was one year old [16]. That is the case for and against a documentation framework in the same two dates, and PNC did not build this one as an extension of the old one: Carroll says the bank was deliberate about not conflating model risk management with generative AI risk management and wrote a separate policy, because the deterministic models had a ground truth in structured data and probabilistic systems have none [6][7].

The verb doing the work here is attest. The owner of the data signs for it, so a bank leaning on credit bureaus, market data providers and other vendors calls on them to attest to accuracy and timeliness [4], and Carroll's framing of goodness as a function of inputs, timeliness, quality and currency, weighted by what the data is used for [3], specifies a warranty rather than a measurement. This is a procurement instrument. If the attestations come back as boilerplate that every vendor signs, the label changes no behaviour, and the first thing to fail is Carroll's own test, that without clear accountability the ability to manage quality is compromised [14]. What would break the reading is a bank repricing or dropping a feed over a failed attestation. American Banker's account describes none.

The banks are declining to do something with their own asset. Carroll is explicit that OpenAI and Anthropic would like PNC to open its data for training and will not get it, because the IP is embedded in policies, procedures and unstructured sources [13]; frontier models trained on all of the internet generally lack data quality controls [11], so PNC and some other banks add their own knowledge and context on top [12]. Every quality claim a bank can honestly make about a generative output is therefore a claim about that added layer and the vendor feeds beneath it, not about the model, which is why the label's useful question to a model provider, what nutrition went into your model [15], has no answer at the frontier [11].

As for supervision, the account attributes the label to the industry workgroup and quotes no regulator adopting, requiring or citing it [17][1]. It is industry attestation carrying industry leverage, which for the banks big enough to make a bureau sign is leverage, and for everyone else is a form somebody else fills in.

What to watch

  • Whether the FSSCC publishes the label's actual field set, and whether vendor attestations show up as contract clauses rather than questionnaires.
  • Any supervisory reference to data provenance attestation in exam guidance or a supervisor's remarks, which is what would convert procurement leverage into a requirement.
  • Whether other large banks copy PNC's split and write separate generative AI policies, or fold generative systems into existing model risk management.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories