Science1 publisher3 min readPublished
ChatGPT told users Biden was still running 10 days before the 2024 election
A peer-reviewed audit scored the free ChatGPT model's election answers on one day in late October 2024, and it counted hedging and omission as failure categories alongside outright error.
The Scientist · Science desk

What happened
- Ten days before the 2024 election, and more than three months after Harris took the top of the Democratic ticket, ChatGPT was still telling users that Joe Biden was running for a second term.
- Of 592 verifiable claims the model made on Oct. 25, 2024, 9.6% were inaccurate, among them a statement that term-limited North Carolina Governor Roy Cooper was seeking reelection.
- Separately from outright error, the researchers recorded hedging in 18% of answers and logged responses that omitted critical information about efforts to overturn the 2020 election.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- exposure The tested build was the free default, so the error rate describes the answers available to anyone who simply typed a question; a paid or retrieval-backed election feature has to be measured on its own.
- decision Anyone testing an assistant on electoral questions has to sample low-salience races, where the training data is thinnest, and not only the presidential contest.
- constraint Jang locates the remedy outside the model: "We need journalism that not only delivers truth, but also helps people evaluate the truth," she said.
The 9.6% works out to about 57 wrong claims in a single day of answers [1]. Two of the errors the study names concern who was on a ballot: Biden running for a second term, and Roy Cooper seeking reelection in North Carolina, where he was term-limited [1][5]. Harris had led the Democratic ticket for more than three months by that date [1].
The design was built so that the conversation could not teach the model. Heesoo Jang and her co-authors opened new chats and used incognito browsing so earlier prompts would not shape later answers, and they wrote prompts in both partisan and nonpartisan framings, covering election information, disinformation and threats, candidates, and Project 2025 [7][6]. The build under test was ChatGPT-4o mini, the free default at the time [3].
That buys a clean reading of one model on one date, a measure of what the model said when prompted. What users actually asked, and what they did with the answers, would be a different study.
Hedging got its own count: qualifying phrases such as "could potentially," "generally aligned" and "typically," returned 18% of the time [9]. The published account does not state the base for that share, so it cannot be stacked on the 9.6% [21]. Omission shows up as answers that left out critical information about efforts to overturn the 2020 election [10].
The third category, which the authors abbreviate as B.S., arrives as examples. Asked about the "DEI hire" pejorative used against Harris, the model returned a job description, defining it as someone hired to oversee diversity, equity and inclusion efforts at an organization [11]. Partisan actors had redefined the term, and a claim-by-claim fact check would not mark that answer false [11].
It also produced policy platforms for North Carolina's statewide school superintendent candidates in generic language that came out nearly identical [12]. "It was trying to fill in the gaps," Jang said [13]. "It can't give us, 'I don't know,' because that's not how it is designed," she said [14].
Jang, who teaches media law and ethics at UMass Amherst, said: "We should be concerned and alarmed by the fact that while ChatGPT is nailing exams, it is also failing in these areas that it wasn't designed for" [19][15]. The study calls the pattern a "careless failure with profound democratic implications" [16].
phys.org reports that ChatGPT models have been updated since 2024 and would likely perform better today, and that Jang is not optimistic about the 2026 midterms, particularly as AI increasingly replaces internet searches and links to primary sources [17]. Her position is that accurate electoral information and the handling of disinformation cannot be trained from past datasets alone. Models will struggle to catch up unless they are designed to prioritize democracy, which she says is difficult when development is driven by commercial incentives [18]. The paper, written with Shannon C. McGregor and Lorcan Neill of UNC-Chapel Hill, is published in Information, Communication & Society [20][8].
What to watch
- Whether the authors rerun the same protocol on a current model before the 2026 midterms, which would show whether the claim-level error rate moved.
- Whether any audit of a paid tier or a competing assistant reports a comparable claim-level error rate on election questions.
- Any response from OpenAI to the specific errors logged on Oct. 25, 2024.