Science1 publisher2 min readPublished
A library dean would have readers classify an AI answer before trying to verify it
The University of Virginia's dean of libraries argues that accuracy is only one test of an AI answer, and sorts replies into four types that each permit a different kind of checking. The demonstration is two Google searches.
The Scientist · Science desk

What happened
- The University of Virginia's university librarian and dean of libraries proposes sorting AI-generated answers into four kinds: factual, interpretive, constructive and strategic.
- Google's AI answer to a question about teenage screen time gave two hours as a limit, then said pediatric guidance weighs the quality and context of use more heavily than the number of hours.
- The American Academy of Pediatrics says there is no exact recommended amount of screen time for teenagers, and emphasises the kind of use and what activities it displaces.
- In a second search about taking a daily aspirin, the system warned about risks, said to consult a medical professional, and offered tailored guidance in exchange for age, cardiovascular history and bleeding risk.
- All four types of answer arrive in much the same fluent, authoritative form, which is why the author says the differences between them are easy to miss.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint The habit most readers have, clicking through to the cited source, settles one of the four types and leaves the other three open.
- decision Opening a source is wasted on a draft eulogy, and reading for voice is wasted on a date of founding, so the type has to be picked before the work starts.
- exposure A reader who verifies only figures is the one exposed here. The two-hour number is the checkable-looking part of a reply whose own next sentence says hours are not the test.
- capability Asking a search engine for the strongest evidence against the answer it just gave is available to anyone, at the price of one more query, and it works without a second source.
The scheme was first proposed in the Journal of Academic Librarianship. The public demonstration of it consists of two queries the author ran on Google search, one about teenage screen time and one about daily aspirin. Nobody sampled users, ran a comparison condition or measured whether people who label an answer before reading it end up better informed. "What interested me was that they were different kinds of replies," the author wrote of the two answers.
Accuracy has been the center of the argument about AI answers, and the author grants that it matters, treating it as one test among several. The four types do not leave a reader the same options. A factual claim can in principle be checked against evidence, and the instruction is to follow the cited link, since the answer on its own is not proof. An interpretive answer rests on which evidence was included, which was left out, and whether another defensible reading exists. A constructive answer is made, and a draft eulogy is judged on purpose, audience and voice. A strategic answer combines information with judgment about goals, risks, trade-offs and personal circumstances.
For interpretive replies the author offers a specific move: ask the search engine "What is the strongest evidence for a different conclusion?" That costs one additional query and depends on the system producing evidence against the answer it has just given.
The aspirin search is the strategic case. The author writes that the system's caution there, including the warning about risks and the referral to a clinician, matches U.S. Preventive Services Task Force guidance. The answer moves with age, cardiovascular history and bleeding risk, and no link settles it for the person asking.
In my view the sort is worth the few seconds it takes, on the narrow ground that it tells a reader which check is even available before any effort goes into checking. The author is explicit that the four categories are not airtight boxes and that a single response can reflect several types at once. "A beautifully written strategic document may not be true," the author wrote.
What to watch
- Any experiment that compares verification behaviour between readers given the four labels and readers given none.
- Whether Google's AI answers keep qualifying their own numbers, as the screen-time reply did, as the product changes.
- Whether library professional bodies build the four-type sort into AI literacy training for staff and students.