Build1 publisher3 min readPublished
ChatGPT's EU search-engine designation moves audit evidence to the retrieval layer
ChatGPT is now a Very Large Online Search Engine under the EU's DSA, designated by the European Commission on Aug 31 with 159.1 million monthly EU users. Qtim CEO Anton Fokin argues that anyone shipping a web-search feature needs logs that can rebuild each answer from its sources.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The Commission called ChatGPT a hybrid service because it both answers prompts and can search the web.
- Once notified, ChatGPT has four months to meet the extra duties, including assessing and mitigating systemic risks, undergoing an independent audit each year and providing data access through set procedures.
- A Microsoft Research preprint from 2026, covering 234,839 public ChatGPT conversations from 2023 to 2025, classed 79% of user inputs as hard to answer with conventional web search.
- On comparable questions that could be searched, ChatGPT's responses covered less diverse information than Google results across most topics.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- precedent With the Commission's label tied to the web-search function, other assistants that browse can expect to be judged against the search-engine threshold once their EU usage passes 45 million.
- constraint For a designated service, Article 34 requires a risk assessment before features likely to have a critical impact on identified risks, so each new browsing capability becomes a release gate.
- exposure Providers based outside Europe are in scope as well, because the DSA covers intermediary services offered to EU recipients regardless of where the provider sits.
Flip on a "search the web" toggle and the interface barely changes, Anton Fokin, chief executive of Qtim, wrote on dev.to [11]. Underneath, the system starts deciding when to browse. It rewrites the user's question, selects sources and compresses them into one answer [11]. His company builds AI chatbots, RAG systems and language-model integrations [10].
Fokin counts at least four decisions behind every generative-search answer: the call to search or not, the queries it sends, the documents and passages it picks, and the policies or filters it applies [13]. The same answer can rest on different sources [14]. And the same model version can give different answers once the web results or an index shift [14]. He wrote that "once a system selects information for users, the final answer stops being enough evidence" [12].
Qtim joins the evidence chain with one request identifier [15]. It ties together the versions of the model and its instructions, each tool call and generated query, the IDs of sources or chunks, policy decisions, the final answer and its citations, plus retries, any human intervention and rollbacks [15]. I think this is the right design for an open-web product. The index moves between the answer and the review, so asking the same question a month later does not reproduce the original decision [14]. Fokin scales the requirement to the system. For a closed knowledge base, document IDs and versioning may be enough [16]. A product that searches the open web needs a log of its decisions, a pipeline for evaluation and a tested means of switching the feature off [16].
The design also declines the obvious shortcut. Keeping every prompt, retrieved page and response forever improves reproduction and creates a much larger privacy and security problem, by Fokin's account [17]. Retention, masking, deletion and access controls belong inside the trace design [17]. "A useful record reconstructs a decision without cloning the conversation database," he wrote [18].
Testing moves down a layer as well. A fixed golden answer breaks whenever the wording changes [19]. Mostly it tests the model's prose style. The steadier checks ask whether the right sources were eligible, whether the expected tool ran, whether a policy fired and whether a high-risk condition stopped the workflow [19].
The Microsoft Research figures describe one product's users. The dataset is observational and does not represent every user [7]. For the diversity gap to transfer, another product's users would need to ask the kinds of questions found in public ChatGPT conversations from 2023 to 2025 [5]. Its retrieval would also need to narrow sources the way ChatGPT's did. Fokin draws the narrower point: a system can accept a broad question and return a narrower range of information [20].
ChatGPT's EU count is about 3.5 times the 45 million threshold [1]. The designation imposes obligations on the named services and does not automatically place every AI product under the DSA [8]. For a team far below that line, the source does not establish a legal duty to build Fokin's request log. His argument is that, outside the DSA, public AI products still need traceability when they choose sources for users [21].
What to watch
- The date of the Commission's formal notification to ChatGPT, which starts its four-month clock for the additional VLOSE obligations.
- Whether the Commission designates other browsing assistants as hybrid services once they report EU usage above 45 million.
- Whether the Microsoft Research diversity finding holds after peer review or on a representative sample of users.