Build1 publisher2 min readPublished
Teams leaving OpenAI's retired Assistants API must rebuild thread history by hand
OpenAI retired the Assistants API on August 26, 2026, without an automatic tool to move old Threads into the Conversation objects that replace them. Teams that relied on managed threads now write that migration and set how much history each turn sends.
The Engineer · Build desk

What happened
- In the Responses API, a file_search tool backed by a vector store succeeds the old Retrieval tool for searching private documents.
- The dev.to post calls the move a non-trivial migration for any app built on the assumption that OpenAI would manage state.
- Managed file_search takes chunking strategy out of the developer's hands, and the post says that can hurt retrieval on highly structured documents.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Every app now sets its own rule for how much history each turn sends, and that rule decides whether the promised cost predictability shows up on the bill.
- cost Migration effort lands on each customer, who has to write a Thread-to-Conversation conversion or start existing users with empty history.
- constraint Teams indexing contracts, tables or other structured documents may still need their own retrieval pipeline, because file_search chunking cannot be tuned.
The post's own code sample shows what a turn looks like after the switch. It creates a conversation with client.conversations.create() and posts the user's question with client.conversations.messages.create(). It then calls client.responses.create() with model "gpt-6.1-sol", the conversation_id and a file_search tool pointed at one vector store [7]. The author describes it as "a conceptual look at what an API call might look like" [8]. The method names are an illustration until someone checks them against OpenAI's reference [8].
The post calls the new model stateless [2]. In the sample, though, the messages go to a Conversation object through OpenAI's API, and the response call points at it by ID [7]. That puts the history with OpenAI. Calling that stateless is a generous use of the word. The developer now owns the lifecycle: the post says Conversation objects are something "you must create and manage," and that keeping context between turns falls to application code [3].
The post's case against the old API was cost. A whole thread "might be re-processed on every turn," it says [4]. A Conversation referenced by ID grows the same way, one message per post [7]. If each response call feeds the full transcript to the model, input per turn climbs with conversation length, as it did under threads. The post says the new design brings more predictable costs [5] but does not show token counts or a trimming method. For that claim to hold in a given app, the app has to cap what each turn carries. A window of recent turns or a running summary would do it, enforced in the team's own code [3].
Where old Thread data is still reachable, the migration is code each team writes [9]. In the sample's terms, that means creating a Conversation for each Thread worth keeping, replaying its messages through conversations.messages.create(), and storing the new ID wherever the old thread ID lived [7].
Retrieval is the part that got easier. A team creates a vector store and uploads files, and OpenAI's backend chunks, embeds and indexes them [11]. At call time the model decides whether to search, pulls the relevant passages and cites them in its answer [12]. That removes an ingestion pipeline from the team's codebase. For prose-heavy document sets, I think it is the right default. Chunking is the part a team gives up [13]. For contracts or table-heavy reports, I'd run the same queries against the managed store and the existing pipeline before deleting the existing one.
What to watch
- Whether OpenAI keeps Assistants API Thread data readable after the August 26 shutdown; without that, old history cannot be carried into Conversations.
- Published token counts for long Conversations passed by ID, showing whether input per turn grows with the full transcript.
- Any chunking controls OpenAI adds to file_search vector stores, the main reason left to keep a self-built RAG pipeline.