Build1 publisher2 min readPublished
Spring AI fixes its hard-coded 60-second stream timeout in a milestone release
Spring AI 2.0.1 ignores configured timeouts and kills any streaming turn longer than 60 seconds. The fix sits in 2.1.0-M1, a milestone built on Spring Boot 4.2.0-M2, so leaving the one-minute ceiling behind means running a pre-release stack.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Affected streams fail with OpenAIIoException: Stream failed, and the post's advice to anyone seeing it is to stop debugging the proxy.
- The hard-coded per-call limit overrode both spring.ai.openai.timeout and spring.ai.openai.chat.timeout.
- Spring AI 2.1.0-M1 shipped on 25 September 2026, and the post says the timeout fix alone justifies the ten minutes the upgrade takes.
- The same release rebuilds AssistantMessage, UserMessage and ToolResponseMessage as ordered lists of typed MessagePart objects.
- It also backfills additionalProperties: false into strict-mode tool schemas, drops an empty text block that drew a Bedrock 400, and closes a TextReader file-descriptor leak.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams stuck at the ceiling have to weigh a hard one-minute limit on 2.0.1 against a milestone whose APIs can still change before release.
- constraint Upgraders who use a persistent chat memory repository get working conversations, but the model starts its reasoning over on every turn, because only InMemoryChatMemoryRepository keeps parts.
- capability Agents on GPT-5.4 or later can combine tool calling with a reasoning effort by setting spring.ai.openai.chat.api=responses. Chat Completions does not support that pairing.
The limit applied per call, according to the dev.to post that documented the regression [1]. On a stream, a per-call limit covers the whole turn. A response still producing tokens at second 59 was cut at second 60 [15].
The post's author wrote that the lesson keeps coming back in payments work: "the shortest timeout in the chain wins, and it's usually one you didn't set" [16]. The usual candidates are a load balancer, a gateway or an SDK default. This time it was a framework regression [18]. Sixty is the kind of number a person types. "When a call dies at a suspiciously round number of seconds, that's rarely a coincidence," the author wrote [17].
The post names 2.1.0-M1 as the release with the fix [4]. It does not mention a 2.0.x patch that carries it. For a service on a GA Spring Boot line, taking the fix puts two pre-release versions in the stack at once: the Spring AI milestone and Spring Boot 4.2.0-M2 [1].
The code change is smaller than the version jump suggests. According to the post, existing code keeps working. getText(), getMedia() and getToolCalls() are now views over the new parts, and the old constructors still produce the legacy order [11]. Streaming subscribers that read getText() see the same deltas as before [11].
The parts design is careful work. OpaquePayload holds provider data that has to be replayed unchanged, such as an Anthropic thinking signature or a Gemini thought signature [13]. Without a defined place for that data, the post says, a framework either drops it or grows a provider-specific hack. If it drops it, the model loses its train of thought partway through a tool loop [13]. UnknownPart keeps the raw JSON of block types the adapter does not model yet [14].
The parts work is also unfinished. In M1, only OpenAiResponsesChatModel produces and consumes parts natively. The other chat models get refactored in RC1 [7]. In my view the milestone is the right choice for a service whose turns regularly run past a minute and whose team already runs Boot milestones. A service that ships only on GA releases is different. There I would keep streaming turns under 60 seconds on 2.0.1 and take the fix from a later release, after the remaining adapters have had their parts refactor [7].
What to watch
- A 2.0.x patch release that carries the timeout fix would let teams get their configured timeouts back without adopting the milestone.
- Spring Boot 4.2 reaching general availability would leave Spring AI's own milestone as the only pre-release in the upgrade path.