Build1 publisher3 min readPublished
OpenAI's Node SDK defaults outlast Vercel's 300-second limit on a hung model call
OpenAI's Node SDK gives each attempt 10 minutes by default, so on Vercel a hung model call is killed at 300 seconds before its catch block can log anything. The call's own limits have to be shorter than every clock in front of it, or the platform ends the request and nothing gets recorded.
The Engineer · Build desk

What happened
- Six of 215 verified launch failures in a dev.to review shared one pattern: an AI call with no timeout, no budget and no fallback.
- Pointed at a local stub that accepts connections and never replies, the example Next.js route kept curl waiting for at least 15 minutes.
- When the stub returned 500 instead, the SDK retried for a few seconds and the route answered 200 with its fallback sentence, so monitors recorded a success.
- At OpenAI's hard spend limit, requests get a 429 with code project_spend_limit_exceeded, and the SDKs retry it even though it will not clear until the limit resets.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint On Vercel the first default attempt is only half finished when the function dies, so the SDK's retry count does nothing for a hung call there.
- exposure An app that logs only inside its catch block records nothing when the platform kills the process, and a stalled model call can run unnoticed, as it did in most of the six cases.
- cost Because spend limits are monthly and project-wide, one runaway agent job can cut off every user at once unless the app enforces its own per-user or per-run cap.
- decision Apps calling both OpenAI and Anthropic need spend-limit handling for two different status codes, a retried 429 and a 400, before blanket retry rules make a spent limit worse.
Every layer between the browser and the model keeps its own clock. In the default setup, the one in the app's own code is the longest [12]. The post's example route builds its client with a bare `new OpenAI()` [5]. Its authors say most generated code ships in that shape [22]. Vercel's 300-second cutoff lands halfway through the first 600-second attempt [1]. The two retries never get a turn.
Vercel terminates a function that runs past its `maxDuration`. The `catch` block never runs, and neither does any logging inside it [16]. Behind Cloudflare, the user gets a 524 even sooner [9]. In most of the six cases, nobody noticed for a while because the app did not log model calls or their outcomes [4].
The 15-minute local wait has a stated cause. Both official Node SDKs retry connection errors, 408, 409, 429 and 5xx twice with a short backoff, and they retry timed-out requests too [11]. The post counts three attempts of at least 300 seconds each [11]. At the documented 10-minute default, three attempts can run to 30 minutes before backoff is added [2].
The fallback also hides errors in the cost data. Failed calls use few tokens, so the cost chart stays flat [17].
Streaming fixes a separate failure. A non-streamed response sends nothing until the model finishes, so every proxy in front of it sees an idle connection [13]. Anthropic's SDK docs recommend streaming for long requests. The SDK also throws an error for a non-streaming request expected to run past roughly 10 minutes unless a timeout is set [14]. That check is good design. It makes the caller pick a number before a long request starts. The same SDK raises its default timeout above 10 minutes, up to 60, for non-streaming requests with a large `max_tokens` [15].
Spend limits look different at each provider. Anthropic answers a reached workspace limit with a 400 `invalid_request_error`, and retry logic that only looks for 429 will not recognise it [20]. Provider limits are monthly and cover a whole project or workspace. Nothing caps spend per user or per run [18]. The post describes an agent or background job that burns a week's budget in a day, trips the limit and cuts off every user at once [21].
How often this happens outside the sample has not been measured. If the 4% share of AI failures applies to the 215 verified cases, that is about nine cases, and six of them fit this pattern [3]. The sample is builders who posted publicly about apps that broke at or after launch [1]. The reproduction is the stronger evidence. The SDK reads `OPENAI_BASE_URL`, so a local server that accepts connections and never replies can stand in for a stalled provider [6].
In my view, those clocks set the order of work. I would give the call a timeout short enough that all its attempts together fit inside the platform's `maxDuration`. I would log the fallback as a failure. And I would keep a per-user or per-run budget in the app, since the provider's cap is monthly and project-wide. That order fits a serverless route with a hard ceiling, which is the setup the post tested on Vercel [8].
What to watch
- The remaining behaviours of the post's fake-provider test suite, which starts with a provider that never answers and an agent that never finishes, and the limits it proposes for each.
- Any change to the OpenAI Node SDK's 10-minute per-attempt default or its retry-on-timeout behaviour, which would move the gap against Vercel's 300-second cutoff.
- A larger failure review than six cases showing whether AI-related 504s in production trace to calls with no timeout of their own.