Build1 publisher3 min readPublished
Your Next.js rate limiter counts per instance, and Server Actions hide behind the page URL
A dev.to field report argues the ten-line in-memory limiter multiplies your limit by warm instance count, and that path-based middleware rules cannot see a Server Action at all.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author writes that moving a simple rate limiter into a Next.js 16 App Router project on Vercel broke it in three unexpected places: where the counter lives, who counts as an identity, and how you attach a limit to a Server Action that has no URL of its own.
- A module-scope Map is private to one function instance, and a serverless platform runs many instances at once, so each instance enforces the full limit on its own; the effective limit is the configured limit multiplied by the number of warm instances.
- With ten warm function instances, a 10-requests-per-minute limit permits 100 requests per minute.
- The ten-instance example represents a ten-fold overshoot of the configured limit.
- The number of warm instances is not something you control or can observe from inside the request.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A set of field notes published on dev.to takes the ten-line rate limiter that most backend engineers have shipped and shows it breaking in three places when moved into a Next.js 16 App Router project on Vercel: where the counter lives, who counts as an identity, and how you attach a limit to a Server Action that has no URL of its own [1]. Both of the first two failures are quiet, which is the consequence worth caring about: the limiter still returns 429s, still looks right in review, and is wrong by a factor nobody in the request can measure.
Start with the arithmetic. A module-scope `Map` is private to one function instance, and a serverless platform runs many instances at once, so each one enforces the full limit on its own [2]. With ten warm instances, a limit of 10 requests per minute permits 100 [3] - a ten-fold error in the only number the policy actually specifies [4]. The instance count is not something you control or can observe from inside the request [5].
Vercel Fluid Compute makes this harder to notice rather than easier, according to the author: it reuses a single instance across concurrent requests instead of spawning one per request, so the `Map` survives far longer than under classic serverless [6]. In local development and in a quiet preview deployment the limiter looks correct, and it only comes apart under traffic spread across enough instances to matter, with nothing in the logs announcing it [7]. The same `Map` never shrinks, so on a long-lived instance every unique key ever seen stays resident until recycling, which the piece describes as a slow memory leak wearing a rate limiter costume [8]. The stated fix is not a cleverer `Map` but a store every instance shares that supports an atomic increment: Redis, or any datastore with a compare-and-set primitive [9].
The second failure is addressing, and it is the one that survives a Redis migration. Every Next.js Server Action POSTs to the URL of the page that called it and carries a build-generated `Next-Action` header, so path-based limiting in middleware cannot tell one action from another [10]. The consequence in the source is blunt: the only reliable place to limit a specific action is inside the action body [11]. A middleware matcher of the shape `['/api/:path*', '/login', '/signup']` [12] therefore never sees a sign-up action invoked from a page outside that list, because the POST is addressed to the calling page [13].
Identity is the third break. `NextRequest.ip` was removed in Next.js 15, and on Vercel the author reads the client address with `ipAddress(request)` from `@vercel/functions` rather than trusting a raw `x-forwarded-for` header [14].
The division of labour that falls out: middleware runs before Next.js resolves the route, so a request rejected there never boots the route's function and never touches your database, which makes it the right place for a coarse, identity-agnostic abuse limit [15]. The route handler knows the authenticated user, the parsed body and the business meaning of the call, which makes it the right place for a per-user quota [16]. Middleware also sees RSC prefetch requests that carry the `RSC: 1` header and that the user never intentionally made [17], so those get counted as traffic unless skipped. Rejections should carry status 429 with a `Retry-After` header [18], and the fail-open versus fail-closed choice should be made per route before Redis has its first outage [19].
Worth checking in your own codebase: whether any limiter state lives in module scope, whether your sensitive Server Actions are limited in the action body or only by a matcher, and whether a hover-triggered prefetch is spending someone's quota.