Skip to content

Build1 publisher2 min readPublished

Firecracker MicroVMs inside Netlify's own network cut Edge Function overhead to 5-6 ms

Netlify moved Edge Functions off hosted V8 isolates onto Firecracker MicroVMs in its own network, cutting warm overhead from 25-40 ms to a 5-6 ms median. A dev.to analysis argues that for other edge teams, removing an external network hop can matter as much as a faster runtime.

The Engineer · Build desk

Illustration accompanying Firecracker MicroVMs inside Netlify's own network cut Edge Function overhead to 5-6 ms

What happened

  • Netlify also reports 47.4 percent faster p99 invocations, 99.998 percent availability and log delivery about five times faster than before.
  • Unikraft, which worked with Netlify on the execution layer, reports a 2 ms p99 for MicroVM startup across Netlify's fleet.
  • Netlify says its Edge Functions handle about a billion invocations a day, so any per-request overhead is paid at very large volume.
  • About 1.2 percent of invocations hit a node that has never served the function and must fetch its images before running anything.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Netlify measured its warm figure after the outside hop was gone. A team that moves to MicroVMs but keeps an external execution provider in the path should budget for that hop on its own.
  • cost Owning the path moves work the hosted provider used to absorb, including provisioning, image distribution, placement and logging, onto Netlify's own operations staff and fleet.
  • capability Because the service ID comes from the spec plus site data, a platform can give every deploy its own MicroVM and still send that deploy's traffic to nodes that already hold its image.

The two speed figures measure different events. Unikraft's startup number is the time to boot a MicroVM [5]. Netlify's headline number is median overhead on a warm invocation [3]. On each matching request, the platform has to pick compute, start or reuse an execution environment, run the customer's code and return the result [6]. In the usual serverless sense, warm means the reuse branch. Boot time should therefore sit mostly outside both the old and new warm figures. The warm saving still works out to roughly 19 to 35 ms per request [1].

The route is the likelier source. Before the move, a matching request left Netlify's network for a hosted V8-isolate service. That service ran the code and handled provisioning and much of the runtime environment [7]. Now the edge node that terminates TLS checks the deployment's routes and hands a match to a regional compute node inside Netlify's network [8]. The dev.to analysis puts placement, network hops, image availability, service lookup, logging and failure recovery on the critical path next to runtime startup [9]. Netlify's reported figures do not break the overhead down by component. They show stronger per-deploy separation and lower latency arriving in the same move, and they do not show which change saved the milliseconds.

Customers write Edge Functions largely as before [2]. The new work sits underneath. Before forwarding, the edge node writes a machine spec with three image layers (the runtime, a platform image and the customer's function image) plus CPU, memory and connection limits [10]. Netlify hashes that spec with site-specific information into a service ID. Different deploys or environment-variable sets get different services and do not share a MicroVM [11]. The analysis describes this as making deployment identity part of the isolation boundary instead of a convention on a shared process [12]. I think it is the strongest part of the design. The service ID that decides which MicroVM a request may share is also the key that rendezvous hashing uses to choose a node [13].

Sticky placement helps on repeat traffic. The same service normally lands on the same node in a region. An image fetched once stays on disk, and a running MicroVM or snapshot stays useful [13]. Sticky placement works well until one tenant has a very good day. Above a threshold, Netlify spreads a hot service across a slice of nodes and gives up some cache warmth for capacity [14].

The slow path that remains is image distribution. Applied to Netlify's daily volume [1], the cold-path share [15] comes to roughly 12 million requests a day that wait for image layers before any customer code runs [2].

What to watch

  • Whether Netlify publishes a breakdown of the warm-overhead drop that separates the removed network hop from runtime and placement changes.
  • Latency on the cold image-fetch path, and whether its roughly 1.2 percent share shrinks as images spread across regional nodes.
  • Whether the threshold at which Netlify spreads a hot service across several nodes becomes documented or configurable per site.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories