Build1 publisher2 min readPublished
Gimlet plans 100 MW of Cerebras inference capacity to share each workload with GPUs
Gimlet Labs and Cerebras plan 100 MW of wafer-scale inference capacity for Gimlet Cloud, run alongside GPUs in the same workload. Buyers can assess the phase-splitting design today, but its one speed figure, 3,000 tokens per second, is a target they cannot price yet.
The Engineer · Build desk

What happened
- Gimlet's software splits inference into prefill, decode, attention and feed-forward phases, then schedules them across GPUs and specialized accelerators.
- Cerebras said an integrated Gimlet system is already serving tokens in private deployments, built from customer work that began in 2025.
- The first Cerebras-powered Gimlet Cloud data center is expected to come online later in 2026.
- Cerebras named Gimlet a launch partner for its CS-4 system, with Gimlet Cloud customers expected to get access in 2027.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams budgeting inference for 2026 still have to size contracts on GPU quotes they can measure; the Gimlet figure enters that comparison only once its test conditions are published.
- constraint Any plan that depends on CS-4 hardware moves to 2027 at the earliest, so 2026 capacity decisions can only cover the first data center and the private deployments.
- exposure Gimlet owns the orchestration layer, so customers calling its standard API absorb its scheduling and handoff choices as latency they cannot tune themselves.
Zain Asgar's argument is that one inference request contains stages with different computing needs, so a cloud built around a single accelerator type leaves performance or capacity unused [8]. Gimlet says the Wafer Scale Engine's on-chip memory and bandwidth suit memory-intensive work, including token generation [11]. GPUs stay in the design for throughput, the property Cerebras co-founder and CTO Sean Lie credited them with. The plan is to run both architectures inside one workload [12].
I think the design is right. Matching each stage to the hardware that suits it is a scheduling problem. Gimlet puts the scheduler behind a standard inference API, so developers do not manage the hardware [10]. The cost is the handoff. When prefill lands on one kind of machine and decode on another, the output of the first stage has to move before the second can start, and that transfer happens inside every split request [9].
That layer is Gimlet's to build. Cerebras brings the wafer-scale systems, while Gimlet contributes orchestration, infrastructure and the developer-facing cloud [13]. Asgar has built cluster software before. He co-founded Pixie Labs, which made Kubernetes observability software and agreed in 2020 to be acquired by New Relic [15]. Gimlet describes its platform as spanning data centers, systems software, workload orchestration and developer APIs [14]. Four product lines is a wide scope for a company that announced an $80 million Series A in March [16].
Asgar said the combined system is planned to reach up to 3,000 tokens per second, a figure the companies present as a target [6]. The materials do not identify the model, prompt length, batch size, concurrency or measurement method, or say whether the figure is per user or for the whole system [7]. For the number to transfer to a buyer's workload, it would have to be per-user output speed on a model the buyer runs, at the buyer's prompt lengths and concurrency [7]. Aggregate throughput across a loaded system is a different quantity. It does not line up against a per-user GPU quote [7].
The 100 MW is a capacity target [2]. A $300 million Series B led by Andreessen Horowitz on September 4 brought Gimlet's disclosed funding to $380 million [17][1].
What to watch
- Published measurements from the private deployments that state the model, batch size, concurrency, and whether tokens per second is per user or per system.
- Location and initial capacity of the first Cerebras-powered Gimlet Cloud data center when it comes online later in 2026.
- How Gimlet moves work between GPUs and wafer-scale systems when one request is split across them, and what that handoff adds to latency.