Skip to content

Build2 publishers3 min readPublished Updated

Redis creator's ds4 engine squeezes a handful of open models into 96GB to 128GB of memory

Salvatore Sanfilippo's open-source ds4 engine runs a short list of large open-weight models locally by compressing their routed experts to about two bits. Even compressed, the supported builds need high-memory Macs or GPU systems that most people do not already own.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Redis creator's ds4 engine squeezes a handful of open models into 96GB to 128GB of memory
Generated illustration

What happened

  • Sanfilippo published the ds4 repository on May 7, 2026, according to independent coverage by RunLocal.
  • Supported models are limited to some DeepSeek V4 and V4.1 models plus GLM 5.x and Qwen3.8 Flash Next; layouts and capabilities differ from one backend to another.
  • ds4 writes prompt-prefix cache state to SSD and restores a matching prefix after a restart instead of recomputing it.
  • A single model state backs three interfaces: a command-line tool, a native agent, and an HTTP server whose endpoints are compatible with OpenAI and Anthropic APIs.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Most teams would have to buy hardware before they could evaluate ds4 at all, so its audience is people who already own high-memory machines.
  • decision Anyone needing a model outside the supported list has to choose between a broader runtime and porting ds4, with Sanfilippo betting coding agents make the port cheap.
  • capability Agent tooling written for OpenAI or Anthropic APIs can drive a local model unchanged, so a long coding session can move off a hosted service.
  • constraint Buyers comparing ds4 with other engines still need their own runs, because the published speeds come from the project and share no common baseline.

The two-bit cut lands on routed experts because they hold much of a large mixture-of-experts model's parameter count [7][18]. Shared components, and other paths the project identifies as critical, keep higher precision [7]. Choosing those paths is a model-specific decision, and the narrow support list is what lets Sanfilippo make it [21]. Sanfilippo, who created Redis, built DwarfStar around that limit: support fewer models, then tune the whole stack for them [1].

The project says it tests the model, prompt handling, tool calls, cache and interfaces together [6]. I think that joint test is the best engineering decision in the repo. A compression scheme is worth little if tool calls or the cache misbehave on top of it, and testing them as one system is how that gets caught. Users pay for that rigour: they cannot assume an arbitrary model file will work [6].

The benchmark table is the project's own, and Runtimewire's coverage says its figures are not independent or apples-to-apples comparisons against other engines [14]. Inside the table, the M5 Max loses about 30% of its generation rate when context grows 32-fold [1]. A 128GB DGX Spark posts 18.1 and 13.8 tokens per second at the same two context lengths, a drop of about 24% [13][2]. The Mac is roughly twice as fast at both [3].

For those figures to transfer, a team needs the same supported build, a machine of the same memory class and a workload dominated by generation. Prompt prefill is reported separately, since ingesting input and producing tokens are different parts of a workload [15]. Runtimewire's write-up does not reproduce the prefill rates or say which model produced the rows. A coding agent that resends a long prompt every turn spends time in prefill that a generation rate does not count. The SSD cache restores only a matching prefix, so its saving depends on how much of each new prompt matches the stored state [9].

Memory sets the floor. Documentation puts supported builds at roughly 96GB to 128GB [8]. DwarfStar's website lists Apple Silicon machines with at least 64GB for some supported configurations, while the baseline DeepSeek V4 Flash Q2 setup targets higher-memory systems [16]. Streaming from SSD extends what fits but does not remove the capacity trade-off, according to the coverage [17]. Both benchmark machines have 128GB [12][13]. By the coverage's own account, the approach needs hardware well above what most people already have [20].

Sanfilippo's reason for accepting that floor is ownership. "AI is too critical to be just a provided service," he wrote about ds4 [4]. In an essay on software distribution, he argued that a repository can be a template that users and their agents adapt to other hardware and needs [11]. He points to DwarfStar itself as an example [11]. In my view the narrow scope is the right call for an engine built around a handful of targets. An agent can port a backend, though it cannot add memory to the machine it tests on. The same write-up treats the agent argument as a wager on how developers will use open-source software, and says it does not mean ds4 already supports every machine or model [19].

What to watch

  • Independent runs of ds4 against other engines on the same 128GB machines, with prefill rates published alongside generation.
  • Agent-made or community ports that extend ds4 to models or devices outside its supported list, which would test Sanfilippo's template argument.
  • Whether SSD streaming brings any supported DeepSeek V4 build below the 64GB Apple Silicon floor.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories