Build2 publishers3 min readPublished
Delos Data raises more than $100M to keep inference running when a link drops
Delos Data's Nonstop AI Reference Architecture is meant to hold a mixed domain of GPUs, CPUs, memory and storage together through hardware and link failures. The 10x figures attached to it describe silicon the company still calls planned.
The Engineer · Build desk

What happened
- Delos Data of Palo Alto says it has raised more than $100 million, and it launched its Nonstop AI Reference Architecture on September 15th with the stated goal of keeping inference running through hardware and link failures.
- The announced set is Nonstop AI Clusters, a disaggregated Nonstop AI Server that mixes processors and switches, the Nonstop AI Data Interface and Mosaic software.
- The Data Interface, the piece meant to link compute, memory and storage, is planned silicon, and Delos Data credits it with 10x lower latency and 10x higher efficiency.
- Named investors include Matrix, Playground Global, Socratic Partners, Capricorn's Technology Impact Fund, Matter Venture Partners and IAG, plus unnamed backers from the computing and networking industries.
- A May 27th, 2026 release described 1,000-GPU scale-up domains as practical and 10,000 GPUs as potentially possible. No independent party has validated either figure.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision A buyer weighing the Data Interface is pricing a design target for a part not yet built, so the evaluation has to wait on delivery of a benchmark with a stated baseline.
- constraint Delos Data does not sell the accelerators or the models, so its revenue depends on customers agreeing to buy a separate data path between parts they already source elsewhere.
- exposure With the press release as the sole source for the money and the product, anyone repeating the 10x figures or the 300x demand projection is passing along a company number that no test configuration or named author stands behind.
A stateless request that meets a broken link costs one retry. Delos Data defines agentic inference as persistent workloads that run across multiple devices and do not stop after one request [4]. Those workloads hold state in the memory of the devices they occupy. A failed device or a failed link takes part of that state with it, so keeping the job alive means the surviving devices can still reach or reconstruct what the dead one held [3].
The 10x latency figure describes a part Delos Data calls planned silicon [14]. That is a design target, and design targets do not automatically transfer. For the number to say anything about a buyer's cluster, the baseline has to be named: latency of which path, measured between which endpoints, at which message size, against which interface in place today. Efficiency needs a unit too. Delos Data's press release is the primary source for the financing and the product claims. It gives no test configuration, no independent benchmark and no named production customer, and it does not say how much was raised, how the round was structured, what the company is worth or who led it [8][15].
The stated benefits of the cluster and server designs are reduced cost per token, improved tokens per watt and increased tokens per second [17]. All three move with the model, the batch size and the sequence length. Comparing them across two vendors means running the same workload on both.
The financing announcement cites an unattributed industry projection of roughly 300x growth in token demand by 2030, with agentic inference accounting for most of the increase [5]. Take 2026 as the baseline and that works out to about 4.2x a year compounded, since 300 to the one-fourth power is 4.16 [20].
"The most expensive idle asset in a data center is a GPU, CPU or an accelerator waiting on the network," Doe said in the September 15th announcement [9]. Doe ran Barefoot Networks, which Intel acquired in 2019, and Barefoot's programmable Tofino Ethernet switch became part of Intel's networking portfolio [10]. Daly came into Intel through its 2011 acquisition of Fulcrum Microsystems, where he was technical director of software and systems and worked on software control planes for high-speed, low-latency switch hardware, according to an Open Networking Foundation biography [11]. Control-plane software for low-latency switching sits close to the problem of holding a distributed job together when a link goes away, and both founders have spent much of their careers on systems that move data between processors [12].
Wen Hsieh, founding managing partner at Matter Venture Partners, described Delos Data as having built software, followed by an AI server [19]. The capital goes to expanding software and hardware engineering teams and to product development and sales [7]. Until the Data Interface silicon exists, the testable claim is the software and cluster architecture: whether a mixed domain of accelerators, memory and storage keeps a persistent inference job running when a device or a link fails [13][21].
What to watch
- A published test configuration and baseline for the Data Interface's 10x latency and efficiency figures.
- A named production customer running a persistent inference job through a device or link failure on Nonstop AI Clusters.
- Any demonstration of the 1,000-GPU scale-up domain by a party other than Delos Data.