Skip to content

Product1 publisher3 min readPublished

Cornelis wants the network to finish the AllReduce before data reaches the GPU

Cornelis Networks launched a fabric that performs reductions and gradient compression in transit, and the only number attached to it, a traffic cut of up to 50%, comes from the company's own pre-production simulations.

The Product Desk · Product desk

Illustration accompanying Cornelis wants the network to finish the AllReduce before data reaches the GPU

What happened

  • Cornelis Networks launched its Active Compute Fabric at the AI Infra Summit, an open architecture that puts programmable compute inside the network for scale-up and scale-out AI clusters.
  • The launch came with a $205 million round led by IAG Capital Partners and a strategy collaboration with Qualcomm Technologies.
  • Chief marketing officer Brandon Draeger said the fabric assembles KV cache data for disaggregated inference and coordinates expert dispatch for mixture-of-experts models as data crosses the network.
  • Draeger put the reduction in overall network traffic at up to 50%, drawn from Cornelis's own pre-production simulations.
  • Cornelis says the fabric runs on Ethernet and UALink for scale-up and Ultra Ethernet for scale-out, so buyers can keep their existing compute architectures.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Anyone sizing a fabric refresh now sets an evidence bar: accept a vendor simulation as grounds for a cluster-scale purchase, or hold the order until a customer publishes measured job times.
  • constraint By Cornelis's own account the gain varies with model, cluster size and stack, so finance is handed a range with no floor and no fixed number to put in a payback case.
  • capability In-path reduction frees accelerator cycles that combining results at both ends currently consumes, and that is capacity the buyer already paid for.
  • precedent Qualcomm appearing on stage behind an open scale-up and scale-out fabric sets the reference other accelerator vendors get compared against.

The operator this lands on already owns the accelerators. Their question is narrower than the launch: does moving reductions into the switch change job completion time on the cluster they run this quarter, and can they show it before the purchase order goes in.

Brandon Draeger, Cornelis's chief marketing officer, described the fabric doing work in the path the data was already taking. "The payload does not arrive the way it left," Draeger said. "In a collective operation, partial results from many endpoints are combined inside the fabric, so a single reduced result lands at the destination instead of thousands of separate contributions. Compressed gradients move as a fraction of their original size. The work happens once, in the path the data was already taking, instead of consuming accelerator cycles at both ends." [9]

"Accelerator utilization in large AI deployments commonly sits near half of installed capacity, and the architecture is designed to return a meaningful share of what is currently wasted," Draeger said [11]. On a 2,048-accelerator cluster at that utilization, roughly 1,024 accelerators' worth of capacity is unused at any moment [16]. A meaningful share of that could be five accelerators or five hundred. Draeger put the variability plainly: "How much depends on the model, the cluster size, and the customer's stack, so improvements will vary by workload and deployment" [12]. SiliconANGLE did not report pricing or availability, and Cornelis put no number on the utilization it expects to give back [17].

Halving the bytes on the wire cannot buy back more time than the jobs currently lose waiting for them. Cornelis's own framing of the problem is that expensive accelerators sit idle waiting for data [4], so the figure a buyer needs is the share of step time their jobs spend blocked on collectives and synchronisation, measured on their stack, not on a simulator.

Chief executive Lisa Spelman argued the network has fallen behind the rest of the stack, saying compute, memory and storage have grown more workload-aware over the past couple of years while traditional networks have not [5]. "AI infrastructure is reaching a point where faster endpoints alone are not enough," she said. "The fabric has to become an active part of the compute system" [13]. The hedge against stranding a purchase is the standards list: Ethernet and UALink for scale-up, Ultra Ethernet for scale-out [6]. Qualcomm shares that open scale-up and scale-out position and will join Cornelis on stage [15].

SiliconANGLE frames Cornelis as a rival to Cisco and Arista [3]. The comparison Cornelis itself makes is with traditional networks in general terms [5]. For a Monday decision, read the fraction of step time your training runs lose to communication, and check whether your collective library can call an in-fabric offload without an application rewrite. Cornelis quantified traffic; the number that decides the order is step time, and you can pull that one out of your own profiler.

What to watch

  • A named customer publishing measured step-time or job-completion numbers on Active Compute Fabric, not simulation results.
  • Whether the Qualcomm collaboration becomes a design win or product commitment with disclosed terms.
  • Pricing, availability, and which collective libraries can call the in-fabric offload without an application rewrite.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories