Skip to content

Build1 publisher3 min readPublished

Cognichip's one detailed ACI example divides out to 40 to 50 times faster

Faraj Aalaei's Cognichip argues that semiconductor physics belongs inside the model, with agents left to orchestrate. The one speedup it has documented in detail is a 55-page spec handled in days against a baseline Cognichip supplied itself.

The Engineer · Build desk

Illustration accompanying Cognichip's one detailed ACI example divides out to 40 to 50 times faster

What happened

  • Forbes reported on September 14th that Cognichip's Artificial Chip Intelligence spans microarchitecture, RTL design, functional verification, power-performance-area optimization and FPGA implementation.
  • One engineer processed a 55-page specification in a few days, against the roughly four to five months Cognichip says comparable work would normally require, covering microarchitecture through PPA optimization.
  • Forbes reports early results pointing to a roughly 100x chip-design speedup, while the figures Cognichip has published do not yet share one consistent benchmark.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision A design manager staffing a project has to pick a number to plan against, and the only worked example supports roughly 41 to 51 times, not 100.
  • constraint Because the training set was curated in-house and topped up with synthetic examples, a buyer cannot check ACI against a shared public corpus; the comparison has to run on the buyer's own designs.
  • exposure The customer-controlled deployment option moves the diligence question to where weights run and what telemetry leaves the firewall, since the design files are the asset being protected.
  • cost Design-side speed is cheaper to demonstrate than it is to bank: application code can be patched after deployment, and silicon carries a higher proof burden, so the saving only counts after verification signs off.

Read "a few days" as three days, which is generous to the claim. Four months is about 122 days and five months about 152, so the 55-page specification example divides out to between 41 and 51 times faster [17][22]. Forbes reports that early ACI results point to a roughly 100x chip-design speedup [14]. The report does not provide the methodology, baseline staffing model or independent replication that would reconcile the example with that headline figure [18].

ACI Enterprise includes an agent framework for orchestrating tools and tasks [10], so the claim is about where the domain knowledge sits. Cognichip says its underlying models already understand hardware abstractions and physical constraints, and that the agents coordinate work instead of compensating for a general-purpose model that treats RTL as another programming language [11]. I would expect that difference to show up as refusals. A model holding timing and power constraints declines a design it cannot close, while a prompted coding model emits valid RTL and leaves synthesis to find the problem.

The reason to build it that way is the data. Semiconductor businesses closely guard design files, verification results and manufacturing constraints, and the public corpus is thin next to the software repositories available to train coding models [7]. The company paired experienced chip designers with AI researchers: the designers produced and curated domain data, and the researchers generated synthetic examples and built models grounded in logic, timing, power and physical constraints [8]. ACI also supports customer-controlled data and deployment configurations intended to keep sensitive intellectual property inside customer firewalls [9].

For 100x to transfer, the denominator has to describe your team. The four-to-five-month figure is Cognichip's own estimate of what comparable work would normally require [17], and the published figures do not yet share one consistent benchmark [16]. Manoher Bommena, a Renesas vice president of engineering, said the system "simply 'speaks chip.'" [15] Cognichip's case rests on its own measurements and on customer testimony of that kind [20].

The FPGA work has the shortest feedback loop. Forbes says ACI can move an FPGA concept to working code in a few hours [13], and Cognichip is extending ACI into FPGAs for robotics, automotive systems and industrial equipment, where shorter implementation cycles could matter [12]. A production tape-out would provide a harder test [23]. In April, TechCrunch reported that Cognichip's evidence centered on collaborations and demonstrations rather than an identified chip designed through ACI [19].

Aalaei founded and led Centillium Communications and Aquantia through public listings, then oversaw Marvell's networking and automotive segment after Marvell acquired Aquantia [4]. He founded Cognichip in 2024, after advances in generative AI made a different design workflow plausible [5]. His CTO, Ehsan Kamalinejad, holds a Ph.D. in applied mathematics from the University of Toronto and worked on machine learning at Apple and Amazon; chief architect Simon Sabato held chip and systems roles at Google, Cisco and Cadence [6].

What to watch

  • A named production tape-out of a chip designed through ACI, with the customer identified.
  • A published methodology behind the roughly 100x figure: baseline headcount, spec complexity and verification depth.
  • Whether the AI Infra Summit sessions on September 15th to 17th include benchmark definitions rather than more demonstrations.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories