Skip to content

Build1 publisher3 min readPublished

Dnotitia's retrieval ASIC needs 1.73x its FPGA prototype to reach a 10x target

The first VDPU samples are back from fab and go into system evaluation in the fourth quarter. Everything Dnotitia has published so far, including the 5.77x throughput claim, was measured on a four-card FPGA rig.

The Engineer · Build desk

Photograph accompanying Dnotitia's retrieval ASIC needs 1.73x its FPGA prototype to reach a 10x target
Photo: yahoo.com

What happened

  • Dnotitia said on September 18th that the first samples of its VDPU vector-search chip had come back from fabrication and that it has begun characterizing the silicon.
  • The company plans to put ASIC-based systems into evaluation during the fourth quarter.
  • The 5.77x vector-search throughput figure Dnotitia has been quoting came from a four-card FPGA evaluation system, not from the chip that returned from the fab.
  • Dnotitia showed a four-card VDPU server at the AI Infra Summit in Santa Clara, California, held from September 15th through 17th.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Sizing a rack or a budget takes more than a throughput ratio: latency at a given recall, board power and card price only exist after the ASIC is tested.
  • decision Because the FPGA ran FAISS, Milvus and hnswlib, a Q4 trial is a card evaluation inside a stack the team already runs, so the buying decision does not start with a database migration.
  • cost If the index-build savings survive the move to silicon, the host capacity an infrastructure owner has to reserve for rebuilds gets smaller. Reserved host capacity is the line item the card would pay for.
  • exposure Dnotitia produced both the methodology and the numbers, so the only measurements a buyer weighing the card in Q4 can check are Dnotitia's own.

A card that owns retrieval has to own the index as well as the distance math. Dnotitia's FPGA build supported brute-force KNN, IVF, NSW and HNSW indexes under FAISS, Milvus and hnswlib [6]. The silicon has to serve those families. An FPGA can be reprogrammed when a fifth one arrives. A tapeout cannot, so a new index family becomes host-side software running on the cores the card was bought to free. Dnotitia says it plans to add further libraries and databases when the ASIC enters evaluation [18].

The 5.77x ratio is a claim about somebody else's server. The baseline was a dual-socket, CPU-only machine running the same software stack, and Dnotitia says the four-card FPGA system matched or improved on its recall [3][4]. Recall parity is the condition that makes a throughput ratio mean anything, and Dnotitia states it up front. The measured workload was 4,096-dimensional and multimodal [5]. Narrower embeddings move less data per comparison, and the throughput advantage shrinks with them. The comparison is against CPU-only retrieval [3], so a team already pushing index work onto spare GPU memory is asking a different question. Dnotitia supplied the benchmark methodology and the results [7].

Dnotitia is targeting as much as a 10x improvement over a CPU-based server once the ASIC is installed in a system, and calls that figure a design target [9]. Divide it by what the prototype produced and the production silicon has to run about 1.73 times faster than the FPGA rig [19]. Fourth-quarter testing is also where the numbers a purchase order needs come from: latency at useful recall levels, sustained query throughput, board power, system cost and performance on datasets that resemble production traffic [8].

The stated job of the card is to keep host processors and GPU memory available for application code and model execution [13]. Index construction is when that is hardest to honour, because a build wants the same cores the model server is using. Dnotitia's index-build figures are the part of the FPGA data that speaks to that, and they are the part most exposed to how the silicon behaves under sustained load. The VDPU sits alongside Seahorse, Dnotitia's vector database and AI storage platform, which lets the company tune the database for its own accelerator and sell a combined system [17].

First samples give Dnotitia working material to characterize, and nothing yet about yield, production readiness or economics [11]. A roughly $63M Series A in April 2026, led by Elohim Partners, pays for the rest of the cycle [16]. Moo-Kyoung Chung founded the company in 2023 after serving as chief technology officer of SAPEON, SK Telecom's former AI chip subsidiary, according to The Elec [14]. CTO Se-Hyun Yang, who leads the hardware side, has spent more than 20 years designing processors, including work at Samsung Electronics on mobile CPUs and GPUs and on AI processors for data centres and supercomputers [15].

What to watch

  • Whether Dnotitia publishes latency at a stated recall target, on top of the throughput ratios, once ASIC systems enter evaluation.
  • Which libraries and databases beyond FAISS, Milvus and hnswlib arrive with the ASIC evaluation.
  • Any measurement of the card by someone other than Dnotitia, which supplied the current benchmark methodology.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories