Product1 distinct publisher3 min readUpdated
The French startup's only public number is 3,000 tokens per second on a 2-billion-parameter model. Its 30x claim for real LLMs has not been shown yet.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
French startup Kog is arguing that the next real drop in inference cost comes from low-level software on the datacenter GPUs enterprises already own rather than from purpose-built silicon, and it ran its May tech preview on AMD MI300X and Nvidia H200 cards to make the point [1][2][3]. For anyone holding an unsigned accelerator purchase order, that is a claim with budget consequences, and it is testable.
The context is that markets welcomed Cerebras and its purpose-built chips at its IPO debut in May, according to TechCrunch [4]. Kog's counter-position is that newer GPUs keep gaining memory bandwidth that is not being used, and that the idea GPUs are poorly suited to decoding has become a misconception, in CEO Gael Delalleau's words [5][6]. He told TechCrunch the preview produced 200 tangible business leads [7].
Now the arithmetic. The demo showed 3,000 per-request tokens per second, but on a purpose-built model of roughly 2 billion parameters, since open sourced as Laneformer 2B [8][9]. The company's headline promise is 30x faster LLM inference [10]. Those are not the same result: the published figure comes from a small custom model, and no equivalent number for large models has been shown [11]. Delalleau says he is confident the approach carries over to LLMs [12]. Kog has also learned that prospective customers are not prepared to fine-tune small models, and has since focused on accelerating larger ones [13]. That is the honest reading of the state of play: the demand is for the thing not yet demonstrated.
Note also what is being measured. The preview was framed around extremely fast single-request decoding, and the metric quoted is per-request throughput [2][8]. That is a latency claim, which fits the customers Kog says it is chasing first: software engineering, where Claude Code users sometimes wait hours for results, and where Anthropic charges a price multiple for Fast Mode [14][15][16]. Design partners generating games and apps from a prompt would convert speed into revenue [17]. Latency per request and cost per token are different lines on a bill, and a buyer should ask which one moves.
The method explains both the appeal and the ceiling. Delalleau studied solid-state physics at Ecole Polytechnique, worked in offensive cybersecurity, and was a four-time DEFCON CTF finalist, and he describes the work as reverse-engineering down to assembly and binary to use hardware for purposes it was not designed for [18][19]. That is hands-on and slow: several weeks or even months of GPU engineering research per new chip, with a team of 11, which he says limits how many chips Kog can support for now [20][21]. The longer-term plan is to push the methodology into agent-based pipelines to cover more chips and models [22]. Kog is not alone in the general thesis; fellow French company ZML shipped hardware-agnostic software that bypasses CUDA, while Delalleau positions Kog closer to Stanford's Hazy Research lab, only deeper [23][24]. The seed round was co-led by Varsity VC, whose Kamel Zeroual was Delalleau's co-founder at his 2009-vintage startup Stribe, and Kog has support from Scaleway and backing from Bpifrance [25][26][27].
What to watch: a per-request number on a large open-weight model on an unmodified H200 or MI300X, the list of GPUs officially supported as headcount stays near 11, and whether any of those 200 leads becomes a named production deployment.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Kog's stated promise is "30x faster LLM inference", which TechCrunch describes as leaving the company a huge leap to make.
Kog is a French startup betting that there is a lot more power to be squeezed out of conventional GPUs through software.
Kog published a tech preview in May aimed at proving that extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own; the preview reached the front page of Hacker News.
The GPUs used for Kog's demo were the AMD MI300X and the Nvidia H200.
According to TechCrunch, markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May.
Kog CEO Gael Delalleau said newer GPUs have more and more memory bandwidth that only begs to be unlocked, and that GPUs have a bright future.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Thin: one publisher, one vendor-run benchmark on a bespoke 2B model
All substantive detail comes from a single TechCrunch interview with the founder. The only quantitative result is a self-reported 3,000 tokens/second single-request decode on a purpose-built ~2B model on MI300X/H200, with no stated baseline, no independent reproduction and no LLM-scale measurement. Partial mitigation: the demo model (Laneformer 2B) is open sourced and the preview was publicly scrutinised on Hacker News, and the publisher itself flags the gap to the 30x claim.
Pre-revenue: leads and unnamed design partners, no disclosed deployments
Disclosed traction is 200 self-reported business leads plus unnamed design partners; no customer names, contract values, production deployments or revenue appear in the source. Kog also reports that prospective customers will not fine-tune small models, which forced a pivoted roadmap toward larger models still under development. An 11-person team with weeks-to-months of engineering per GPU further caps near-term reach.
Marketed 30x outruns the demonstrated 2B-model result
The company markets '30x faster LLM inference' while the only public number is a single-request decode figure on a 2-billion-parameter purpose-built model, and the founder's own near-term milestone is 10x on a first major model expected later. That is a large distance between claim and shown result, which the publisher acknowledges rather than obscures — hence a clearly positive but not extreme gap.
Founder-sourced narrative tied to an upcoming Series A
Nearly every claim originates with the CEO of a pre-Series A startup that explicitly links its next raise to demonstrating a first major model at 10x speed, giving strong incentive to present preview results favourably. The seed round was co-led by the firm of the founder's former co-founder, and public backing (Scaleway, Bpifrance, French Tech 2030) plus a European sovereignty framing add further promotional pull. Countervailing: the publisher discloses these relationships and the demo's small-model caveat.
Moderate-low: internally consistent but unverified and single-sourced
The account is detailed, self-consistent and transparent about its main weakness (a 2B-model demo behind a 30x claim), which supports confidence in what was said. But there is no second publisher, no independent benchmark, no named customer and no disclosed round size, so confidence in the underlying performance and traction claims stays low.
build
Developer habit, priced at $965B: what Anthropic's run actually proves1 distinct publisher
build
OpenAI puts latency on the price list: 750 tokens/sec, gated by workload fit3 distinct publishers
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 14, 2026