Product1 publisher3 min readPublished
Kog's pitch: the cheapest inference upgrade is the H200s you already bought
The French startup's only public number is 3,000 tokens per second on a 2-billion-parameter model. Its 30x claim for real LLMs has not been shown yet.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Kog is a French startup betting that there is a lot more power to be squeezed out of conventional GPUs through software.
- Kog published a tech preview in May aimed at proving that extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own; the preview reached the front page of Hacker News.
- The GPUs used for Kog's demo were the AMD MI300X and the Nvidia H200.
- According to TechCrunch, markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May.
- Kog CEO Gael Delalleau said newer GPUs have more and more memory bandwidth that only begs to be unlocked, and that GPUs have a bright future.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
French startup Kog is arguing that the next real drop in inference cost comes from low-level software on the datacenter GPUs enterprises already own rather than from purpose-built silicon, and it ran its May tech preview on AMD MI300X and Nvidia H200 cards to make the point [1][2][3]. For anyone holding an unsigned accelerator purchase order, that is a claim with budget consequences, and it is testable.
The context is that markets welcomed Cerebras and its purpose-built chips at its IPO debut in May, according to TechCrunch [4]. Kog's counter-position is that newer GPUs keep gaining memory bandwidth that is not being used, and that the idea GPUs are poorly suited to decoding has become a misconception, in CEO Gael Delalleau's words [5][6]. He told TechCrunch the preview produced 200 tangible business leads [7].
Now the arithmetic. The demo showed 3,000 per-request tokens per second, but on a purpose-built model of roughly 2 billion parameters, since open sourced as Laneformer 2B [8][9]. The company's headline promise is 30x faster LLM inference [10]. Those are not the same result: the published figure comes from a small custom model, and no equivalent number for large models has been shown [11]. Delalleau says he is confident the approach carries over to LLMs [12]. Kog has also learned that prospective customers are not prepared to fine-tune small models, and has since focused on accelerating larger ones [13]. That is the honest reading of the state of play: the demand is for the thing not yet demonstrated.
Note also what is being measured. The preview was framed around extremely fast single-request decoding, and the metric quoted is per-request throughput [2][8]. That is a latency claim, which fits the customers Kog says it is chasing first: software engineering, where Claude Code users sometimes wait hours for results, and where Anthropic charges a price multiple for Fast Mode [14][15][16]. Design partners generating games and apps from a prompt would convert speed into revenue [17]. Latency per request and cost per token are different lines on a bill, and a buyer should ask which one moves.
The method explains both the appeal and the ceiling. Delalleau studied solid-state physics at Ecole Polytechnique, worked in offensive cybersecurity, and was a four-time DEFCON CTF finalist, and he describes the work as reverse-engineering down to assembly and binary to use hardware for purposes it was not designed for [18][19]. That is hands-on and slow: several weeks or even months of GPU engineering research per new chip, with a team of 11, which he says limits how many chips Kog can support for now [20][21]. The longer-term plan is to push the methodology into agent-based pipelines to cover more chips and models [22]. Kog is not alone in the general thesis; fellow French company ZML shipped hardware-agnostic software that bypasses CUDA, while Delalleau positions Kog closer to Stanford's Hazy Research lab, only deeper [23][24]. The seed round was co-led by Varsity VC, whose Kamel Zeroual was Delalleau's co-founder at his 2009-vintage startup Stribe, and Kog has support from Scaleway and backing from Bpifrance [25][26][27].
What to watch: a per-request number on a large open-weight model on an unmodified H200 or MI300X, the list of GPUs officially supported as headcount stays near 11, and whether any of those 200 leads becomes a named production deployment.