French Startup Kog Aims for Enhanced AI Performance Through Conventional GPUs
The quest for rapid AI inference intensifies as French startup Kog emerges, advocating for the untapped potential of standard GPUs over newly launched chips like those from Cerebras, which made a notable debut through its IPO in May. Kog gained attention on Hacker News with a tech preview confirming that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own,” using AMD MI300X and Nvidia H200 GPUs for demonstrations.
While some users expressed disappointment that this approach does not apply to laptop GPUs, many saw promise in Kog’s software optimizations aimed at enhancing existing hardware. CEO Gaël Delalleau stated, “We had 200 tangible business leads,” highlighting significant interest in overcoming current limitations of inference speed and cost.
Kog has identified software engineering as the primary use case based on early feedback. Targeting users who face lengthy wait times in AI workflows, the startup plans to assist design partners in creating applications and games more swiftly through its Kog Inference Engine (KIE), enhancing revenue opportunities.
Recognizing the immaturity of the market, Kog has shifted focus from small model fine-tuning to developing larger models to align with observed demand. The startup aims to achieve “30x faster LLM inference” as demonstrated by its current capabilities, achieving 3,000 per-request tokens per second with a smaller model, the Laneformer 2B, now open-sourced.
Delalleau remains confident despite skepticism, asserting that the same techniques will be effective for larger models, stating, “GPUs have a bright future.” He insists misconceptions about GPUs’ unsuitability for decoding are unfounded, emphasizing increased memory bandwidth in newer models waiting to be utilized.
While Kog is not alone in its belief in software optimization’s potential, with other firms like ZML offering solutions that bypass Nvidia’s CUDA, Delalleau argues that Kog shares a more profound focus on GPU acceleration akin to Stanford’s Hazy Research lab.
Delalleau’s background, having studied solid-state physics at École Polytechnique and working in offensive cybersecurity, informs Kog’s methodology. He encourages his team to understand both the scientific laws governing GPUs and their practical applications through reverse-engineering techniques learned during his experiences at competitions like DEFCON.
This hands-on, detail-oriented method requires extensive research on new GPUs, making scalability a challenge for Kog’s small team of 11. Looking ahead, the startup aims to integrate its methodologies into agent-based pipelines to support a greater range of chips and models, in tune with Europe’s push for increased technological sovereignty.
Proving its approach’s effectiveness on LLMs is critical for Kog’s future, particularly in attracting further investment. Delalleau anticipates demonstrating the performance of a major model at 10x speed by September, which he sees as essential for establishing customer traction and advancing to a Series A funding round.


