Inside Kog’s Quest to Make AI Inference 30x Faster Using Software Alone.
Kog Wants to Prove Your GPU Has More Speed Left in the Tank:
The French startup is betting that software, not new silicon, is the fastest path to faster AI inference — and what that means for the future of enterprise AI performance.
30x: faster LLM inference Kog is targeting
3,000 TPS: achieved in its 2B-parameter demo model
200 tangible business leads generated after launch
1: The Chips You Already Own May Have More to Give:
While Cerebras rode a warm IPO welcome for its purpose-built AI chips in May, French startup Kog is chasing a different bet entirely: that the standard GPUs enterprises already own have far more inference speed locked inside them than anyone is using.
Kog made its case with a tech preview aimed at proving that extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own — AMD MI300X and Nvidia H200 chips, the kind already sitting in production environments, not exotic new hardware. The preview hit the front page of Hacker News in May.
Some readers were disappointed the approach didn't extend to consumer laptop GPUs. Others saw the bigger opportunity: with inference speed and cost now a genuine bottleneck for AI deployment, a software-only route to unlocking existing hardware got real attention. Founder and CEO Gaël Delalleau told TechCrunch the launch generated 200 tangible business leads.
2: Where the Demand Is Actually Coming From:
The first real use case taking shape is software engineering — a field acutely familiar with AI's speed problem.
Veteran Claude Code users know the wait: results that can take hours to return. Anthropic itself prices around this reality, charging a premium for Claude's Fast Mode. Kog is targeting exactly those users — professionals whose AI workflows are held up by latency — alongside design partners building prompt-to-game and prompt-to-app tools, where faster generation translates directly into more revenue.
But Kog also learned something about the market's maturity along the way: prospective customers aren't prepared to fine-tune small models themselves. That finding reshaped the roadmap. Since launch, the company has been fully focused on accelerating larger models rather than smaller, easier-to-optimize ones — meeting the demand it actually saw, not the demand it expected.
3: The Gap Between the Demo and the Promise:
Kog's headline claim is 30x faster LLM inference. Its demo, so far, has proven something narrower.
The showcased result was 3,000 per-request tokens per second — an impressive number, but generated by Laneformer 2B, a purpose-built, now open-sourced small model with roughly 2 billion parameters. Scaling that kind of speed to genuine large language models, the ones enterprises actually deploy, is a considerably bigger technical leap.
"GPUs have a bright future."— Gaël Delalleau, CEO of Kog, on the idea that GPUs are poorly suited for LLM decoding
Delalleau pushes back on the idea that GPUs are inherently limited for decoding work, pointing to newer chips' growing memory bandwidth as untapped headroom. Kog isn't alone in the software-optimization camp — fellow French startup ZML has released hardware-agnostic software that bypasses Nvidia's CUDA entirely to support fast inference across competing chips. Delalleau positions Kog's approach as closer to Stanford lab Hazy Research: an even deeper, lower-level focus on squeezing performance directly out of the GPU itself.
4: A Hacker's Approach to Hardware:
Delalleau's path to Kog runs through physics and offensive cybersecurity, not conventional AI research.
He studied solid-state physics at France's École Polytechnique before moving into white hat hacking, becoming a four-time finalist at DEFCON's CTF tournament. That combination shapes how he describes Kog's engineering culture: a physics mindset focused on understanding the underlying laws of the GPU, paired with a hacker's instinct for reverse-engineering systems down to assembly language and binary code — then repurposing them for goals they weren't originally designed for.
The trade-off is that this approach is slow and hands-on. Every new GPU generation requires several weeks or months of dedicated, low-level research before Kog can support it. With a team of just 11 people, that inherently limits how many chips the company can realistically cover in the near term.
5: What's Next — and the Bigger Picture:
Kog's longer-term plan is to feed its GPU methodology into agent-based pipelines that can scale support across more chips and models automatically, rather than relying purely on manual, months-long engineering per chip.
There's also a sovereignty angle: as Europe pushes to build independent AI infrastructure and model capability, Kog's homegrown approach — already backed by Scaleway, France's Bpifrance, and the French Tech 2030 program — fits a broader national push. But for now, the company's near-term milestone is proving the approach works on real large language models, not just small demo models.
Delalleau expects to hit 10x speed on a major model by September, a result he's counting on to demonstrate customer traction and unlock a Series A raise.
Speed Is a Feature — Make Sure Your AI Has It.
Kog's bet is that most enterprises are sitting on more AI performance than they're using — the hardware just needs the right layer of intelligence on top of it. That's exactly the gap Agent+ closes for businesses that don't have an 11-person GPU research team on staff. Agent+, Otherworlds AI's enterprise platform, deploys purpose-built agents — powered by Google Opal automation — that are tuned to run fast and lean from day one, starting at $297/month.
No months of low-level engineering required, no waiting hours for results: just AI that performs the way your business actually needs it to.
See how Agent+ or a custom enterprise build can speed up your AI workflows at otherworldsai.com.








