VIDRAFT · Korean Pre-AGI AI startup · 2026-08-05

VIDRAFT Claims World No. 1 in Google's Fast Gemma Challenge

A Korean AI startup just proved that inference efficiency — not model size — is the new competitive frontier.

TL;DR: VIDRAFT, a Korean AI startup led by CEO Kim Min-sik, has claimed the top verified position in The Fast Gemma Challenge, a global inference optimization competition co-hosted by Google's Gemma team and Hugging Face. The company's autonomous AI agent vidraft-darwin achieved a verified throughput of 510.58 tokens per second (TPS) with a perplexity score of 2.39 — roughly six times faster than the baseline inference speed of the original model. Results are publicly verifiable on the official Fast Gemma Challenge leaderboard.

VIDRAFT, the Korean AI startup headquartered in Seoul, announced on August 2, 2026, that it has taken the top verified spot in The Fast Gemma Challenge, a globally competitive AI inference optimization contest organized jointly by Google's Gemma team and Hugging Face. The achievement marks a significant milestone not only for the company but also for the broader conversation about what true AI competitiveness looks like in a hardware-constrained world.

What VIDRAFT Announced

The Fast Gemma Challenge tasks participating developers worldwide with a deceptively simple goal: take the same model, run it on the same hardware, and see who can make it run the fastest — without sacrificing output quality. Every participant receives Google's open multimodal model Gemma-4-E4B-it and an identical NVIDIA A10G GPU. The rules explicitly prohibit swapping out the model or disabling any of its text, image, or audio capabilities. Only software-level optimization skill counts.

What makes the competition especially credible is its verification mechanism. Rankings are not determined by self-reported scores. Instead, the organizers independently re-run each submission using undisclosed test prompts and measure speed and quality themselves. Because only scores that the judges can independently reproduce are accepted, there is no room for inflated or gamed results.

Against that rigorous backdrop, VIDRAFT's autonomous AI agent, vidraft-darwin, posted a verified throughput of 510.58 tokens per second and a perplexity (PPL) score of 2.3929 — securing the highest verified performance among all entries that maintained acceptable output quality. The throughput figure represents approximately six times the baseline inference speed of the original unoptimized model.

VIDRAFT attributes the result to optimization know-how accumulated through two internal technology pillars: its proprietary VK-series inference engine and POCKET, an on-device AI platform designed to run large models without a dedicated GPU. The company stated that the techniques developed across these platforms directly contributed to the record-setting run in the challenge.

VIDRAFT CEO Kim Min-sik commented: "AI competitiveness is determined not by model size but by how quickly and economically you can deliver a service on the same hardware. Our technology for improving efficiency while preserving quality has now been validated in an open international competition."

Why It Matters

The implications of a six-times inference speedup extend well beyond a leaderboard trophy. VIDRAFT pointed out that if a model runs six times faster on the same GPU, an operator can serve the same workload using roughly one-sixth the number of GPUs — a direct and dramatic reduction in infrastructure costs. In a landscape where GPU access is both expensive and geopolitically sensitive, that kind of efficiency translates into what the company describes as "AI infrastructure sovereignty": the ability to do more with less, independent of raw compute supply.

The competition's strict, judge-verified format also lends the result unusual credibility. Unlike benchmark results published unilaterally by AI labs, The Fast Gemma Challenge score is reproducible and publicly visible on an official leaderboard, making it a transparent, third-party-validated proof point.

More broadly, the result reinforces a growing argument in AI development circles: that the next wave of differentiation will come from inference optimization — squeezing more performance out of existing hardware — rather than from simply training ever-larger models. VIDRAFT's stack, which spans foundation model development (the open AETHER model), a family of evolving models under the Darwin line, inference engine technology, on-device AI, and research in physics, chemistry, life sciences, and drug discovery, positions the company as a vertically integrated player in that efficiency-first paradigm.

Key Takeaways

Frequently Asked Questions

Q: What is The Fast Gemma Challenge?

A: The Fast Gemma Challenge is a global AI inference optimization competition co-organized by Google's Gemma team and Hugging Face. Participants must maximize the inference speed of Google's open Gemma-4-E4B-it model on a standardized NVIDIA A10G GPU using only software optimizations, with rankings determined by scores that organizers independently verify and reproduce.

Q: What score did VIDRAFT's vidraft-darwin achieve in the competition?

A: vidraft-darwin recorded a verified throughput of 510.58 tokens per second and a perplexity score of 2.3929 — approximately six times the original model's baseline inference speed — earning the top verified position on the leaderboard.

Q: How does faster inference translate into real-world business value?

A: According to VIDRAFT, a six-times inference speedup means operators can run the same AI service on approximately one-sixth the number of GPUs, substantially cutting infrastructure costs and reducing dependence on large hardware supplies — what the company frames as AI infrastructure sovereignty.


Source: AI타임스 (2026-08-03) — original article

Published by VIDRAFT · All posts