VIDRAFT · Korean Pre-AGI AI startup · 2026-08-06

VIDRAFT Claims World No. 1 Verified Score in Google & Hugging Face Fast Gemma Challenge

A Korean AI startup just proved that inference optimization—not model size—may be the real competitive frontier.

TL;DR: VIDRAFT, a Seoul-based Pre-AGI AI startup led by CEO Minsik Kim, secured the world's top verified ranking in "The Fast Gemma Challenge," a global inference optimization competition co-hosted by Google's Gemma team and Hugging Face. The company's autonomous AI agent, vidraft-darwin, achieved 510.58 tokens per second with a perplexity score of 2.3929 under official re-verification conditions—representing more than a sixfold speed improvement over the baseline model while preserving output quality. The result was confirmed on the competition's official leaderboard.

VIDRAFT, the Korean AI deep-tech startup headquartered in Seoul, has claimed the top verified position in a high-profile global inference optimization competition, according to a report published by 전자신문 on August 3, 2026. The achievement comes from "The Fast Gemma Challenge," a contest jointly organized by Google's Gemma team and Hugging Face, designed to test pure software optimization skill rather than hardware or model-swapping advantages.

What VIDRAFT Announced

The Fast Gemma Challenge operates under strict conditions intended to create a genuinely level playing field. Every participant receives identical hardware—an NVIDIA A10G GPU—and works with the same publicly available multimodal model, Google's gemma-4-E4B-it. Competitors are prohibited from swapping out the model or disabling any of its text, image, or audio capabilities. The goal is straightforward: make the software run as fast as possible through optimization alone, without cutting corners on output quality.

What makes the leaderboard particularly credible is its verification system. Rankings are determined solely by "VERIFIED" scores—results that the organizers themselves reproduce using undisclosed test prompts, independently confirming both speed and quality. This structure makes it structurally impossible for participants to inflate or game their results.

VIDRAFT's autonomous AI agent, vidraft-darwin, cleared this verification process and posted a throughput of 510.58 tokens per second alongside a perplexity score of 2.3929. According to the company, this represents a more than sixfold increase in inference speed compared to the model's default baseline—while maintaining answer quality, a balance that many optimization attempts sacrifice in the pursuit of raw speed.

The company attributes the result to its proprietary VK-series inference engine and the optimization expertise developed through POCKET, its on-device AI platform capable of running large models without a dedicated GPU. VIDRAFT's broader technical portfolio also includes the fully open foundation model AETHER and the Darwin family of evolutionary models, giving the company an end-to-end stack that spans model development, compression, and deployment.

CEO Minsik Kim commented that AI competitiveness is no longer defined by model scale alone, but by how efficiently and economically a model can be served on the same hardware—and that this competition validated VIDRAFT's approach to raising efficiency without compromising quality.

Why It Matters

The significance of this result extends well beyond a leaderboard position. The AI industry's competitive axis is shifting. Building a frontier model matters, but so does the ability to serve it cheaply and quickly at scale. A sixfold improvement in inference throughput has a direct operational translation: the same workload can theoretically be handled with a fraction of the GPU resources previously required. In an environment where cloud compute costs are a major constraint for AI deployment, inference optimization is rapidly becoming as strategically important as model training.

For Korean AI, the result carries additional weight. VIDRAFT's verified No. 1 ranking in an international open competition—run by two of the most prominent names in the global AI ecosystem, Google and Hugging Face—provides independent, third-party validation of a Korean startup's technical capabilities on a world stage. It signals that inference engineering, a discipline that often operates in the shadow of headline-grabbing model releases, is an area where Korean deep-tech firms can compete at the highest level.

VIDRAFT describes itself as a Pre-AGI and AI infrastructure research company, with work spanning AI models, inference engines, on-device AI, and quantum computing—applied across domains including physics, chemistry, life sciences, and drug discovery.

Key Takeaways

Frequently Asked Questions

Q: What is The Fast Gemma Challenge?

A: It is a global inference optimization competition co-run by Google's Gemma team and Hugging Face, where all participants use identical hardware and the same open multimodal model, competing solely on software optimization to maximize speed without degrading output quality.

Q: How did VIDRAFT's vidraft-darwin perform?

A: It achieved a verified throughput of 510.58 tokens per second with a perplexity score of 2.3929, earning the top verified ranking on the official leaderboard and representing more than a sixfold speed improvement over the model's default baseline.

Q: Why does inference optimization matter for AI businesses?

A: Faster inference on the same hardware directly reduces the compute resources—and therefore the cost—needed to run AI services at scale, making inference efficiency a critical factor in AI deployment economics and infrastructure strategy.


Source: 전자신문 (2026-08-03) — original article

Published by VIDRAFT · All posts