VIDRAFT · Korean Pre-AGI AI startup · 2026-08-07

VIDRAFT Claims Global No. 1 in Google's Fast Gemma Challenge

A Korean AI startup just beat the world at inference speed — and proved it with verified results.

TL;DR: VIDRAFT, a Korean AI deep-tech startup, ranked first globally in the verified records category of The Fast Gemma Challenge, a competitive AI inference benchmark co-hosted by Google's Gemma team and Hugging Face. The company's autonomous AI agent, vidraft-darwin, achieved 510.58 tokens per second — more than six times the speed of the baseline model — while fully preserving output quality. Results were announced on August 3, 2026.

Korean AI startup VIDRAFT made headlines on August 3, 2026, when it announced it had taken the top global spot in the verified records category of The Fast Gemma Challenge, a competitive AI inference optimization event organized jointly by Google's Gemma team and Hugging Face.

What VIDRAFT Announced

The Fast Gemma Challenge tasks participants with pushing the inference speed of Google's multimodal model, gemma-4-E4B-it, to its absolute limits — using only software-level optimization, with every team working on identical hardware. The benchmark runs on NVIDIA A10G GPUs, meaning no competitor can gain an edge through custom silicon or infrastructure upgrades. Performance alone determines the ranking.

VIDRAFT entered the competition with its autonomous AI agent, vidraft-darwin. After undergoing the organizers' own rigorous, private verification process — reserved exclusively for "VERIFIED" results that have been independently tested for both speed and output quality — vidraft-darwin clocked 510.58 tokens per second (TPS) and a perplexity (PPL) score of 2.3929. According to VIDRAFT, that throughput figure represents more than a sixfold improvement over the unoptimized baseline model, achieved without any degradation in response quality.

The company attributes this result to its proprietary VK inference engine and the optimization techniques embedded in POCKET, its on-device platform. Together, these technologies enabled the kind of software-only performance gains that the challenge is specifically designed to surface and reward.

VIDRAFT CEO Kim Min-sik was quoted by 동아일보 as emphasizing that delivering fast, economical AI services on the same hardware as everyone else is the core competitive differentiator of the coming AI era, and that the challenge result represented an objective, internationally verified validation of the company's capabilities.

Why It Matters

On the surface, finishing first in a speed benchmark sounds like a trophy for engineers. In practice, the implications for AI economics are substantial. When a model runs at more than six times its baseline inference speed, the cost structure of deploying it shifts dramatically. The same volume of requests can be served by a fraction of the compute resources — VIDRAFT estimates the efficiency gain translates to roughly one-sixth the infrastructure overhead for an equivalent service workload. For businesses and developers building on top of large language models, that kind of efficiency improvement has a direct and immediate impact on operating costs.

What also distinguishes this result is the verification mechanism. The leaderboard ranking was calculated solely from results that the organizers themselves tested privately before certifying as VERIFIED — a stricter standard than self-reported benchmarks. For an independent startup competing on a global stage against well-resourced teams, earning the top verified rank carries considerably more credibility than a community-submitted score.

VIDRAFT positions itself as a deep-tech AI company spanning foundation model development and AI infrastructure. Its foundation model, AETHER, anchors the company's longer-term research roadmap. More recently, the company has released Darwin Family — a model-merging AI framework designed to enhance LLM inference performance without additional training — and MARL, a runtime middleware built to reduce hallucination in AI outputs. The Fast Gemma Challenge result adds an external, third-party validation milestone to that growing portfolio.

Key Takeaways

Frequently Asked Questions

Q: What is The Fast Gemma Challenge?

A: It is a competitive AI inference benchmark co-organized by Google's Gemma team and Hugging Face, where participants optimize the inference speed of Google's gemma-4-E4B-it multimodal model using software techniques only, on identical hardware.

Q: What performance did VIDRAFT's agent achieve?

A: vidraft-darwin recorded 510.58 tokens per second (TPS) with a perplexity score of 2.3929, representing more than six times the speed of the unoptimized baseline model while maintaining full output quality.

Q: Why does faster inference speed matter for AI costs?

A: Higher inference throughput means more requests can be handled on the same hardware. VIDRAFT notes that a sixfold speed improvement allows an equivalent service to be run on approximately one-sixth the infrastructure, directly reducing operational expenses.


Source: 동아일보 (2026-08-03) — original article

Published by VIDRAFT · All posts