A Korean AI startup just topped a rigorous, host-verified inference benchmark — and the margin is hard to ignore.
TL;DR: VIDRAFT, a Korean AI startup led by CEO Kim Min-sik, has achieved the top verified score in the inference optimization competition "The Fast Gemma Challenge," co-hosted by Google and Hugging Face. The company's vidraft-darwin agent recorded 510.58 tokens per second (TPS) with a perplexity (PPL) of 2.3929 on a standardized hardware setup. The result represents approximately six times the inference speed of the baseline model while maintaining answer quality.
VIDRAFT, a Korean deep-tech AI startup, made international headlines on August 3, 2026, when it announced a top verified ranking in "The Fast Gemma Challenge" — a competitive inference optimization event jointly organized by Google and Hugging Face. The Seoul-based company, helmed by CEO Kim Min-sik, says its vidraft-darwin agent outpaced all other verified entrants on the leaderboard, demonstrating both raw throughput and strong output quality under identical hardware conditions.
The Fast Gemma Challenge tasks participating teams with squeezing maximum inference performance out of Google's gemma-4-E4B-it model running on a standardized NVIDIA A10G GPU environment. Critically, the competition rules prohibit hardware swaps or custom silicon advantages — every optimization must come purely from software tuning. Scores are not self-reported; the organizers run their own closed-environment reproduction tests, and only results that pass this independent verification earn the official "VERIFIED" designation on the public leaderboard.
Under those conditions, VIDRAFT's vidraft-darwin agent clocked 510.58 tokens per second alongside a perplexity score of 2.3929. According to VIDRAFT, this throughput figure is roughly six times faster than the unoptimized baseline model — all without sacrificing measurable output quality, as reflected in the PPL metric.
The company attributed the achievement to two proprietary technologies applied during the optimization process: its internally developed VK-series inference engine and its large-model serving platform called POCKET. Neither system's internal specifications were disclosed.
CEO Kim Min-sik commented that the public competition format gave the company a valuable opportunity to validate the efficiency and competitiveness of its technology against real-world, independently verified benchmarks.
The verified result is publicly accessible through the official Fast Gemma Challenge leaderboard.
Benchmark victories at industry competitions carry weight precisely because of their verification rigor — and The Fast Gemma Challenge's host-side reproduction requirement makes gaming the results considerably harder than in self-reported evaluations. Earning a top "VERIFIED" score signals that the performance is reproducible under neutral conditions, which is a meaningful credibility signal for enterprise customers and investors evaluating AI infrastructure vendors.
The timing is also notable. As generative AI deployment scales globally, the economics of inference — how much compute is consumed per output token — have become a central concern for businesses running large language models in production. A team that can deliver roughly six times the throughput on the same GPU hardware, without degrading output quality, is essentially multiplying the effective capacity of existing infrastructure. That translates directly into lower per-query costs and better service responsiveness.
VIDRAFT operates at the intersection of AI modeling, inference engine development, and quantum-computing-based research — a broad technical scope that positions the company as more than a single-product play. This competition win adds a publicly verifiable data point to its credentials in the inference optimization space specifically, which is increasingly where competitive differentiation in the AI industry is being decided.
For Korean AI startups more broadly, a top finish in a globally open competition hosted by two of the industry's most prominent platforms — Google and Hugging Face — offers meaningful international visibility at a moment when Korean deep-tech is actively seeking recognition on the world stage.
Q: What is The Fast Gemma Challenge, and who runs it?
A: The Fast Gemma Challenge is an inference optimization competition co-hosted by Google and Hugging Face. Teams compete to achieve the fastest processing speeds on Google's gemma-4-E4B-it model within a fixed, standardized hardware environment, using software optimization only. Only results independently verified by the organizers receive official recognition.
Q: What scores did VIDRAFT's vidraft-darwin agent achieve?
A: The vidraft-darwin agent recorded 510.58 tokens per second (TPS) and a perplexity (PPL) of 2.3929 in the host-verified evaluation, which VIDRAFT states is approximately six times faster than the baseline model's unoptimized inference speed.
Q: What technologies did VIDRAFT use to achieve this result?
A: VIDRAFT applied its proprietary VK-series inference engine and its large-model serving platform, POCKET, to optimize the model's inference performance. The optimization was purely software-based, with no changes to the standardized hardware environment specified by the competition rules.
Source: 한국경제TV (2026-08-03) — original article