VIDRAFT · Korean Pre-AGI AI startup · 2026-09-04

VIDRAFT's VKAE System Delivers Up to 23x GPU Performance Gains

Korean AI startup VIDRAFT is rewriting the economics of AI inference — without touching the hardware.

TL;DR: VIDRAFT, a Korean AI startup, has introduced VKAE, an inference acceleration system that can boost GPU performance by up to 23 times in certain scenarios on existing hardware. Tests conducted on an NVIDIA B200 accelerator showed throughput of over 10,000 tokens per second under high parallel load for the Qwen3.5-35B-A3B model, with no measurable drop in output quality. The findings were reported by Uzbekistan-based outlet Zamin.uz on July 6, 2026, citing Russian tech publication Ixbt.com.

VIDRAFT, the Korean AI startup, made headlines this week after Uzbekistan's Zamin.uz reported on its newly unveiled VKAE inference acceleration system — a software-layer solution that claims to multiply the effective performance of existing GPU hardware by as much as 23 times in targeted workloads, all without requiring any hardware upgrades.

What VIDRAFT Announced

At the center of the announcement is VKAE, a system VIDRAFT describes as a software extension for AI accelerators. Rather than waiting for the next generation of chips, VKAE is designed to extract dramatically more throughput from hardware that organizations already own and operate.

According to Zamin.uz — citing Ixbt.com — benchmark testing was carried out on an NVIDIA B200 graphics accelerator. The results reportedly exceeded the developers' own expectations. Across multiple tested models, speeds several times higher than baseline systems were recorded. Crucially, VIDRAFT stated that no degradation in response quality or model accuracy was observed during these tests, meaning the performance gains come without the typical trade-off of reduced output reliability.

The most striking figure in the report centers on the Qwen3.5-35B-A3B model. Under high-level parallel load conditions, VKAE demonstrated a generation throughput of over 10,000 tokens per second. Under real-world, mixed-query conditions, that figure settled at approximately 455 tokens per second — a distinction the developers themselves highlighted to underscore that efficiency is inherently tied to the nature of the workload being run.

On the integration side, VIDRAFT has positioned VKAE for low-friction deployment. The system is described as fully compatible with OpenAI API interfaces, meaning organizations could, in principle, layer it on top of existing AI infrastructure with minimal reconfiguration. The developers have also released a container bundling model weights and an optimized runtime environment, enabling independent verification of the claimed results — something VIDRAFT frames as a foundational trust criterion for the technology.

A detailed scientific paper on the underlying methodology is expected to be published, though the precise technical mechanisms behind VKAE's performance gains remain undisclosed for now.

Why It Matters

The timing of VIDRAFT's announcement reflects a broader shift in how the AI industry is thinking about cost and scale. Training a large language model is a one-time capital expenditure, but inference — the continuous process of generating responses for real users — is an ongoing operational cost that compounds at scale. For cloud providers and enterprise AI platforms alike, inference efficiency is increasingly the determining factor in whether an AI product is economically viable.

VKAE enters this conversation as a software-only answer to a problem the industry has largely been trying to solve with ever-more-powerful — and ever-more-expensive — silicon. If the benchmark results hold up under independent scrutiny, the implications are significant: organizations could multiply the effective capacity of their current GPU fleets without procurement cycles, capital outlays, or data center reconfigurations.

Zamin.uz also noted particular relevance for emerging technology markets such as Uzbekistan, where AI adoption is growing but access to cutting-edge server hardware remains constrained. A software acceleration layer that makes existing infrastructure perform like more advanced hardware could meaningfully lower the barrier to entry for local startups and enterprises building AI-powered services.

The reproducibility emphasis from VIDRAFT — providing a verifiable container environment and committing to an upcoming research publication — is notable in an industry where performance claims often outrun independent confirmation.

Key Takeaways

Frequently Asked Questions

Q: What is VIDRAFT's VKAE system?

A: VKAE is a software-based inference acceleration system developed by Korean AI startup VIDRAFT. It is designed to increase the effective performance of existing GPU hardware — by up to 23 times in certain scenarios — without requiring any hardware changes.

Q: Does VKAE reduce AI output quality in exchange for faster performance?

A: According to VIDRAFT's own testing, no decrease in response quality or model accuracy was observed. The company states that reliability is maintained even as throughput increases significantly.

Q: How can organizations verify VKAE's performance claims?

A: VIDRAFT has released a container that includes model weights and an optimized runtime environment, allowing third parties to independently reproduce and verify the reported results. A detailed scientific paper on the technology is also expected to be published.


Source: Zamin.uz (우즈베키스탄) (2026-07-06) — original article

Published by VIDRAFT · All posts