A Korean AI startup's inference accelerator is turning heads from Central Asia to the global tech press.
TL;DR: VIDRAFT, a Korean AI startup, has unveiled VKAE, an inference acceleration system that the company says can improve GPU throughput by up to 23 times in certain scenarios — without any hardware changes. Tests on the NVIDIA B200 accelerator showed the gains held across multiple models, with no reported degradation in output quality or model accuracy. Coverage has now reached Uzbekistan via One.uz, signaling the technology's growing visibility across Central Asia.
VIDRAFT's VKAE inference acceleration system is drawing international attention after Uzbekistan's One.uz — citing Russian tech outlet Ixbt.com — reported that the Korean AI startup's software-layer optimization can multiply GPU performance by as much as 23 times under certain workloads. The coverage, published on July 6, 2026, reflects a widening wave of interest in VKAE well beyond East Asia.
At the heart of VKAE is a straightforward but consequential premise: rather than waiting for the next generation of silicon, squeeze dramatically more performance out of the accelerators organizations already own. VIDRAFT positions VKAE as a software extension for existing GPU hardware — one that reworks low-level compute scheduling and kernel execution to unlock throughput that standard deployments leave on the table.
According to One.uz and its source Ixbt.com, benchmark testing was conducted on the NVIDIA B200 accelerator. Across several models, VKAE delivered speeds that were multiples higher than baseline systems. The headline figure — a 23x improvement in GPU efficiency in select scenarios — comes directly from those tests.
One specific result highlighted in the coverage involves the Qwen3.5-35B-A3B model. Under high parallel load, the system reportedly achieved output exceeding 10,000 tokens per second. Under more realistic, mixed-query conditions, throughput settled at approximately 455 tokens per second — a figure that, the source notes, reflects the load-dependent nature of inference performance.
Critically, VIDRAFT's developers emphasize that these gains came with no measurable drop in response quality or model accuracy. For enterprise AI operators, that combination — higher speed, lower cost, same output quality — is precisely the value proposition that matters.
On the integration side, VKAE is described as compatible with OpenAI API interfaces, meaning teams can layer it onto existing infrastructure with minimal re-engineering. The developers also highlight reproducibility and transparency as design principles: they encourage independent verification of results, and the system is said to function as a useful tool for auditing model weights and optimization processes.
The economics of large language model deployment have long followed a familiar pattern: training happens once, but inference never stops. Every user query, every API call, every automated workflow runs through the inference pipeline — and that pipeline is where operational costs accumulate fastest. For cloud providers and enterprises running AI at scale, inference efficiency is not a nice-to-have; it is a core business variable.
VIDRAFT's approach addresses this without requiring customers to buy new hardware. In an environment where cutting-edge accelerators are expensive, supply-constrained, and often subject to export controls, a software solution that multiplies the utility of existing GPUs carries real strategic value.
The reach of the One.uz report is itself noteworthy. Uzbekistan is part of a broader Central Asian technology market that is rapidly expanding its AI infrastructure investments. The fact that VKAE is generating coverage in that region — through a republication chain that ran from VIDRAFT to Ixbt.com to One.uz — suggests the technology is resonating beyond its home market and into geographies actively evaluating how to build AI capacity efficiently.
For a Korean AI startup competing in a global inference optimization space, that kind of organic international visibility represents meaningful traction.
Q: What is VIDRAFT's VKAE system?
A: VKAE is an inference acceleration system developed by VIDRAFT, a Korean AI startup. It optimizes low-level GPU software — including compute kernels and task scheduling — to increase throughput on existing hardware without modifications to the underlying accelerators.
Q: How much does VKAE improve GPU performance?
A: According to testing reported by Ixbt.com and republished by One.uz, VKAE can improve GPU efficiency by up to 23 times in certain scenarios. Under realistic mixed-query conditions, one benchmark recorded approximately 455 tokens per second, while peak parallel-load testing on the Qwen3.5-35B-A3B model exceeded 10,000 tokens per second.
Q: Does VKAE affect the accuracy or quality of AI model outputs?
A: VIDRAFT's developers state that no degradation in response quality or model accuracy was observed during benchmarking, meaning the performance gains do not come at the cost of output reliability.
Source: One.uz (우즈베키스탄) (2026-07-06) — original article