VIDRAFT · Korean Pre-AGI AI startup · 2026-09-06

VIDRAFT's VKAE Boosts AI Inference Speed 23x Without New Hardware

As AI compute costs soar, software-level optimization is emerging as a powerful alternative to buying more chips.

TL;DR: VIDRAFT, a Korean AI startup, has published benchmark results for its inference acceleration system VKAE, which reportedly delivers up to 23× throughput gains on existing GPU hardware without any hardware changes. Testing was conducted on a single GPU, and developers state that no degradation in model output quality or accuracy was observed across any of the measured runs. The results include performance data for the Qwen3.5-35B-A3B model, which surpassed 10,000 tokens per second under high-concurrency conditions.

VIDRAFT, the Korean AI startup specializing in language models and AI infrastructure, has released benchmark data for VKAE, its inference acceleration system, showing performance improvements of up to 23 times over baseline inference configurations — all without replacing or upgrading any physical hardware. The findings were reported on July 6, 2026, by Russian regional technology outlet vTambove, reflecting growing international attention to software-driven approaches to AI efficiency.

What VIDRAFT Announced

VIDRAFT's VKAE system is designed to extract significantly more performance from existing GPU accelerators through low-level software optimization — including improvements to compute kernels and task scheduling mechanisms — rather than relying on next-generation hardware. The company positions VKAE as a kind of "software extension" for accelerators that enterprises and cloud providers already own.

According to the published benchmark data, testing was performed on a single GPU unit, with throughput in the optimized mode compared directly against a baseline inference setup using the same measurement environment. Across all tested runs, the developers explicitly report that no reduction in output quality or model accuracy was detected.

One of the headline results involves the Qwen3.5-35B-A3B model. Under high-concurrency load — meaning many simultaneous requests — the system demonstrated aggregate throughput exceeding 10,000 tokens per second. The developers are transparent about an important nuance here: that figure is workload-dependent. Under more varied, real-world query patterns, the same model delivered approximately 455 tokens per second. VIDRAFT presents both numbers, making clear that peak throughput and realistic throughput reflect different operating conditions.

The acceleration factor of up to 23× is not described as a universal property of the system. VIDRAFT acknowledges that gains vary considerably depending on the architecture of the model being served. Some models see dramatic speedups; others show more modest improvements. The company attributes this variance to differences in computational bottlenecks, memory organization, and the internal structural characteristics of individual model architectures.

Reproducibility appears to be a stated priority. VIDRAFT notes that benchmarks can be reproduced via a ready-to-deploy container, lowering the barrier for independent verification by researchers, enterprise evaluators, and potential partners.

Why It Matters

The economics of AI deployment make inference optimization an increasingly critical frontier. Training a large language model is a one-time event, but inference — generating responses for users, APIs, and enterprise systems — is continuous and cumulative. For cloud AI services, corporate AI platforms, and consumer-facing products, inference costs are the dominant operational expense, not training. Any system that meaningfully reduces the compute required per inference request translates directly into lower costs and greater capacity without capital expenditure on new hardware.

VIDRAFT's approach sits in a growing category of inference optimization technologies that aim to close the gap between the raw theoretical performance of GPU accelerators and what typical serving frameworks actually extract from that hardware. While chipmakers advance through successive hardware generations, software-layer optimizations can deliver substantial gains to operators who are locked into current infrastructure due to supply constraints, procurement cycles, or budget limitations — a reality that is particularly acute given ongoing global GPU scarcity.

The fact that these results are drawing coverage from international media, including a Russian regional outlet with a technology readership, signals that VIDRAFT's work is resonating beyond its home market. For a pre-AGI Korean AI startup operating in a field dominated by U.S. and Chinese players, that kind of cross-border visibility is meaningful.

Key Takeaways

Frequently Asked Questions

Q: What is VIDRAFT's VKAE and what does it do?

A: VKAE is an inference acceleration system developed by VIDRAFT, a Korean AI startup. It optimizes low-level software — including compute kernels and scheduling — to increase AI model throughput on existing GPU hardware without requiring any new equipment.

Q: How much faster is VKAE compared to standard inference setups?

A: According to VIDRAFT's published benchmarks, VKAE achieved up to 23× faster throughput compared to baseline configurations in some scenarios, though gains vary depending on the architecture of the model being run.

Q: Does VKAE affect the quality or accuracy of AI model outputs?

A: VIDRAFT states that across all measured benchmark runs, no degradation in output quality or model accuracy was observed.


Source: vTambove (러시아) (2026-07-06) — original article

Published by VIDRAFT · All posts