VIDRAFT · Korean Pre-AGI AI startup · 2026-08-16

VIDRAFT AX-Ray Detects "Causal Leakage" Flaws in Two Public AI Models

A Korean AI startup's new safety diagnostic tool is redefining how the world evaluates large language models.

TL;DR: VIDRAFT, a Korean deep-tech AI startup, publicly released its AI safety diagnostic platform AX-Ray on Hugging Face on August 13, 2026, alongside an evaluation dataset and leaderboard. Using AX-Ray, VIDRAFT successfully detected and reproduced "causal leakage" vulnerabilities in two general-purpose large language models, including one developed by NVIDIA — marking the first confirmed identification of this flaw in real-world LLMs.

VIDRAFT, the Seoul-based deep-tech AI startup, made a significant advance in AI safety on August 13, 2026, by publicly releasing its AX-Ray safety diagnostic platform and evaluation dataset on Hugging Face — and simultaneously announcing that the tool had identified dangerous "causal leakage" signals in two widely used, publicly available AI models, one of which was built by NVIDIA.

What VIDRAFT Announced

VIDRAFT's AX-Ray is not a performance benchmark in the traditional sense. Rather than measuring how capable or intelligent a model is, it is designed to uncover how dangerous or unpredictable a model might be in real-world deployment. The platform was released publicly on Hugging Face together with a leaderboard, allowing the broader AI research community to examine and compare model safety profiles.

The central finding accompanying the launch was the detection and successful reproduction of "causal leakage" in two general-purpose large language models. Causal leakage refers to a phenomenon in which an AI model's reasoning or behavior is influenced not by the intended logical pathway, but by hidden information or unintended causal cues embedded within its architecture. In practical terms, this means a model could be quietly nudged toward decisions or actions that its designers never anticipated — and that safety guardrails might not catch.

Until now, causal leakage had largely remained a theoretical concern within AI safety discussions. Identifying it experimentally in real-world, production-grade LLMs had been considered an open and difficult challenge. VIDRAFT's AX-Ray claims to have done exactly that, capturing signals in two models and reproducing the conditions under which those signals appeared.

The AX-Ray diagnostic framework covers 117 distinct safety evaluation criteria. What makes it especially notable is its legal and normative grounding: each criterion is mapped to existing laws and regulatory frameworks across multiple countries. For regions including Arab nations, the framework goes further, incorporating religious and social norms — such as Sharia-derived legal standards — alongside civil law, enabling culturally and jurisdictionally nuanced AI safety assessments. This design philosophy reflects VIDRAFT's ambition to build a globally applicable diagnostic standard, not just a technically oriented research tool.

Why It Matters

The launch of AX-Ray arrives at a moment when the AI industry's center of gravity is shifting. For years, the dominant question driving AI development has been: how smart is this model? Increasingly, regulators, enterprises, and researchers are asking a different question: how safe and controllable is this model?

This shift is not merely philosophical. As large language models and autonomous AI agents are deployed in high-stakes sectors — finance, healthcare, robotics, national defense, and public services — the consequences of unpredictable or guideline-violating behavior grow more severe. A model that quietly routes around its safety constraints, or one that responds to hidden causal triggers in ways its operators cannot anticipate, poses risks that raw benchmark scores cannot capture.

The hypothesis that causal leakage could underlie some of AI's most alarming failure modes — including autonomous agents taking sudden dangerous actions or AI systems accessing resources they were never intended to reach — has been circulating in safety research circles. VIDRAFT is now claiming to have moved that hypothesis closer to empirical confirmation, at least in the controlled diagnostic context of AX-Ray.

VIDRAFT CEO Kim Min-sik stated that the next competitive frontier in AI is safety, not intelligence, and that the company intends to develop AX-Ray into a comprehensive global AI safety diagnostic system that maps dangerous causal relationships inside models to the legal and social norms of individual countries.

Beyond AX-Ray, VIDRAFT operates at the intersection of multiple deep-tech domains. The company is developing its own foundation models, a quantum operating system, and an autonomous platform for drug and materials discovery. It is a resident company at the Seoul AI Hub and a recipient of government-backed advanced computing support.

Key Takeaways

Frequently Asked Questions

Q: What is causal leakage in AI, and why is it dangerous?

A: Causal leakage occurs when an AI model's decisions or behaviors are driven by hidden information or unintended causal cues rather than its designed reasoning path. This can lead to unpredictable actions, including bypassing safety guardrails or accessing systems in ways developers never intended.

Q: What is VIDRAFT's AX-Ray, and where can it be accessed?

A: AX-Ray is an AI safety diagnostic platform developed by VIDRAFT that evaluates large language models across 117 safety criteria mapped to laws and norms in multiple countries. It was released publicly on Hugging Face on August 13, 2026, along with an evaluation dataset and a model leaderboard.

Q: Which AI models were found to have causal leakage signals by AX-Ray?

A: VIDRAFT reported detecting causal leakage signals in two general-purpose publicly available AI models, one of which was developed by NVIDIA. The company has not publicly disclosed the full list of tested models beyond this.


Source: 전자신문 (2026-08-14) — original article

Published by VIDRAFT · All posts