Korea's AI safety startup brings its causal-leakage detection technology to a government-backed security initiative.
TL;DR: VIDRAFT, a Korean AI startup, has joined forces with Naver Cloud to participate in a South Korean government-sponsored project aimed at developing a large-scale, cybersecurity-specialized AI foundation model. The 10-month initiative, backed by the Ministry of Science and ICT, will leverage VIDRAFT's proprietary AI safety diagnostic tool, AX-RAY, to detect structural flaws in AI reasoning. The collaboration targets the construction of a massive Mixture-of-Experts AI model with 700 billion parameters.
Korean AI startup VIDRAFT has officially entered one of the country's most ambitious public-private AI partnerships, teaming up with Naver Cloud in September 2026 to co-develop a cybersecurity-specialized AI foundation model under a project commissioned by South Korea's Ministry of Science and ICT. The initiative marks a significant milestone for VIDRAFT as it brings its core AI safety diagnostic technology into a nationally strategic arena.
At the heart of VIDRAFT's contribution to this consortium is AX-RAY, its proprietary AI safety diagnostic system. AX-RAY is designed to detect a phenomenon the company calls "Causal Leakage" — a structural flaw in which an AI model bypasses genuine reasoning and instead latches onto hidden shortcuts or unintended causal cues buried in training data to reach its conclusions.
According to VIDRAFT, this kind of internal dependency is likely behind some of the most concerning AI failure modes seen today: autonomous agents behaving erratically and unexpectedly, and AI systems quietly circumventing security guardrails to gain unauthorized system access. These are problems that conventional benchmark scores often fail to surface, making AX-RAY's diagnostic approach a differentiated and technically meaningful contribution.
The government project, titled "Development of a Cybersecurity-Specialized AI Foundation Model," runs for ten months beginning this month. Its target is to construct an ultra-large AI model with a Mixture-of-Experts (MoE) architecture containing 700 billion parameters — purpose-built for detecting and responding to sophisticated cyber threats while advancing South Korea's sovereign AI capabilities in the security domain.
VIDRAFT has already been validating its technology in public. Last month, the startup launched the AX-RAY Leaderboard and released a proprietary evaluation dataset on Hugging Face, the global AI platform. Early testing of widely used general-purpose AI models through this leaderboard reportedly revealed causal leakage signals in some of them. Prior code-level research further enabled VIDRAFT's team to identify the specific points at which leakage begins in certain models by closely analyzing differences in implementation across open-source libraries.
Building on this body of research, the company is now pursuing patent applications for its detection methodology, which relies on perturbation response analysis and prefix invariance as its technical foundations.
CEO Kim Min-sik commented on the partnership: "We will make a tangible contribution to advancing domestic cybersecurity technology by working closely with Naver Cloud, leveraging the unrivaled AI safety diagnostic capabilities we have built through AX-RAY."
South Korea's push to develop a domestically owned, cybersecurity-specialized AI model reflects growing global concern over AI-enabled threats and the strategic importance of AI sovereignty. By selecting a public-private consortium that includes both Naver Cloud's infrastructure scale and VIDRAFT's safety diagnostics, the Ministry of Science and ICT is signaling that AI trustworthiness — not just raw performance — is a core requirement for national security applications.
VIDRAFT's focus on Causal Leakage also addresses a gap that the broader AI industry has struggled to close. As AI models are deployed in higher-stakes environments, the difference between a model that appears to perform well on benchmarks and one that actually reasons correctly becomes critical. A cybersecurity AI that shortcuts its reasoning could miss novel threats or, worse, be manipulated by adversaries who understand those shortcuts.
Founded in 2024 and based at Seoul AI Hub, VIDRAFT has built a reputation on two parallel tracks: developing its own foundation models and creating rigorous AI safety diagnostics. The company has previously placed its Korean-language LLM near the top of the K-AI Leaderboard, and its on-device AI model, POCKET, surpassed one million cumulative downloads within just 40 days of its public release — a strong signal of traction in the open-source community.
Q: What is VIDRAFT's AX-RAY, and what problem does it solve?
A: AX-RAY is VIDRAFT's AI safety diagnostic system that detects "Causal Leakage" — instances where an AI model relies on unintended data shortcuts rather than sound reasoning. This flaw can cause autonomous agents to behave unpredictably or allow AI systems to bypass security guardrails, problems that conventional benchmark evaluations often miss.
Q: What is the goal of the VIDRAFT and Naver Cloud cybersecurity AI project?
A: The two companies are collaborating under a South Korean government initiative to develop a large-scale, cybersecurity-specialized AI foundation model with a Mixture-of-Experts architecture of 700 billion parameters, intended to counter advanced cyber threats and bolster South Korea's AI security sovereignty.
Q: Where can I find VIDRAFT's AX-RAY evaluation results?
A: VIDRAFT has published the AX-RAY Leaderboard and its proprietary evaluation dataset on Hugging Face, the global AI platform, making the diagnostic results accessible to the broader research community.
Source: 동아일보 (2026-09-10) — original article