VIDRAFT · Korean Pre-AGI AI startup · 2026-08-01

VIDRAFT POCKET-35B: Run a 35B AI Model Without a GPU

A Korean AI startup's on-device model is turning heads from Reddit to China's tech community.

TL;DR: VIDRAFT's POCKET-35B is a 35-billion-parameter on-device language model released on Hugging Face by community team FINAL-Bench in GGUF quantized format, designed to run on CPUs without a dedicated GPU. The model supports PCs, mini-PCs, Android phones, iPhones, iPads, and Macs across multiple quantization tiers. A Reddit r/LocalLLaMA post reporting up to 59 tokens per second on CPU alone helped the release spread rapidly, reaching Chinese-language tech audiences by July 26, 2026.

VIDRAFT's POCKET-35B made a remarkable leap from a Korean AI startup's release into the global local-inference conversation — and then into China's tech press — largely on the strength of a single Reddit post. Chinese outlet 桃子快讯 reported on July 26, 2026, that FINAL-Bench, the community team behind the release, had published a full suite of GGUF-format quantized models on Hugging Face, each targeting a different class of consumer hardware. The defining promise: no GPU, no CUDA, no cloud required.

What VIDRAFT Announced

The POCKET-35B release centers on native compatibility with llama.cpp, the widely used inference engine that enables large language models to run on ordinary CPU hardware. FINAL-Bench published several quantization tiers under the POCKET-35B-GGUF banner, calibrated to match different memory footprints and quality requirements.

The flagship tier, Q4_K_M, produces a 21 GB file and is aimed at PCs or servers with at least 32 GB of RAM. It carries the lowest perplexity score in the lineup — 5.79 — making it the highest-quality option for users whose hardware can accommodate it.

The Q2_K tier, flagged as the recommended choice, weighs in at 13 GB and is designed to run comfortably on a GPU-free mini-PC as an everyday workhorse. Its perplexity score of 6.49 represents a reasonable tradeoff between compression and output quality.

The most compressed option, IQ1_M, fits in 8.2 GB and targets devices with as little as 16 GB of RAM. While its small footprint is notable, its perplexity score rises to 9.69 — a meaningful quality trade-off that the source acknowledges openly.

Beyond the core 35B line, FINAL-Bench also released several language- and platform-specific variants. A POCKET-KR-GGUF IQ2_M (5.1 GB) targets Korean-language users on Android devices with 8 GB or more of RAM. An equivalent POCKET-KR-MLX 2-bit (5.1 GB) is optimized for Apple Silicon via the MLX inference path, covering iPhone, iPad, and Mac. On the English side, a POCKET-EN-GGUF iPhone-mix (5.3 GB) is built to work with the PocketPal app on iOS, while a POCKET-EN-GGUF PC-mix (6.8 GB) serves both PC and Android users. Perplexity data for the English variants was not published in the release.

The Reddit r/LocalLLaMA post announcing the model claimed a generation speed of 59 tokens per second on CPU. The source notes, however, that the testing conditions behind that figure were not fully disclosed in the post itself, and recommends consulting the Hugging Face repository directly for complete details.

Why It Matters

The local-inference space has grown rapidly, but 35-billion-parameter models have historically demanded high-end GPU hardware to run at practical speeds. POCKET-35B's positioning — a full-scale 35B model that fits on a consumer laptop or a smartphone — pushes that boundary further than most comparable releases. The Q2_K tier in particular, at 13 GB with a perplexity of 6.49, offers a compelling middle ground for users who want near-full model capability without dedicated accelerator hardware.

The speed at which this Korean AI startup's release crossed language and geographic borders is itself notable. Within days of the Reddit post, Chinese-language tech media was covering the model in detail, reflecting both the appetite for CPU-friendly large models and the reach that a well-placed community post can generate. For VIDRAFT, the organic spread from r/LocalLLaMA to China's 桃子快讯 represents grassroots validation that is difficult to manufacture through conventional PR.

That said, the release comes with caveats worth noting. The source of the original base model and the full technical specifications were not comprehensively disclosed in the Reddit announcement. The IQ1_M tier's perplexity jump is significant and suggests that aggressive quantization at this scale still carries real quality costs.

Key Takeaways

Frequently Asked Questions

Q: What hardware do I need to run VIDRAFT POCKET-35B?

A: Requirements vary by tier. The Q4_K_M version needs at least 32 GB of RAM on a PC or server; the recommended Q2_K tier runs on a GPU-free mini-PC; and the IQ1_M version targets devices with 16 GB of RAM. Mobile variants exist for Android, iPhone, iPad, and Mac.

Q: Is POCKET-35B open source and free to use?

A: The model was published on Hugging Face by community team FINAL-Bench as an open release. For full licensing terms and technical documentation, the Hugging Face repository is the authoritative source.

Q: How fast does POCKET-35B run on CPU?

A: The Reddit r/LocalLLaMA release post claimed up to 59 tokens per second on CPU, though the specific hardware configuration and testing conditions behind that figure were not fully disclosed in the post.


Source: 桃子快讯 (중국) (2026-07-26) — original article

Published by VIDRAFT · All posts