FaceWall MiniCPM lands on the Samsung Z Fold8: The first time a domestic on-device model has been integrated into a global flagship.
Samsung London Galaxy Unpacked unveiled the Z Fold8 series last night. Today, MiniCPM, an on-device large model by Mihomo Intelligence, was officially announced as a deep empowerment partner of Galaxy AI. This marks the first time a domestically developed on-device model has entered the global flagship product line of a leading international manufacturer, signaling the beginning of a reshuffle in the supply landscape of on-device AI models for smartphones.
At last night’s Galaxy Unpacked event in London, Samsung unveiled the Galaxy Z Fold8, Z Fold8 Ultra, and Z Flip8 all at once. This morning, Mianbi Intelligence officially announced that its MiniCPM series of on-device large models will deeply empower Samsung Galaxy AI, deployed on the Z Fold8 and Z Fold8 Ultra.
To sum up the importance of this in one line—a domestically developed on-device large model has entered the global flagship lineup of a top international smartphone brand for the first time. This isn’t a China-exclusive edition or a sub-flagship trial—it’s Samsung’s two most expensive foldables this year, running a model built by a Beijing company.

This didn’t happen overnight
The groundwork for this was laid back on July 15, when the Cyberspace Administration of China (CAC) released the latest batch of generative AI service filings, approving seven mobile on-device AIs at once: Apple Intelligence, Huawei Xiaoyi, OPPO Andes GPT, vivo BlueHeart, Xiaomi Pengpai AI, Nubia Doubao Mobile LLM, and—Samsung "Galaxy AI."
Later, 36Kr’s “Emergent Intelligence” column exclusively revealed that Galaxy AI’s on-device large model capability was powered by Mianbi’s MiniCPM series, covering flagship models like the Galaxy S26 Ultra, Z Fold7, and Z Flip7. At the time, the industry reaction was mostly “let’s wait and see”—after all, a filing is one thing, but actually shipping in global flagships is another, especially since Samsung and Google had been deeply tied on Gemini, with Korean media still reporting close collaboration between the two ahead of Galaxy Unpacked.
The result: Samsung did both—Gemini continues to run in the cloud, while MiniCPM powers the on-device side. In the newly released Z Fold8 series, the officially described Galaxy AI capabilities involve dual engines for text understanding and multimodal perception, powered by Mianbi’s models.
The division of labor is very clear: Gemini handles tasks that need internet access, huge knowledge, and complex reasoning; MiniCPM handles tasks that must run locally with low latency and privacy sensitivity—translation, summarization, image recognition, UI understanding—things users do dozens of times a day. Smartphone makers all know users have a low tolerance for lag in “AI assistants”; if there’s even a loading spinner, they’ll complain. On-device models are a necessity, not a backup.
Why Samsung chose Mianbi
Let’s start with the hardware. The Z Fold8 adopts Samsung’s first 4:3 wide fold design with an 8-inch inner screen, more like a small tablet; the Z Fold8 Ultra is the first Ultra in the foldable lineup, with a 200MP main camera, 5000mAh battery, and a starting price of 14,999 RMB. Both devices are positioned for productivity—large screens, multitasking, side-by-side writing and viewing. This form factor naturally fits on-device AI: a larger screen means higher information density, multitasking means the AI assistant is constantly active, and the local model can’t be a toy.
So Samsung needed a model that runs stably on mobile SoCs, is capable enough to represent Galaxy AI, and won’t overheat or drain the battery. That’s a small pool to pick from.
Mianbi made it in thanks to years of betting on “knowledge density.” In 2024, Mianbi, together with the Tsinghua team, proposed the Large Model Densing Law: the capability density of open-source large models doubles roughly every 3.5 months, meaning the parameter scale required for the same level of intelligence decreases exponentially. This idea was heavily questioned at the time—mainstream belief was still that “scaling laws never sleep,” focused on growing parameters, compute, and GPU counts. Mianbi went the other way.
Subsequent model iterations validated the theory:
- MiniCPM5-1B: Released in May 2026, with just 1 billion parameters, scoring 17.9 on the Artificial Analysis Intelligence Index, outperforming many much larger baselines.
- MiniCPM-V 4.6: 1.3B parameters, runs smoothly on a phone with just 6GB of memory, supporting iOS, Android, and HarmonyOS.
- MiniCPM-o 4.5: 9B parameters enabling full-duplex multimodal interaction across speech, video, and text.
1.3B parameters and 6GB memory—these numbers mean everything to smartphone engineers. You don’t have to equip next-gen flagships with 24GB of RAM just to fit the model, nor worry that mid-range models can’t support it at all. For Samsung, a company that ships over 200 million phones a year with extremely complex SKUs, “can run” matters far more than “runs flashy.”
Mianbi’s CTO Zeng Guoyang once said in an interview with 36Kr: “On-device is a hard limit—if the model’s too big to run, no subsidy fixes that; if power draw makes it overheat, you can’t subsidize it with an ice pack.”
In other words: in the cloud, AI can burn cash to subsidize compute; on-device, physics doesn’t give discounts.
Chip-level integration already done
Algorithms alone aren’t enough; on-device models must run across wildly different SoCs. Mianbi has spent years doing the hard work—normalizing chip adaptation:
- Overseas chips: Qualcomm, MediaTek, Intel, Rockchip, NVIDIA, AMD
- Domestic chips: Huawei Ascend, Cambricon, even launching China’s first ternary quantized large model BitCPM-CANN, reducing inference memory usage by about 6×.
Samsung’s Galaxy S and Z series use Qualcomm Snapdragon or in-house Exynos chips depending on the region, and MiniCPM already covers both—a prerequisite for integration. Anyone who’s deployed on-device knows: the same model can display vastly different power usage, tokens/s, and first-token delays across Snapdragon and MediaTek; each platform needs separate tuning.
Mianbi’s open-source community stats also back up its engineering maturity—the MiniCPM series has surpassed 38 million cumulative downloads on GitHub and Hugging Face. That scale suggests the models don’t just run but are actively used by developers for real applications.
The supply landscape of mobile AI is being reshaped
Zooming out, this partnership marks a key moment in a larger trend.
For the past two years, domestic smartphone makers’ on-device AI was primarily self-developed: Huawei Xiaoyi, OPPO AndesGPT, vivo BlueHeart, Xiaomi Pengpai. The logic was simple—if the phone is yours, the model should be too. But that logic has started to loosen in the first half of this year:
- Alibaba’s Qwen integrated into Apple Intelligence (China edition), covering iOS, iPadOS, macOS, and visionOS.
- Mianbi’s MiniCPM deployed in Samsung’s flagships, including the Fold8 and Fold8 Ultra, plus the previously filed S26 Ultra and Z Fold7/Flip7.
The message is clear: smartphone OEMs no longer need to do everything in-house; model companies are emerging as independent suppliers. This mirrors the paths once taken by chipsets, screens, and camera modules—initially every big brand tried full-stack self-development, but over time specialized suppliers outperformed, and OEMs began sourcing instead.
For model companies, this is a far better business path than “building an app.” Consumer AI assistants must fight for attention against WeChat or TikTok—not promising odds; cloud API models are crushed by hyperscalers’ price wars. On-device AI, by contrast, has hard technical barriers—chip adaptation, power optimization, and long-term supply relationships—forming a deeper moat.
Mianbi has already proven the model on the automotive side. Its self-developed on-device agent SuperMate is expected to be delivered in over 300,000 mass-production vehicles by the end of 2026, covering Geely, SAIC Volkswagen, GAC, and Mazda. The automotive chain is even longer than phones—from model adaptation to functional safety to supply chain and OTA—and having run that end-to-end, entering smartphones was natural.
Some unanswered questions
Of course, there are still several open points.
First, which Galaxy AI features exactly run on MiniCPM, and which on Gemini or Samsung’s own models? Samsung’s official line broadly mentions “text understanding and multimodal perception,” but specific user-facing features—real-time translation, gallery search, screen reading—haven’t been clarified.
Second, the monetization mechanism. On-device models don’t have the clear per-token billing of cloud APIs. How model companies get paid per phone shipped—flat licensing, per-device royalties, or revenue-sharing based on active users—has no industry standard yet. This will define the ceiling for the business.
Third, exclusivity. Is this Samsung–Mianbi partnership exclusive, or will it coexist with others? If MiniCPM runs on the Z Fold8 Ultra, what about the S26 lineup later this year? The S27 next January? Will Samsung bring in more Chinese models later? These are things to watch in the coming months.
Some perspective
From Mianbi’s full-on bet on on-device AI starting in late 2023, to raising over 5 billion RMB by mid-2026 and reaching a 20 billion valuation—the highest among domestic on-device AI unicorns—to landing inside Samsung’s flagship tonight, the journey took about two and a half years.
Li Dahai previously said, “2026 will be the first year of large-scale on-device intelligence deployment.” Judging by today’s milestone, that doesn’t sound exaggerated. The CAC filing gave Mianbi the qualification; the Galaxy Z Fold8 series gave it the order.
For developers, on-device and cloud models are no longer either-or—they’ll coexist for the long term in real applications. Cloud models handle heavy reasoning and multi-turn agent workflows; on-device models handle privacy-sensitive, low-latency, always-on, high-frequency use. OpenAI Hub aggregates API access to GPT, Claude, Gemini, DeepSeek, etc., all via a single key compatible with OpenAI SDKs; meanwhile, on-device models like MiniCPM have weights available on Hugging Face and GitHub so both can collaborate within a single app.
The next Galaxy Unpacked will be in January for the S26 lineup. By then, we’ll see whether this collaboration was a one-time deal—or the starting point of Chinese models going global.



