Xiaomi AI Cube runs local LLMs on three in-house Xring chips
Xiaomi has unveiled AI Cube, an experimental mini-PC combining three Xring processors to run AI models with up to 120 billion parameters locally. The prototype signals the company’s expansion beyond mobile chips, although pricing and release plans remain undisclosed.

<p>Xiaomi has demonstrated AI Cube, a prototype mini-PC built to run large language models directly on the device. Local processing could reduce reliance on cloud infrastructure while keeping model data and parameters within the system.</p><p>The platform combines three members of Xiaomi’s in-house Xring family: the O3, O100 and D100. Their integration turns AI Cube into a showcase for the company’s broader semiconductor strategy, which now spans mobile devices, compact computers, neural accelerators and automotive systems.</p><p>Xring O3 serves as the primary SoC and is manufactured by TSMC on a 3-nanometer process. It combines a 10-core CPU, a 16-core G2-Ultra NX GPU and an NPU rated by Xiaomi at up to 200 TOPS, alongside support for LPDDR6 memory delivering as much as 113.8 GB/s of bandwidth.</p><p>The 6-nanometer O100 acts as a dedicated, high-bandwidth AI accelerator. Its three-dimensional packaging and tightly spaced interconnects help provide a claimed throughput of 1.22 TB/s and inference performance reaching 330 tokens per second.</p><p>The third component is Xring D100, originally developed for intelligent-driving workloads. It features a 20-core CPU, a 16-core NPU and support for up to 160 GB of unified memory; while the platform can theoretically handle 200-billion-parameter models, Xiaomi specifies a working range of 3B to 120B for the AI Cube prototype.</p><p>The system is housed in a chassis machined from a single aluminum block, with 33,874 ventilation holes and cooling designed to sustain 150 watts under full load. Users can switch between faster and slower operating modes according to workload complexity, but AI Cube currently remains a technology demonstrator with no announced price or launch date; its chips are Xiaomi designs, though manufacturing is still outsourced to TSMC.</p>
Read also

Qwen3.8-Max: long-running agents enter the spotlight
Qwen is moving beyond traditional coding benchmarks, focusing on agents capable of working autonomously toward the same objective for extended periods: planning, writing code, executing it, receiving feedback, and continuously iterating.

OpenAI Disbands Risk Team
OpenAI has disbanded its preparedness team, responsible for assessing and mitigating risks associated with AI models. The team's role was to identify potential risks and develop strategies to address them.

LLMs Mirror Human Brain
Large Language Models develop a modular architecture mirroring the human brain, with tasks recruiting overlapping neurons across cognitive domains. This discovery sheds light on the fundamental principles of intelligent systems.