Xiaomi AI Cube and Xring O100: 1.22 TB/s, 330 Tokens/s and 120B Local AI

· AiCybr ·

5 min read Original article ↗

Xiaomi has shown the AI Cube Prototype, a compact local-AI system built around three Xring processors: O3, O100 and D100.

The system sustains up to 150 W and runs a 120B + 3B dual-model configuration locally. Its dedicated AI accelerator, Xring O100 (玄戒 O100), combines 6 nm logic with vertically stacked DRAM and delivers 1.22 TB/s of near-memory bandwidth.

O100 reaches up to 330 tokens/s on Xiaomi MiMo 3B.

Xiaomi AI Cube specifications

Specification Xiaomi AI Cube Prototype
Compute design Three-chip collaborative computing
Chips Xring O3 + Xring O100 + Xring D100
Sustained performance envelope Up to 150 W
Local model configuration 120B + 3B
Model operation Fast/slow dual-model switching
O3 10-core CPU, 16-core G2-Ultra NX GPU, 200 TOPS NPU
O100 6 nm 3D-stacked high-bandwidth AI accelerator
D100 3 nm, 20-core CPU, 16-core NPU, up to 160 GB unified memory support
Chassis Aerospace-grade aluminium unibody
Ventilation 33,874 CNC-machined perforations

O3 provides the general-purpose CPU/GPU/NPU platform, O100 handles high-bandwidth AI inference, and D100 adds Xiaomi's large-memory high-compute architecture.

Xring O100 specifications

Specification Xring O100
Purpose High-bandwidth AI accelerator for on-device large models
Logic process 6 nm
Packaging 3D wafer-level vertical stacking
Stacking method Wafer on Wafer
Bonding Hybrid Bonding
Physical interconnect Face-to-Face metal-layer connection
Stack 1 × 6 nm logic die + 2 × DRAM dies
Bonding pitch 1.4 μm
TSV diameter About 0.7 μm
Effective data lines 28,672
NPU 14-core high-bandwidth NPU
Internal interconnect Xuanwu high-bandwidth matrix bus
Near-memory bandwidth 1.22 TB/s
Bandwidth increase vs Xiaomi flagship-phone reference 16×
Demonstrated model Xiaomi MiMo 3B
MiMo 3B inference Up to 330 tokens/s
Prototype cooling 10 W-class active-air heat dissipation
Commercial rollout 2027

3D-stacked logic and DRAM

O100 places DRAM directly above the AI logic to create a short, wide path between model weights and the NPU.

The stack combines:

  1. one 6 nm logic die;
  2. two DRAM dies;
  3. Wafer-on-Wafer vertical assembly;
  4. Hybrid Bonding;
  5. Face-to-Face metal connections;
  6. 28,672 effective data lines.

The bonding pitch is 1.4 μm and the TSV diameter is about 0.7 μm. Xiaomi's packaging comparison increases the data-path count from 96 lines in a conventional PoP-style reference to 28,672 lines in O100.

1.22 TB/s near-memory bandwidth

O100 provides 1.22 TB/s between its stacked memory and AI compute subsystem.

Large-model inference repeatedly reads model weights during token generation, making memory bandwidth a major part of inference throughput.

Platform Published memory bandwidth
RTX 5090 1.792 TB/s GDDR7
Xring O100 1.22 TB/s near-memory
Apple M3 Ultra 819 GB/s unified memory
Radeon AI PRO R9700 640 GB/s GDDR6
NVIDIA DGX Spark 273 GB/s coherent LPDDR5X
AMD Ryzen AI Halo 256 GB/s LPDDR5X
Xring O3 113.8 GB/s LPDDR6

O100 reaches its bandwidth through the vertical memory stack and dense interconnect rather than a conventional external memory bus.

330 tokens/s on Xiaomi MiMo 3B

Xring O100 runs Xiaomi MiMo 3B at up to 330 tokens per second.

A 3-billion-parameter model occupies approximately:

Weight precision Raw weight size
16-bit 6 GB
8-bit 3 GB
4-bit 1.5 GB

Xiaomi also demonstrated an O3 + O100 dual-chip prototype, with O3 handling system-level work and O100 handling large-model inference.

Xuanwu high-bandwidth matrix bus

O100 uses Xiaomi's Xuanwu high-bandwidth matrix bus to connect the NPU cores and memory subsystem.

It supports parallel processing across the 14 NPU cores, dynamic routing and high-throughput data movement between compute and the stacked DRAM.

120B + 3B dual-model AI Cube

The AI Cube runs two model sizes locally:

  • 120B-class large model
  • 3B-class small model

Xiaomi describes the system as a fast/slow dual-model design. The smaller model handles fast-response workloads while the 120B model provides the larger parameter budget.

Weight precision Approx. raw weight size
BF16 / FP16 240 GB
8-bit 120 GB
6-bit 90 GB
5-bit 75 GB
4-bit 60 GB
3-bit 45 GB

A 120B model at 4-bit uses about 60 GB for weights. Inference also uses memory for KV cache, runtime buffers and activations.

AI Cube compared with other local-AI systems

The closest product-class comparisons are compact high-memory systems such as NVIDIA DGX Spark and AMD Ryzen AI Halo. Mac Studio and discrete GPU systems provide useful bandwidth and memory reference points.

Platform Key local-AI specifications
Xiaomi AI Cube O3 + O100 + D100; 120B + 3B local models; O100 at 1.22 TB/s; up to 150 W sustained
NVIDIA DGX Spark GB10 Grace Blackwell; 128 GB coherent LPDDR5X; 273 GB/s; up to 1 PFLOP FP4; 140 W GB10 TDP; models up to 200B
AMD Ryzen AI Halo Ryzen AI Max+ 395; 128 GB LPDDR5X; 256 GB/s; Radeon 8060S; XDNA 2; 120 W
Apple Mac Studio M3 Ultra Up to 512 GB unified memory; 819 GB/s; up to 32-core CPU and 80-core GPU
RTX 5090 32 GB GDDR7; 1.792 TB/s; 575 W board power
Radeon AI PRO R9700 32 GB GDDR6; 640 GB/s; 191 TFLOPS peak FP16 matrix; 300 W board power

Four RTX 5090s

A four-card RTX 5090 workstation provides:

  • 128 GB aggregate VRAM
  • 7.168 TB/s aggregate local VRAM bandwidth
  • 2,300 W total GPU board power

Each card has its own 32 GB VRAM pool, with multi-GPU model partitioning handled across PCIe.

Four Radeon AI PRO R9700s

A four-card R9700 workstation provides:

  • 128 GB aggregate GDDR6 VRAM
  • 2.56 TB/s aggregate local memory bandwidth
  • 764 TFLOPS aggregate peak FP16 matrix throughput
  • 1,200 W total GPU board power

R9700 is a general-purpose workstation GPU for ROCm and professional GPU workloads, while O100 is a dedicated high-bandwidth inference accelerator.

2027 rollout

Xring O100 is planned for commercial use in 2027. Xring D100 is also scheduled for 2027, while O3 reaches the Xiaomi 18 Fold and Pad 9 Pro Max in September 2026.

Chip Role
Xring O3 Flagship phone and tablet SoC
Xring O100 High-bandwidth local AI accelerator
Xring D100 Intelligent-driving and large-memory AI processor

The AI Cube is the first public system combining all three.

Read next: Xiaomi Xring O3: full specifications, LPDDR6 and benchmarks and Xiaomi Xring D100: 3nm, 160GB unified memory and 200B models.