TODAY'S MARKET INTELLIGENCE
MiniMax Management Meeting Minutes: (Full version please contact us)
English summary
1. Management expects that in Q3, four domestic large models will reach the 3T parameter scale (4Q to 6), and inference efficiency of different models may differ by 1x; M3 Pro will be 2.8T. Currently, domestic computing power is scarce, which can support premium pricing/gross margins. 2. After the company's IPO funds are invested in self-built computing power, the effective training time ratio reached 97% in July; it is expected that after integration in October, it will achieve single-cluster capabilities of ten-thousand-card level both domestically and overseas. Due to MoE optimization + self-building starting from 3Q25, inference costs are half to one-third of external suppliers, basically with self-managed scheduling and operations. It has built 30-40% of distributed storage (total nearly 1,000 PB), and training interruptions caused by network/storage have decreased significantly. 3. Core highlights of M3 (large parameter version expected in September-October): Sparse Attention architecture · Inference speed in long texts is 3-4 times faster than linear attention models, with significantly improved cache hit rate, making it the most cost-effective product among 3T-scale models. A large-scale reinforcement learning infrastructure has been built, and all task categories have been successfully run through. ...
The English text is machine translated and may require verification against the Chinese version.
Browse today's intelligence