Miles 团队在 Blackwell 架构上实现了两种原生低精度强化学习方案:端到端 MXFP8 和 MoE 专家权重的逐 token NVFP4。在 8x B200 上对 Qwen3-30B-A3B 的消融实验中,BF16 与所有五种低精度配置的原始奖励曲线高度重合,且 MXFP8 和 NVFP4 减少了推理时间。
论文
·LMSYS:Blog(Chatbot Arena 团队)
Miles 在 Blackwell 架构上实现端到端 MXFP8 与逐 token NVFP4 强化学习方案
— Blog Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles TL;DR: We implemented two Blackwell-native RL recipes in Miles: end-to-end MXFP8 and per-token NVFP4 for MoE experts. Both are supported by fine-grained precision control across checkpoint conversion,… Ziang Li, humans& and Miles Team