SGLang и RadixArk выпустили открытый RL-фреймворк Miles v0.1
SGLang и RadixArk представили Miles v0.1 — открытый фреймворк для обучения с подкреплением и постобучения языковых и мультимодальных моделей. Он ориентирован на отладку, эффективность оборудования и масштабирование рабочих процессов.
SGLang и RadixArk выпустили Miles v0.1 — открытый фреймворк для обучения с подкреплением при постобучении больших языковых и мультимодальных моделей. Проект должен помочь командам проверять корректность запусков, эффективнее использовать вычислительные ресурсы и масштабировать RL-задачи.
Согласно описанию проекта, Miles объединяет высокопроизводительные запуски моделей через SGLang с масштабируемым обучением на базе Megatron-LM. Фреймворк поддерживает асинхронные RL-сценарии, настраиваемые режимы on-policy и off-policy, а также оборудование NVIDIA и AMD. Заявления о числе участников, охвате моделей и использовании в производственных задачах в основном исходят от команды проекта и её партнёров.
Источники
GitHub - radixark/miles: Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime. · GitHubgithub.com · supporting[2026/08] 🔥 Miles v0.1 is released! Read the blog post here: Miles v0.1: Production-level Post-training. [2026/07] Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles (blog). [2026/07] 🔥 SGLang and Miles add day-0 support for Kimi K3 (blog). [2026/07] On-policy distillation lands in Miles (blog). [2026/07] 🔥 SGLang and Miles add day-0 support for Inkling, a frontier multimodal model (blog). [2026/07] DeepSeek-V4 Flash RL training comes to AMD Instinct MI355X with Miles (blog). [2026/06] SGLang and Miles add day-0 support for NVIDIA Nemotron 3 Ultra (blog). [2026/05] No token left behind: token-in-token-out in Miles (blog). [2026/04] Updating 1 T parameters in seconds: P2P weight transfer in large-scale distributed RL (blog). [...] Fully async RL. Rollout and training workers are decoupled, with configurable on- and off-policy schedules, a pipeline tuned for fewer bubbles, and customizable async rollout and eval modes. See Fully Async RL. Fast
miles/README.md at main · radixark/miles · GitHubgithub.com · supporting[2026/08] 🔥 Miles v0.1 is released! Read the blog post here: Miles v0.1: Production-level Post-training. [2026/07] Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles (blog). [2026/07] 🔥 SGLang and Miles add day-0 support for Kimi K3 (blog). [2026/07] On-policy distillation lands in Miles (blog). [2026/07] 🔥 SGLang and Miles add day-0 support for Inkling, a frontier multimodal model (blog). [2026/07] DeepSeek-V4 Flash RL training comes to AMD Instinct MI355X with Miles (blog). [2026/06] SGLang and Miles add day-0 support for NVIDIA Nemotron 3 Ultra (blog). [2026/05] No token left behind: token-in-token-out in Miles (blog). [2026/04] Updating 1 T parameters in seconds: P2P weight transfer in large-scale distributed RL (blog). [...] ## About Miles is a high-performance, enterprise-ready reinforcement learning framework for large-scale model post-training. It pairs SGLang for high-throughput rollout with Megatron-LM for scalable training, and sh
🚀 Celebrating Partnership Success and Innovation! 🎉 It's ...instagram.com · supportingA huge congratulations to our long-term partner SGLang/RadixArk on the official launch of Miles v0.1! From the M-Series and H3 to Music 3
Yutong Wang's Post - LinkedInlinkedin.com · supportingToday we're launching Miles v0.1, an open-source RL framework for LLMs and multimodal models. RadixArk's mission is to make AI
Guohao Li 🐫 on X: "Huge congrats to the @radixark team on the launch of Miles v0.1! We've been using Miles to train RL agents for terminal coding and ML engineering, and to validate our RL environments across a wide range of training setups, both for open-source research like our SETA project (ht… / Xx.com · supporting## Post ## Post # Guohao Li 🐫 on X: "Huge congrats to the @radixark team on the launch of Miles v0.1! We've been using Miles to train RL agents for terminal coding and ML engineering, and to validate our RL environments across a wide range of training setups, both for open-source research like our SETA project ( and for our lab clients. So far we've trained across thousands of environments with models like DeepSeek V4 Flash, using Miles together with SGLang. Its efficient async RL support lets us iterate faster and focus on improving the quality of our agents and RL environment data. The Miles team has also been incredibly supportive throughout, an amazingly talented group to work with!" @guohao_li @radixark Avatar See this post in the app [...] @guohao_li @radixark Avatar See this post in the app Use the app to view all comments and discover more posts. @guohao_li @radixark @mingyilu123 @zz30gs @BanghuaZ
MiniMax (official) on X: "Congrats to our long-term partner SGLang/RadixArk on the launch of Miles v0.1! 🎉 From M-Series to H3 and Music 3, we’ve been closely building and pushing the OSS ecosystem forward together. Nothing beats the feeling of seeing things we built together keep growing @radix… / Xx.com · supporting## Post ## Post # MiniMax (official) on X: "Congrats to our long-term partner SGLang/RadixArk on the launch of Miles v0.1! 🎉 From M-Series to H3 and Music 3, we’ve been closely building and pushing the OSS ecosystem forward together. Nothing beats the feeling of seeing things we built together keep growing @radixark 🤝 @MiniMax\_AI" @MiniMax_AI @radixark ## Log in or sign up for X See what’s happening and join the conversation ## Relevant people Avatar ## Trending now @MiniMax_AI @radixark @zz30gs @kiri49x86 @eva_morgan_ai