• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    MiMo-V2.6 reportedly scales reinforcement learning to about 2 billion tokens per step

    In a September 16 update, the team behind MiMo-V2.6 said it was scaling compute, agent environments and grading, with plans to open-source details in stages.

    Percy LiangPL
    Lucas Beyer (bl16)LB
    Nathan LambertNL
    38 Sources, ,

    TLDR

    The MiMo-V2.6 team said on September 16 that its reinforcement learning run was underway, using roughly 2 billion tokens per step in a fully asynchronous setup with 1,568 prompts × 16 rollouts. It said the run combines multiple agent tasks and harnesses, with rewards based on test cases and rubrics. The team shared a link to stream the run and said it would open-source details piece by piece over the following weeks.

    Combined views

    4.2M

    38 Sources, first seen 21d ago

    Combined views

    4.2M

    38 Sources, first seen 21d ago

    19.7K likes
    21d ago
    first seen 21d ago
    19.7K likes
    654 comments
    7.8K saves
    3K reposts
    654 comments
    7.8K saves
    3K reposts

    Sentiment

    Positive89.8%10.2%Negative

    Based on 361 sentiment-bearing replies from 334 accounts across 11 conversations.

    Sentiment

    Positive89.8%10.2%Negative

    Based on 361 sentiment-bearing replies from 334 accounts across 11 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    38 Sources

    Fuli Luo@_LuoFuliNearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: https://mimo.xiaomi.com/rl/21d
    Han Xiao@hxiao1. union-alpha is not from xiaomi 2. bold & open to livestream the entire training pipeline 3. should send this gif to cfo21d
    Andrew Curran@AndrewCurran_'We believe RL is one of the most scalable and efficient paths toward self-improvement.'21d
    Brendan O'Donoghue@bodonoghue85RT @_LuoFuli: Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL…21d
    Lisan al Gaib@scaling01this is fucking sick imagine OpenAI and Anthropic had this21d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexReally amazing, we get the baseline for advanced RL costs21d
    Lucas Beyer (bl16)@giffmanaSo they say "three things we scaled". Let me translate: 1. "compute" - yep, that's compute. 2. "environments and harnesses" - actually, also compute. 3. "and grader compute" - you guessed it, that's also compute. joke aside, pretty cool to see their public live dashboard, including "cost so far" (!) (due to recent events/discussions: no, this QT is not sponsored. I don't do sponsor stuff, I'm here for the fun.)21d
    elie@eliebakouchwow insane, they are literally livestreaming the RL training run of Mimo V2.6 Pro (1T, 42B active) and Flash (309B, 15B active) with per batch data/harness composition and a ton of internal training metrics https://mimo.xiaomi.com/rl/21d
    Nathan Lambert@natolambertOne of the coolest at-scale RL resources made public yet! You love to see it.21d
    wh@nrehiew_How much it costs for an RL hero run on a large open-source model (the pro model) - Batch size 25K - $65K per step - 2 hours per step - Average 4K input, 80K output21d

    38 Sources

    Fuli Luo@_LuoFuliNearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: https://mimo.xiaomi.com/rl/21d
    Han Xiao@hxiao1. union-alpha is not from xiaomi 2. bold & open to livestream the entire training pipeline 3. should send this gif to cfo21d
    Andrew Curran@AndrewCurran_'We believe RL is one of the most scalable and efficient paths toward self-improvement.'21d
    Brendan O'Donoghue@bodonoghue85RT @_LuoFuli: Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL…21d
    Lisan al Gaib@scaling01this is fucking sick imagine OpenAI and Anthropic had this21d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexReally amazing, we get the baseline for advanced RL costs21d
    Lucas Beyer (bl16)@giffmanaSo they say "three things we scaled". Let me translate: 1. "compute" - yep, that's compute. 2. "environments and harnesses" - actually, also compute. 3. "and grader compute" - you guessed it, that's also compute. joke aside, pretty cool to see their public live dashboard, including "cost so far" (!) (due to recent events/discussions: no, this QT is not sponsored. I don't do sponsor stuff, I'm here for the fun.)21d
    elie@eliebakouchwow insane, they are literally livestreaming the RL training run of Mimo V2.6 Pro (1T, 42B active) and Flash (309B, 15B active) with per batch data/harness composition and a ton of internal training metrics https://mimo.xiaomi.com/rl/21d
    Nathan Lambert@natolambertOne of the coolest at-scale RL resources made public yet! You love to see it.21d
    wh@nrehiew_How much it costs for an RL hero run on a large open-source model (the pro model) - Batch size 25K - $65K per step - 2 hours per step - Average 4K input, 80K output21d