• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology
Announcement

ai& releases TermGrade with 1,004 graded terminal environments for training AI agents

ai& says six models graded the environments and all 36,144 attempts are included.

Yağız ÇalıkYÇ
Hugging FaceHF
4 Sources, 4h ago, first seen 4h ago

TLDR

ai& says TermGrade includes 1,004 terminal environments for training AI agents, graded by six models, with all 36,144 attempts included. It reports that training on tasks a model solves half the time lifted gemma-4-31B-it 3.1 points on Terminal-Bench 2.1. In a reply, a user says the environments, every trajectory—including 14k failures—the exact training split and the checkpoint are open on Hugging Face.

Combined views

18.5K

4 Sources, first seen 4h ago

161 likes14 comments136 saves46 reposts

Combined views

18.5K

4 Sources, first seen 4h ago

161 likes14 comments136 saves46 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

4 Sources

Yağız Çalık@WeyaxiWe're releasing TermGrade at @aiand_: high-quality RL environments for terminal agents, fully open source, along with the full story of how we built and graded them, and the RL recipe we trained with. That's 1k executable environments, each verified by execution and graded against 6 models, plus all 36k trajectories behind the grades, failures included. And training gemma-4-31B-it on the tasks it solved half the time took it +3.1 on Terminal-Bench 2.1. Thread below 🧵4h
Hugging Face@huggingfaceRT @Weyaxi: We're releasing TermGrade at @aiand_: high-quality RL environments for terminal agents, fully open source, along with the full…2h
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    4 Sources

    Yağız Çalık@WeyaxiWe're releasing TermGrade at @aiand_: high-quality RL environments for terminal agents, fully open source, along with the full story of how we built and graded them, and the RL recipe we trained with. That's 1k executable environments, each verified by execution and graded against 6 models, plus all 36k trajectories behind the grades, failures included. And training gemma-4-31B-it on the tasks it solved half the time took it +3.1 on Terminal-Bench 2.1. Thread below 🧵4h
    Hugging Face@huggingfaceRT @Weyaxi: We're releasing TermGrade at @aiand_: high-quality RL environments for terminal agents, fully open source, along with the full…2h
    Today's Rank

    #19

    Today's Rank

    #19