• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

A proposed method for ranking base models for coding agents before post-training

A post about an NVIDIA paper says its three rankings closely matched post-trained SWE-bench Verified scores across ten model pairs.

elvisEL
1 Source, 23m ago, first seen 23m ago

TLDR

A post describing an NVIDIA paper says five of six base models solved zero SWE-bench Verified tasks. The proposed method starts with tasks a strong post-trained agent solved, identifies the first code edit that makes the tests pass, then checks whether a base model can produce or recognize that fix using the preceding context. The post says three resulting rankings closely matched post-trained scores across ten model pairs, offering a way to choose checkpoints before post-training.

Combined views

395

1 Source, first seen 23m ago

9 reposts

Combined views

395

1 Source, first seen 23m ago

9 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

elvis@omarsar0RT @dair_ai: Super interesting NVIDIA paper on choosing base models for coding agents. It's actually a clever way to rank base checkpoints…23m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    1 Source

    elvis@omarsar0RT @dair_ai: Super interesting NVIDIA paper on choosing base models for coding agents. It's actually a clever way to rank base checkpoints…23m
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet