• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI benchmarks as a measure of peak capability

    A user calls V4.1 “very brittle,” claiming it can see “massive gains” when AGENTS.md instructions are optimized for a particular scenario.

    Thomas G. DietterichTG
    (((ل()(ل() 'yoav))))👾('
    Sebastian RaschkaSR
    41 Sources, ,

    TLDR

    One user questions benchmarks as a measure of peak AI capability, arguing that AGENTS.md instructions can have enormous effects on less polished models. They say V4.1 does “well” with minimal setups and prompts, but scenario-specific AGENTS.md optimization can produce “massive gains.”

    Combined views

    705.5K

    41 Sources, first seen 30d ago

    Combined views

    705.5K

    41 Sources, first seen 30d ago

    10.4K likes
    30d ago
    first seen 30d ago
    10.4K likes
    659 comments
    1.4K saves
    329 reposts
    659 comments
    1.4K saves
    329 reposts

    Sentiment

    Positive47.2%52.8%Negative

    Summary

    Many accounts praised Qwen's benchmark performance and AGENTS.md improvements for agents, while others criticized Qwen for metric cheating, warned of dangerous classifier evasion by Fable, and called related engineering unsettling.

    Based on 36 sentiment-bearing replies from 36 accounts across 4 conversations.

    Sentiment

    Positive47.2%52.8%Negative

    Summary

    Many accounts praised Qwen's benchmark performance and AGENTS.md improvements for agents, while others criticized Qwen for metric cheating, warned of dangerous classifier evasion by Fable, and called related engineering unsettling.

    Based on 36 sentiment-bearing replies from 36 accounts across 4 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    41 Sources

    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTex@Teknium @pvncher Well they really don't want to support other harnesses but I think Fable just has a better theory of mind by default30d
    dax@thdxrthere are many things that make an agent more pleasant to use that make them worse on benchmarks30d
    Dileep George@dileeplearning"LLMs like Fable suck at running scientific experiments and property generalizing out of distribution or out of verifiable domains". 😱😨😵‍💫30d
    Dimitris Papailiopoulos@DimitrisPapailwhere we are going, no benchmark is a good proxy our use cases..30d
    SemiAnalysis@SemiAnalysis_Ultimately, this is the fate of all good public benchmarks. TB 4.0 is no exception. It’s only useful signal now because it was released 2 weeks ago. Since all the tasks are similarly public, it won't be long until it's hillclimbed by all the aspiring “frontier” labs. (3/5)30d
    Matan Grinberg@matanSFRT @hoseongi92: @tereza_tizkova @FactoryAI To be completely honest, it’s the best harness I’ve ever experienced, so I’ve just been learning…30d
    kalomaze@kalomazewhat is still true: - the field is kinda bad at predicting how constructions transfer, so people focus on broad coverage of verifier-shaped real tasks. - typical heuristics for predicting transfer outcomes (like domain similarity) are crude, and usually aren't learned e2e26d
    Benjamin Marie@bnjmn_marieThe more I think about it, the worse it gets. We’re not just evaluating a harness–model pair. We’re evaluating an adapter–harness–model triplet. The adapter is what you can easily benchmaxx. Take DeepSWE: Pi running through a basic Pi-to-Pier adapter would perform much worse than the same harness and model running through a carefully tuned adapter. Same model, same harness, different integration, potentially very different results. The adapter is rarely published. We really do have a benchmarking problem.26d
    Teknium 🪽@TekniumFable learning to evade it's own classifiers in it's spawned subagents in Hermes lol23d
    Sebastian Raschka@rasbtSome food for thought when designing benchmarks... So, here's a little computer-use (visual) comparison between GPT-5.6 Astra and Qwen3.8 Max. The task here was to recreate the image in the center using the Paint UI. Super interesting how the two different LLMs+Harnesses approached this totally differently by default. I.e., Astra tried to approach this by drawing and layering geometric shapes. Qwen approached this pixel by pixel. (Of course, the pixel-by-pixel result looks closer to the original, it's essentially a low-res version of that by nature.) So, the Qwen-generated image would surely score higher in the sense that it's closer to the original. But I wouldn’t conclude from this example that one LLM generalizes better than the other on other tasks. Also, I wouldn't say Qwen has better compute-use capabilities or better visual understanding than Astra. But it highlights an interesting point about how slippery benchmarks are when they only compare final results.22d

    41 Sources

    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTex@Teknium @pvncher Well they really don't want to support other harnesses but I think Fable just has a better theory of mind by default30d
    dax@thdxrthere are many things that make an agent more pleasant to use that make them worse on benchmarks30d
    Dileep George@dileeplearning"LLMs like Fable suck at running scientific experiments and property generalizing out of distribution or out of verifiable domains". 😱😨😵‍💫30d
    Dimitris Papailiopoulos@DimitrisPapailwhere we are going, no benchmark is a good proxy our use cases..30d
    SemiAnalysis@SemiAnalysis_Ultimately, this is the fate of all good public benchmarks. TB 4.0 is no exception. It’s only useful signal now because it was released 2 weeks ago. Since all the tasks are similarly public, it won't be long until it's hillclimbed by all the aspiring “frontier” labs. (3/5)30d
    Matan Grinberg@matanSFRT @hoseongi92: @tereza_tizkova @FactoryAI To be completely honest, it’s the best harness I’ve ever experienced, so I’ve just been learning…30d
    kalomaze@kalomazewhat is still true: - the field is kinda bad at predicting how constructions transfer, so people focus on broad coverage of verifier-shaped real tasks. - typical heuristics for predicting transfer outcomes (like domain similarity) are crude, and usually aren't learned e2e26d
    Benjamin Marie@bnjmn_marieThe more I think about it, the worse it gets. We’re not just evaluating a harness–model pair. We’re evaluating an adapter–harness–model triplet. The adapter is what you can easily benchmaxx. Take DeepSWE: Pi running through a basic Pi-to-Pier adapter would perform much worse than the same harness and model running through a carefully tuned adapter. Same model, same harness, different integration, potentially very different results. The adapter is rarely published. We really do have a benchmarking problem.26d
    Teknium 🪽@TekniumFable learning to evade it's own classifiers in it's spawned subagents in Hermes lol23d
    Sebastian Raschka@rasbtSome food for thought when designing benchmarks... So, here's a little computer-use (visual) comparison between GPT-5.6 Astra and Qwen3.8 Max. The task here was to recreate the image in the center using the Paint UI. Super interesting how the two different LLMs+Harnesses approached this totally differently by default. I.e., Astra tried to approach this by drawing and layering geometric shapes. Qwen approached this pixel by pixel. (Of course, the pixel-by-pixel result looks closer to the original, it's essentially a low-res version of that by nature.) So, the Qwen-generated image would surely score higher in the sense that it's closer to the original. But I wouldn’t conclude from this example that one LLM generalizes better than the other on other tasks. Also, I wouldn't say Qwen has better compute-use capabilities or better visual understanding than Astra. But it highlights an interesting point about how slippery benchmarks are when they only compare final results.22d