• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    GPT-6 Astra's smallest TRACES gains were reportedly in error correction

    A user describes TRACES as a scientific-discovery benchmark that evaluates an AI's process, arguing that correct answers can hide weak investigations.

    RP
    3 Sources, ,

    TLDR

    According to a user, Apodex AI released TRACES to test AI on difficult scientific problems where the correct answer may not already be known. The benchmark separately measures Tools, Repair, Alternatives, Coherence, Evidence and Scope. The user says GPT-6 Astra improved across all six, but Repair—fixing its own mistakes—improved far less than the others. They argue that reaching the right answer while ignoring contradictory feedback could mask behavior that fails on the next task.

    Combined views

    6.7K

    3 Sources, first seen 18d ago

    Combined views

    6.7K

    3 Sources, first seen 18d ago

    16 likes
    18d ago
    first seen 18d ago
    16 likes
    12 comments
    4 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    12 comments
    4 saves
    2 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 Sources

    @rohanpaul_aiA benchmark for difficult scientific discovery from @tianqiao_chen now shows GPT-6 Astra improving across all six capabilities. But fixing its own mistakes improved the least, Repair improved far less than the other dimensions. For AI agents, a correct final answer can hide a weak investigation. An agent might reach the right result after ignoring contradictory feedback. That behavior could fail on the next task, even though the current answer passes its check. @Apodex_AI released TRACES, a new benchmark for testing AI systems on difficult scientific problems where the correct answer may not already be known. TRACES separately measures Tools, Repair, Alternatives, Coherence, Evidence, and Scope instead of reducing discovery capability to one score. Instead of scoring only the final answer, TRACES also evaluates how the AI works through the problem, including its tool use, error correction, evidence, and reasoning process.

    3 Sources

    @rohanpaul_aiA benchmark for difficult scientific discovery from @tianqiao_chen now shows GPT-6 Astra improving across all six capabilities. But fixing its own mistakes improved the least, Repair improved far less than the other dimensions. For AI agents, a correct final answer can hide a weak investigation. An agent might reach the right result after ignoring contradictory feedback. That behavior could fail on the next task, even though the current answer passes its check. @Apodex_AI released TRACES, a new benchmark for testing AI systems on difficult scientific problems where the correct answer may not already be known. TRACES separately measures Tools, Repair, Alternatives, Coherence, Evidence, and Scope instead of reducing discovery capability to one score. Instead of scoring only the final answer, TRACES also evaluates how the AI works through the problem, including its tool use, error correction, evidence, and reasoning process.