GPT-6 Astra's smallest TRACES gains were reportedly in error correction
A user describes TRACES as a scientific-discovery benchmark that evaluates an AI's process, arguing that correct answers can hide weak investigations.
TLDR
According to a user, Apodex AI released TRACES to test AI on difficult scientific problems where the correct answer may not already be known. The benchmark separately measures Tools, Repair, Alternatives, Coherence, Evidence and Scope. The user says GPT-6 Astra improved across all six, but Repair—fixing its own mistakes—improved far less than the others. They argue that reaching the right answer while ignoring contradictory feedback could mask behavior that fails on the next task.
Combined views
6.7K
3 Sources, first seen 18d ago