OpenAI Engineer Endorses Astra Over Flawed Evals
OpenAI research engineer Ted Sanders posted that evals often miss practical usefulness while endorsing Astra.
TLDR
Ted Sanders, a research engineer at OpenAI, wrote that model evaluations frequently prove sloppy and fail to capture what matters in actual use. He gave the example of asking clarifying questions, which can lower scores even when that behavior proves helpful in real interactions. Sanders stated that Astra stands apart as the real deal. The model remains imperfect, he noted, yet it has crossed his personal trust threshold in a way no earlier model achieved. His comments appear in a single public post and reflect his direct assessment rather than any broader announcement from OpenAI.
Combined views
7.3K
1 Source, first seen 27d ago