OpenAI Engineer Shares Astra Benchmark Scores
OpenAI researcher retweets claim of model performance on a browser benchmark.
TLDR
Ted Sanders works as a research engineer at OpenAI where he focuses on large language models, prompting, and model improvements. He retweeted a post by gregpr07. That post claimed Astra just obliterated the hardest benchmark. It showed Browser Use Benchmark v2 results listing Astra medium at 77.3 percent and Opus 5 at 50.5 percent. The message used an eye emoji to highlight the scores. The visible posts consist only of this retweet and the quoted claim. No independent verification or additional context is provided in the packet.
Combined views
5.5K
2 Sources, first seen 25d ago