• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    OpenAI Engineer Endorses Astra Over Flawed Evals

    OpenAI research engineer Ted Sanders posted that evals often miss practical usefulness while endorsing Astra.

    TS
    1 Source, 27d ago, first seen 27d ago

    TLDR

    Ted Sanders, a research engineer at OpenAI, wrote that model evaluations frequently prove sloppy and fail to capture what matters in actual use. He gave the example of asking clarifying questions, which can lower scores even when that behavior proves helpful in real interactions. Sanders stated that Astra stands apart as the real deal. The model remains imperfect, he noted, yet it has crossed his personal trust threshold in a way no earlier model achieved. His comments appear in a single public post and reflect his direct assessment rather than any broader announcement from OpenAI.

    Combined views

    7.3K

    1 Source, first seen 27d ago

    Combined views

    7.3K

    1 Source, first seen 27d ago

    249 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    249 likes
    14 comments
    17 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    14 comments
    17 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @sandersteddon't trust evals. they're often sloppy and don't measure what matters. little things like asking clarifying questions can tank a score, even if helpful in real life. But Astra is the real deal - it's imperfect, but it's crossing my trust threshold like no model before it.

    1 Source

    @sandersteddon't trust evals. they're often sloppy and don't measure what matters. little things like asking clarifying questions can tank a score, even if helpful in real life. But Astra is the real deal - it's imperfect, but it's crossing my trust threshold like no model before it.