• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Astra and Sol Top Fidelity Benchmark

    Pseudonymous developer @shinboson states Astra and Sol alone earned full marks on his fidelity benchmark.

    𝞍S
    1 Source, 25d ago, first seen 25d ago

    TLDR

    @shinboson posted that Astra along with Sol are the only western frontier models to receive full marks on his fidelity bench. The benchmark measures how often models refuse a request or argue that an unrequested alternative is better. The tweet includes an attached dark-themed table. The account belongs to a pseudonymous AI developer known for creating LLM benchmarks and the OpenPlanter project.

    Combined views

    2.4K

    1 Source, first seen 25d ago

    Combined views

    2.4K

    1 Source, first seen 25d ago

    57 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    57 likes
    14 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    14 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @shinbosonAstra along with Sol are the only western frontier models to get full marks on my fidelity bench, which measures how often the model refuses to do what you ask it or tries to persuade you that the different thing it did that you didn't ask for was actually *better* than what you asked for and you should be grateful https://model-pareto-frontier.pages.dev

    1 Source

    @shinbosonAstra along with Sol are the only western frontier models to get full marks on my fidelity bench, which measures how often the model refuses to do what you ask it or tries to persuade you that the different thing it did that you didn't ask for was actually *better* than what you asked for and you should be grateful https://model-pareto-frontier.pages.dev