• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Sonnet 5 Matches Opus 5 Speed and Cost on Workflow

    SaaStr founder Jason Lemkin shares results from his model evaluation test.

    J✨
    1 Source, 29d ago, first seen 29d ago

    TLDR

    Jason Lemkin posted results from his first evaluation of cheaper models versus Sonnet 5 on a complex Connect workflow. The task required completion in under 30 seconds and under $0.25. He stated that Sonnet 5 performed as well as Opus 5 while running faster and cheaper, meeting both criteria for speed and cost with no loss in quality or acceptance. Lemkin concluded there was no need to use Opus. The tweet included a clean table screenshot showing the comparison.

    Combined views

    7.7K

    1 Source, first seen 29d ago

    Combined views

    7.7K

    1 Source, first seen 29d ago

    12 likes
    12 likes
    16 comments
    6 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    16 comments
    6 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @jasonlkRan my first eval of cheaper models vs Sonnet 5 for a complex Connect workflow. Has to finish in under 30 seconds and stay well under $0.25. Sonnet 5 already performed as well as Opus 5 here and is faster and cheaper so no need to use Opus. Goal was faster, cheaper, same quality / acceptance: Only Sonnet met both criteria (speed + cost) and quality: - Grok completed but took 166 seconds a pair. Disqualifying. - Gemini Flash timed out at 45s every attempt. - Haiku, DeepSeek and Qwen all returned valid output and failed the validator the same way: citing the wrong side’s facts Will do more, add more models, I guess tune the prompts for each model (which I didn’t want to do … but I guess need to do) but didn’t get the quick win moving off Sonnet 5 was hoping for Two learnings: #1. Speed eliminates more models than cost does #2. And the cheap ones just cited the wrong facts. But against a validator we’d tuned around Sonnet ... rebuilding the validator to handle many models is a project. Maybe the meta learning even for us is yes, tuning the harness around a SOTA LLM is most important of all. And doing this right is a lot of work.

    1 Source

    @jasonlkRan my first eval of cheaper models vs Sonnet 5 for a complex Connect workflow. Has to finish in under 30 seconds and stay well under $0.25. Sonnet 5 already performed as well as Opus 5 here and is faster and cheaper so no need to use Opus. Goal was faster, cheaper, same quality / acceptance: Only Sonnet met both criteria (speed + cost) and quality: - Grok completed but took 166 seconds a pair. Disqualifying. - Gemini Flash timed out at 45s every attempt. - Haiku, DeepSeek and Qwen all returned valid output and failed the validator the same way: citing the wrong side’s facts Will do more, add more models, I guess tune the prompts for each model (which I didn’t want to do … but I guess need to do) but didn’t get the quick win moving off Sonnet 5 was hoping for Two learnings: #1. Speed eliminates more models than cost does #2. And the cheap ones just cited the wrong facts. But against a validator we’d tuned around Sonnet ... rebuilding the validator to handle many models is a project. Maybe the meta learning even for us is yes, tuning the harness around a SOTA LLM is most important of all. And doing this right is a lot of work.