• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Santiago Valdarrama on Three Model Selection Mistakes

    Santiago Valdarrama outlines three fixed strategies for calling models in apps.

    RP
    SA
    TH
    5 Sources, 29d ago, first seen 29d ago

    TLDR

    Santiago Valdarrama posted advice on building AI apps. He listed three strategies that lead to trouble: always calling the best model, which becomes very expensive; always calling the cheapest model, which he called dumb; and always calling a mid model, which produces mediocre results. Valdarrama stated the proper solution is to use a router that decides which model to call for each request. The post comes from the computer scientist and ML educator who teaches production AI/ML engineering.

    Combined views

    382.9K

    5 Sources, first seen 29d ago

    Combined views

    382.9K

    5 Sources, first seen 29d ago

    862 likes
    862 likes
    80 comments
    395 saves
    132 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    80 comments
    395 saves
    132 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @tomas_hkToday we’re releasing our methodology for evaluating model routing with interactive benchmarks, which represent agent cost accumulation better than static benchmarks do. Across leading benchmarks, we achieve Pareto-dominance, exceeding Opus xhigh quality at 20–80% lower cost.
    @rohanpaul_aiNot Diamond just released the methodology behind their model routing which gets Opus xhigh quality while cutting agent costs 20–80%. Model routing becomes much harder decision for long-running coding agents. So Not Diamond is benchmarking the router across simulated user turns, changing task complexity, and response delays that can expire the KV cache. And harder than it seems, switch models too aggressively and you can throw away the KV cache. Pick one model from the opening prompt and you are assuming task complexity stays fixed for the rest of the session. So Not Diamond is treating routing as a sequential decision problem instead. At each step, its router predicts future reward and cost for a particular model and reasoning effort, using current and previous session state, message and token counts, task complexity, KV-cache state, and intermediate reward signals. Read more detail on their technical report.
    @svpinoThree ways you get in trouble when building an app: 1. Always calling the best model 2. Always calling the cheapest model 3. Always calling a mid model The first strategy will be very expensive. The second strategy is just dumb. The third strategy is just mediocre. The proper solution is to use a router and decide which model is best for a particular request. Check out these guys. Their platform lets you select a model and a specific reasoning effort for each request. This keeps the quality of the best models while costing much, much less.

    5 Sources

    @tomas_hkToday we’re releasing our methodology for evaluating model routing with interactive benchmarks, which represent agent cost accumulation better than static benchmarks do. Across leading benchmarks, we achieve Pareto-dominance, exceeding Opus xhigh quality at 20–80% lower cost.
    @rohanpaul_aiNot Diamond just released the methodology behind their model routing which gets Opus xhigh quality while cutting agent costs 20–80%. Model routing becomes much harder decision for long-running coding agents. So Not Diamond is benchmarking the router across simulated user turns, changing task complexity, and response delays that can expire the KV cache. And harder than it seems, switch models too aggressively and you can throw away the KV cache. Pick one model from the opening prompt and you are assuming task complexity stays fixed for the rest of the session. So Not Diamond is treating routing as a sequential decision problem instead. At each step, its router predicts future reward and cost for a particular model and reasoning effort, using current and previous session state, message and token counts, task complexity, KV-cache state, and intermediate reward signals. Read more detail on their technical report.
    @svpinoThree ways you get in trouble when building an app: 1. Always calling the best model 2. Always calling the cheapest model 3. Always calling a mid model The first strategy will be very expensive. The second strategy is just dumb. The third strategy is just mediocre. The proper solution is to use a router and decide which model is best for a particular request. Check out these guys. Their platform lets you select a model and a specific reasoning effort for each request. This keeps the quality of the best models while costing much, much less.