• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    The distinction between harder AI tests and more realistic ones

    The post argues that more realistic tests often make answers harder to judge: real work involves deciding what problem to solve, which constraints matter and what a useful answer looks like.

    SH
    1 Source, 17d ago, first seen 17d ago

    TLDR

    A post about AI evaluation explores a tradeoff: tests that better reflect real work often become harder to grade confidently. It suggests science may offer a way around that problem, because predictions—such as how much drag a wing produces—can be checked against physical measurements. But the author notes that those measurements are expensive. Using a fluid-dynamics simulation instead introduces a distinction the post highlights: verification asks whether the chosen equations are solved correctly; validation asks whether the model adequately represents the actual flow for its intended use.

    Combined views

    3.3K

    1 Source, first seen 17d ago

    Combined views

    3.3K

    1 Source, first seen 17d ago

    34 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    34 likes
    2 comments
    11 saves
    3 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 comments
    11 saves
    3 reposts

    1 Source

    @shyamalanadkatin his recent conversation with @dwarkesh_sp, @schuman separates two axes of AI progress: difficulty and realism: a task can become arbitrarily hard while remaining unrealistically well specified. real work often requires figuring out what problem to solve, which constraints matter, and what would count as a useful answer. this is roughly why i helped incubate applied evals at oai. we started in domains where answers are verifiable, then work out how to measure usefulness in the messier settings where people actually do their jobs. but this introduces a tradeoff: as the eval becomes more realistic, it often becomes harder to determine whether the answer is good. we gain fidelity to the work while losing confidence in the grading! science seems to offer a way around this tradeoff. the problems are real and messy, yet physics still provides an external check. if a model predicts that a wing produces a certain amount of drag, we can measure whether it does. we appear to get both realism and verifiability. the difficulty is that physical measurements are expensive. If we evaluate an AI's drag predictions against a CFD simulation, our immediate reference is a numerical model of physics. CFD practitioners distinguish two questions: verification asks whether we are solving the chosen equations correctly; validation asks whether the model adequately represents the actual flow for its intended use. eg: a turbulence model can be implemented correctly and solved accurately while still giving poor predictions near flow separation. an AI evaluated against that simulator could therefore improve its benchmark score by learning to reproduce the simulator's mistakes more faithfully. agreement would increase without a corresponding improvement in our ability to predict the world. maybe the evaluation would have no way to distinguish the two. this is the subtlety in saying "physics is the verifier". we have to establish where it agrees with measurements, how uncertain its predictions are, and whether that uncertainty matters for the decision. a benchmark can measure agreement with its answer key. establishing that the answer key deserves our trust is additional scientific work.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @shyamalanadkatin his recent conversation with @dwarkesh_sp, @schuman separates two axes of AI progress: difficulty and realism: a task can become arbitrarily hard while remaining unrealistically well specified. real work often requires figuring out what problem to solve, which constraints matter, and what would count as a useful answer. this is roughly why i helped incubate applied evals at oai. we started in domains where answers are verifiable, then work out how to measure usefulness in the messier settings where people actually do their jobs. but this introduces a tradeoff: as the eval becomes more realistic, it often becomes harder to determine whether the answer is good. we gain fidelity to the work while losing confidence in the grading! science seems to offer a way around this tradeoff. the problems are real and messy, yet physics still provides an external check. if a model predicts that a wing produces a certain amount of drag, we can measure whether it does. we appear to get both realism and verifiability. the difficulty is that physical measurements are expensive. If we evaluate an AI's drag predictions against a CFD simulation, our immediate reference is a numerical model of physics. CFD practitioners distinguish two questions: verification asks whether we are solving the chosen equations correctly; validation asks whether the model adequately represents the actual flow for its intended use. eg: a turbulence model can be implemented correctly and solved accurately while still giving poor predictions near flow separation. an AI evaluated against that simulator could therefore improve its benchmark score by learning to reproduce the simulator's mistakes more faithfully. agreement would increase without a corresponding improvement in our ability to predict the world. maybe the evaluation would have no way to distinguish the two. this is the subtlety in saying "physics is the verifier". we have to establish where it agrees with measurements, how uncertain its predictions are, and whether that uncertainty matters for the decision. a benchmark can measure agreement with its answer key. establishing that the answer key deserves our trust is additional scientific work.