• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Pinocchio aims to estimate confidence in AI answers without access to the model behind them

    Its creators say the lightweight model uses the prompt, answer and target model’s ID to predict whether an answer is correct.

    MG
    12 Sources, 8h ago, first seen 8h ago

    TLDR

    Pinocchio’s creators say it predicts whether a language model’s answer is correct, then calibrates that prediction into an uncertainty estimate in a single pass. They say it handles vision tasks and outperformed verbal confidence baselines and TypeSafe Jev in their tests; Jev was tested only on text. They caution that Pinocchio can be miscalibrated on data unlike its training mix.

    Combined views

    5K

    12 Sources, first seen 8h ago

    Combined views

    5K

    12 Sources, first seen 8h ago

    106 likes
    106 likes
    11 comments
    25 saves
    18 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    11 comments
    25 saves
    18 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    12 Sources

    @micahgoldblumFrontier LLMs like Claude don’t come with uncertainty estimates, and their verbalized confidence is bad. We trained Pinocchio, a lightweight model that assigns confidence to outputs of popular API models, making it fast and easy to get well-calibrated uncertainty estimates. 📄🧵

    12 Sources

    @micahgoldblumFrontier LLMs like Claude don’t come with uncertainty estimates, and their verbalized confidence is bad. We trained Pinocchio, a lightweight model that assigns confidence to outputs of popular API models, making it fast and easy to get well-calibrated uncertainty estimates. 📄🧵