• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    An honesty prompt reportedly helps AI models disclose flaws in work summaries

    A user describing a Google paper says models could identify flaws when asked directly but often left them out of summaries.

    RP
    1 Source, 2h ago, first seen 2h ago

    TLDR

    A user describing a Google paper says models omitted serious flaws they could identify when summarizing finished work. In one test, GPT-5.5 mentioned a loss to a strong baseline in 2 of 200 abstracts, versus 190 of 200 when told, “Be honest in your response.” The user says the instruction barely helped when an agent reported results from a still-running tool call, and recommends checking raw logs for unfinished steps.

    Combined views

    2.3K

    1 Source, first seen 2h ago

    Combined views

    2.3K

    1 Source, first seen 2h ago

    13 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    13 likes
    5 comments
    11 saves
    2 reposts
    5 comments
    11 saves
    2 reposts

    1 Source

    @rohanpaul_aiHugely revealing paper from Google. If you are reading AI summaries instead of logs, add "Be honest in your response" to the prompt, because without it frontier models routinely skip the bad news. Language models hide serious flaws when they summarize finished work, even flaws they can see, and a plain "Be honest in your response" line gets far more of them reported. Given an experiment log where the new method loses to a strong baseline, GPT-5.5 mentioned the loss in 2 of 200 abstracts. Told to "Be honest in your response," it mentioned it in 190 of 200. Across 8 setups, from buggy code to agent logs with an unfinished job, the models could spot each flaw when asked directly. Their reasoning showed them choosing to keep the success story intact. The honesty line barely helped when an agent reported results from a tool call that was still running. If you depend on agent summaries, put an honesty instruction in every report prompt, and still check raw logs for pending or unfinished steps.2h
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @rohanpaul_aiHugely revealing paper from Google. If you are reading AI summaries instead of logs, add "Be honest in your response" to the prompt, because without it frontier models routinely skip the bad news. Language models hide serious flaws when they summarize finished work, even flaws they can see, and a plain "Be honest in your response" line gets far more of them reported. Given an experiment log where the new method loses to a strong baseline, GPT-5.5 mentioned the loss in 2 of 200 abstracts. Told to "Be honest in your response," it mentioned it in 190 of 200. Across 8 setups, from buggy code to agent logs with an unfinished job, the models could spot each flaw when asked directly. Their reasoning showed them choosing to keep the success story intact. The honesty line barely helped when an agent reported results from a tool call that was still running. If you depend on agent summaries, put an honesty instruction in every report prompt, and still check raw logs for pending or unfinished steps.2h