• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI models reportedly score higher through APIs than chatbot interfaces

    Researchers auditing seven systems across ChatGPT, Claude and Gemini say adjusting system prompts, sampling parameters and reasoning settings did not reliably reproduce chatbot behavior through the API.

    CM
    DH
    CL
    6 Sources, ,

    TLDR

    The research team reports that, across seven systems and nine benchmarks, the same models scored about 3.4 points higher through APIs—which let software interact with models—than through chatbot interfaces. The researchers argue that evaluation reports should specify how a model was accessed, alongside its name and the date. They warn that an API audit may certify a different system from the one most people use.

    Combined views

    34K

    6 Sources, first seen 15d ago

    Combined views

    34K

    6 Sources, first seen 15d ago

    204 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    15d ago
    first seen 15d ago
    204 likes
    20 comments
    66 saves
    81 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    20 comments
    66 saves
    81 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 Sources

    @jennjwangThird-party auditors often audit models via the API, but do those findings transfer to chatbot interfaces? We audited seven systems across ChatGPT, Claude, and Gemini and found that they don’t. (EMNLP 2026)
    @ChengleiSiRT @jennjwang: Third-party auditors often audit models via the API, but do those findings transfer to chatbot interfaces? We audited seven…
    @chrmanningRT @jennjwang: Third-party auditors often audit models via the API, but do those findings transfer to chatbot interfaces? We audited seven…
    @sanmikoyejoEver noticed your LLM behaving differently depending on how you reach it? Turns out the access surface matters, and more than I'd assumed. Across 7 systems and 9 benchmarks, the same models score about 3.4 points higher through the API than through the chatbot interface. We tried to close the gap from the API side: system prompts, sampling parameters, reasoning settings. Nothing reliably reproduced interface behavior. Whatever sits between the endpoint and the deployed product isn't something an auditor can reconstruct. This matters for eval reports, which should include the access surface alongside the model name and date. This also matters for audits: access granted at the endpoint may certify a different system from the one most people use. Great work by @jennjwang with @joabaum, Dan Ho, and me (EMNLP 2026): https://arxiv.org/abs/2609.08861
    @dhadfieldmenellRT @sanmikoyejo: Ever noticed your LLM behaving differently depending on how you reach it? Turns out the access surface matters, and more t…

    6 Sources

    @jennjwangThird-party auditors often audit models via the API, but do those findings transfer to chatbot interfaces? We audited seven systems across ChatGPT, Claude, and Gemini and found that they don’t. (EMNLP 2026)
    @ChengleiSiRT @jennjwang: Third-party auditors often audit models via the API, but do those findings transfer to chatbot interfaces? We audited seven…
    @chrmanningRT @jennjwang: Third-party auditors often audit models via the API, but do those findings transfer to chatbot interfaces? We audited seven…
    @sanmikoyejoEver noticed your LLM behaving differently depending on how you reach it? Turns out the access surface matters, and more than I'd assumed. Across 7 systems and 9 benchmarks, the same models score about 3.4 points higher through the API than through the chatbot interface. We tried to close the gap from the API side: system prompts, sampling parameters, reasoning settings. Nothing reliably reproduced interface behavior. Whatever sits between the endpoint and the deployed product isn't something an auditor can reconstruct. This matters for eval reports, which should include the access surface alongside the model name and date. This also matters for audits: access granted at the endpoint may certify a different system from the one most people use. Great work by @jennjwang with @joabaum, Dan Ho, and me (EMNLP 2026): https://arxiv.org/abs/2609.08861
    @dhadfieldmenellRT @sanmikoyejo: Ever noticed your LLM behaving differently depending on how you reach it? Turns out the access surface matters, and more t…