• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Researcher Urges Stop to GPT-4o Medical Evaluations

    Tanishq Abraham questions reliability of studies using outdated LLMs for medical tasks.

    TM
    MC
    MN
    4 Sources, 39d ago, first seen 39d ago

    TLDR

    Tanishq Mathew Abraham posted a plea to halt use of GPT-4o, Llama-3, and Command-R in medical LLM evaluations. He stated the models were never strong at medical capabilities and that conclusions from studies relying on them cannot be trusted. Mark Cuban retweeted a post describing a randomized control study that asked whether an LLM's diagnostic and management performance holds when patient input is added. A generated summary in the packet reported the study found a 60 percent drop in diagnostic accuracy and a 12 percent decline in appropriate management decisions.

    Combined views

    57.7K

    4 Sources, first seen 39d ago

    Combined views

    57.7K

    4 Sources, first seen 39d ago

    550 likes
    550 likes
    36 comments
    57 saves
    84 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    36 comments
    57 saves
    84 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    @iScienceLuvrPLEASE IM BEGGING YOU TO STOP USING GPT-4O FOR EVALUATIONS On top of that, using Llama-3 and Command-R?! 🤮 LLMs aren't without limitations, but I can't trust any conclusions made by studies using such old models. I mean these models themselves were honestly not great at medical capabilities to begin with, especially relative to current models. I'm sorry, but this study is absolutely useless. And it is unfortunate people keep bringing these studies up as evidence LLMs are bad. AAAAHHHHH
    @mcubanRT @YounisJoseph: This is a well-designed randomized control study that asks the question: Does an LLM's ability to diagnose and manage the…
    @menhguin@iScienceLuvr they putting models i dont even know how to access anymore

    4 Sources

    @iScienceLuvrPLEASE IM BEGGING YOU TO STOP USING GPT-4O FOR EVALUATIONS On top of that, using Llama-3 and Command-R?! 🤮 LLMs aren't without limitations, but I can't trust any conclusions made by studies using such old models. I mean these models themselves were honestly not great at medical capabilities to begin with, especially relative to current models. I'm sorry, but this study is absolutely useless. And it is unfortunate people keep bringing these studies up as evidence LLMs are bad. AAAAHHHHH
    @mcubanRT @YounisJoseph: This is a well-designed randomized control study that asks the question: Does an LLM's ability to diagnose and manage the…
    @menhguin@iScienceLuvr they putting models i dont even know how to access anymore