Researcher Urges Stop to GPT-4o Medical Evaluations
Tanishq Abraham questions reliability of studies using outdated LLMs for medical tasks.
TLDR
Tanishq Mathew Abraham posted a plea to halt use of GPT-4o, Llama-3, and Command-R in medical LLM evaluations. He stated the models were never strong at medical capabilities and that conclusions from studies relying on them cannot be trusted. Mark Cuban retweeted a post describing a randomized control study that asked whether an LLM's diagnostic and management performance holds when patient input is added. A generated summary in the packet reported the study found a 60 percent drop in diagnostic accuracy and a 12 percent decline in appropriate management decisions.
Combined views
57.7K
4 Sources, first seen 39d ago
