Researcher Urges Stop to GPT-4o Medical Evaluations
Tanishq Abraham questions reliability of studies using outdated LLMs for medical tasks.
Tanishq Mathew Abraham posted a plea to halt use of GPT-4o, Llama-3, and Command-R in medical LLM evaluations. He stated the models were never strong at medical capabilities and that conclusions from studies relying on them cannot be trusted. Mark Cuban retweeted a post describing a randomized control study that asked whether an LLM's diagnostic and management performance holds when patient input is added. A generated summary in the packet reported the study found a 60 percent drop in diagnostic accuracy and a 12 percent decline in appropriate management decisions.
PLEASE IM BEGGING YOU TO STOP USING GPT-4O FOR EVALUATIONS On top of that, using Llama-3 and Command-R?! 🤮 LLMs aren't without limitations, but I can't trust any conclusions made by studies using such old models. I mean these models themselves were honestly not great at…
This is a well-designed randomized control study that asks the question: Does an LLM's ability to diagnose and manage the same medical scenarios degrade when a patient is the communicator versus a physician. The result: 60% drop in diagnostic accuracy, 12% drop in appropriate…


