WearableQA tests AI reasoning on real users’ wearable data, a post says
A post summarizing Meta researchers’ work describes 4,084 multiple-choice questions based on 200 people’s wearable records, testing both data analysis and health interpretation.
TLDR
The post says WearableQA uses hundreds of days of wearable measurements per person, supplemented by blood biomarkers and demographic information. It reports accuracy ranging from 19.6% to 72.9% across evaluated models, against a 10% chance baseline. Data reasoning was harder than health reasoning for most models, the summary says, while reasoning across multiple signals remained challenging across model scales. In the September 7, 2026 post, the user said the benchmark had not yet been released.
Combined views
4.9K
1 Source, first seen 23d ago