Grok 4.7 and Claude Opus 5.5 score zero on CancerBench's cancer-cure metric
CancerBench's creator says the two models have been added and now tie for first and last place. The benchmark's only metric counts how many types of cancer a model has cured.
TLDR
CancerBench's creator says they made the benchmark in response to AI lab CEOs talking about curing cancer. After adding Grok 4.7 and Claude Opus 5.5, the creator says both score zero on its sole metric: the number of cancer types a model has cured.
Combined views
6.9K
1 Source, first seen 5h ago
likes
