Legal AI benchmark adds hallucination gate; Grok 4.7 leads at 9.4%
Artificial Analysis says the updated benchmark only credits tasks meeting every rubric criterion without a material hallucination.
TLDR
Artificial Analysis says Harvey LAB-AA v1.1, developed with Harvey, adds a hallucination check to its legal-agent benchmark. A task counts toward the headline score only if it meets every rubric criterion without a material hallucination. Grok 4.7 (xhigh) leads at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%. More than 60% of otherwise passing results across models tested at launch contained a material hallucination, it says.
Combined views
1.1M
9 Sources, first seen ago
