CRYPTONEWSFREE ← Back to Live Stream
Crypto Briefing • October 8th 2026, 6:38 PM

Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

Key Summary

Harvey's Legal Agent Benchmark, LAB-AA, now includes hallucination checks to evaluate model performance in a legal context. The new feature assesses whether a model's claims are supported by underlying source documents, enhancing the benchmark's reliability. This update aims to reduce errors in legal workflows and improve the accuracy of AI models.

Please see our real time news feed on our Home Page

What Changed in LAB-AA v1.1

Harvey announced the v1.1 release of LAB-AA, which adds a layer of hallucination checking to how models are scored. Outputs are now judged on whether they have explicit support in the underlying source documents.

Background

LAB-AA is the version of the benchmark run by Artificial Analysis, the independent AI evaluation firm. It launched around July 7, 2026, using a private set of 120 tasks and a single-judge scoring system. Keeping the task set private serves a practical purpose. Models cannot be quietly tuned to ace questions they have never seen, which keeps the leaderboard honest.

The New Hallucination Metric

The new hallucination metric sits on top of the existing scoring approach. Previously, the focus was on whether a model's answer satisfied expert-written rubric criteria. Now the benchmark also asks whether each claim is grounded in what the documents actually say.

History of LAB

Harvey open-sourced its Legal Agent Benchmark, known as LAB, on May 6, 2026. The release included more than 1,200 tasks spread across 24-plus practice areas. Those tasks are graded against more than 75,000 rubric criteria written by legal experts. Artificial Analysis followed roughly two months later with LAB-AA, carving out its private 120-task subset and applying its own evaluation setup.

Impact of the Update

Adding hallucination checks raises the bar further. A model could, in principle, hit rubric criteria while padding its answer with claims the documents do not support. The v1.1 update is designed to catch exactly that behavior. By Harvey's figures, foundation models land between 0.7% and 1.9% in terms of hallucination rates. This could translate into a meaningful difference in the number of errors a reviewing attorney has to catch in a legal workflow that touches thousands of documents.
#AI#LegalTech#US#Crypto

Latest Related Headlines