Harvey LAB-AA v1.1 加入模型幻覺檢查,提升智能體法律工作的評測標準
Announcing Harvey LAB-AA v1.1: adding hallucination checks to raise the bar for agentic legal work
Harvey 的 LAB-AA v1.1 新增按任務來源文件核查模型幻覺,重大幻覺會令任務在 Hallucination-Gated All-Pass Rate 中得零分。
Score
Changes in Harvey LAB-AA v1.1
Hallucinations
In human preference studies run by Harvey, human experts flagged hallucinations as a primary determining factor in which response they prefer out of two otherwise comprehensive answers. Our hallucination check focuses on errors that could materially affect real legal work, and takes a conservative approach to flagging them. It distinguishes unsupported task-specific claims from general legal knowledge and places the latter out of scope, so models are not penalized for drawing on case law or statutes beyond the source files.
Every deliverable a model submits is audited for hallucinations against the task's source documents. A task with no usable submission scores zero and is not audited nor used in the hallucinations per task calculation. A hallucination is a claim in the work product that the sources do not support, and each falls into one of three categories:
- A contradiction of the source materials.
- Fabricated source content.
- A specific assertion with no support in the record.
Each hallucination is graded material or minor. A material hallucination would mislead a reader on a substantive point, such as a wrong contractually required date. A minor hallucination is a real error that is unlikely to meaningfully affect the legal interpretation of a deliverable. Only material hallucinations affect scoring: having one or more zeroes a task's score in the Hallucination-Gated All-Pass Rate. Criterion Pass Rate measures rubric coverage before the hallucination gate. Material and minor hallucination counts are reported per task the check ran on; minor hallucinations do not affect scores.
The GPT-6 models hallucinate least: GPT-6 Astra averages 0.03 material hallucinations per task (4 hallucinations across all 120 tasks) and GPT-6 Sol 0.07 (8 hallucinations across all 120 tasks), while Gemini 3.8 Flash averages the most in the launch set, at 13.96 per task. Completing the criteria and not hallucinating are different skills: Muse Spark 1.3 passes the most criteria (96.0%) but averages 1.68 material hallucinations per task, against 0.03 for GPT-6 Astra. Every open-weights model averages at least 2.09 material hallucinations per task. A model's hallucination rate depends far more on the model than on the practice area.
How we chose the hallucination checker
We use GPT-6 Sol (high) for both passes of the hallucination check, separately from the three-judge rubric panel. We compared six candidates for the hallucination judge: GPT-6 Sol, GPT-6 Luna, Grok 4.7, Claude Opus 5.5, Claude Sonnet 5.5 and Gemini 3.8 Flash, each at high reasoning effort, on the same 20 tasks and deliverables from eight evaluated models.
Across this subset of tasks, GPT-6 Sol and GPT-6 Luna generally identified more material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash identified far fewer. Claude Opus 5.5 fell between Grok 4.7 and Claude Sonnet 5.5 in every model row. In total, Opus upheld 99 material hallucinations, compared with 219 for Grok, 57 for Sonnet and 470 for GPT-6 Sol. GPT-6 Sol identified more material hallucinations despite the check's conservative approach to flagging errors, which excludes general legal knowledge not contained in the source documents.
All six checkers found no material hallucinations in GPT-6 Astra's outputs, while GPT-6 Sol ranged from 0 to 0.20 per task. Both were among the models with the fewest material hallucinations under every checker.
Harvey LAB-AA v1.1: Hallucination Judge Model Comparison
Mean upheld material hallucinations per task · Subset of 20 tasks and 8 models · Lower is better
Mean material hallucinations per task
0481217
| MODEL BEING CHECKED | HALLUCINATION JUDGE | |||||
|---|---|---|---|---|---|---|
| GPT-6 Sol(high) | GPT-6 Luna(high) | Grok 4.7(high) | Claude Opus 5.5(high) | Claude Sonnet 5.5(high) | Gemini 3.8 Flash(high) | |
| GPT-6 Astra (max) | ||||||
| GPT-6 Sol (max) | ||||||
| Grok 4.7 (xhigh) | ||||||
| Claude Opus 5.5 (max with fallback) | ||||||
| Claude Fable 5.1 (max with fallback) | ||||||
| Muse Spark 1.3 (max) | ||||||
| Kimi K3 (max) | ||||||
| Gemini 3.8 Flash (high) | ||||||
| Total material hallucinations per task | 23.50 | 19.35 | 10.95 | 4.95 | 2.85 | 1.40 |
Hallucination-Gated Near-Pass Rate
Cost
Token Usage
Speed
Turns
Model Size (Open Weights Models Only)
Score vs. Release Date
Browse representative Harvey LAB tasks from the public task set, the reference files each model was given, and the deliverables it produced.
Harvey LAB-AA resources
來源:Artificial Analysis 完整文章 · artificialanalysis.ai