HUMAN EDGE · LEGAL AGENT BENCHMARK

AGENT EVALUATION · INTERNATIONAL ARBITRATION

This report presents the overall evaluation of agentic legal performance on international-arbitration work. The benchmark uses structured tasks, controlled matter rooms, reference answers and expert rubrics to assess both answer quality and agentic execution, and will expand as new tasks and models are evaluated.

Scope. We built the evaluation harness, task preparation, scoring implementation and LLM judge independently. Results are specific to the evaluated task packages, candidate models and judge configuration.
EXECUTIVE SUMMARY

What the overall evaluation indicates

Quality

Claude Opus 5 had the highest overall equal-weight mean at 79.8%. The next-highest model was GPT-6 Astra at 73.3%.

Efficiency trade-off

Opus achieved the strongest quality but used 106.7 mean model turns and 139.5 mean tool calls. Among models averaging no more than 20 turns per task, GPT-6 Astra provided the strongest quality-efficiency balance: 73.3% with 13.3 turns and 66.3 tool calls.

Task sensitivity

Scores depended strongly on the task. Across the fully covered models, 001L and 002L were comparatively strong, while 002M produced the lowest cross-model performance. Model ranking should therefore be read task by task, not as one universal capability score.

Overall model comparison

This table puts overall quality beside the amount of agent activity required to produce it. The score is the equal-weight mean of the task-level weighted scores currently included in the evaluation; turns and tool calls are averaged across those same tasks.

ModelOverall meanMean turnsMean tool calls
Claude Opus 579.8%106.7139.5
GPT-6 Astra73.3%13.366.3
Gemini 3.6 Flash71.3%27.626.6
GPT-5.6 Sol71.0%6.736.9
Claude Sonnet 569.1%93.8128.5
Grok 4.669.0%10.327.7
GPT-5.6 Luna58.0%6.625.6
Limitations
  • This is an independently implemented agentic legal evaluation. Results are specific to this harness, task package and judge configuration.
  • The default judge was Claude Sonnet 5. Its assessment of Sonnet candidate outputs creates a potential same-family evaluation effect.
  • Three requested runs per task provide an initial consistency signal, not a statistically robust estimate. Reported scores use completed, fully judged runs only.
  • Candidate models used the same task materials, agent tools and output contract, but provider-specific reasoning and runtime settings were not identical.
  • Rubric categories summarize criterion outcomes; they do not replace expert review of the generated legal work products.
METHODOLOGY

How the agent harness and evaluation worked

The tasks were derived from real-world international arbitration matters and organized into controlled matter rooms. Client-facing case descriptions and downloadable work products in this report are anonymized.

1. Task construction

Under this agentic legal evaluation structure, each task combined a client-style instruction, a controlled matter room, a reference answer and an expert rubric. The current release covers deliverables across Cases 001–003; later versions can extend the task set without changing the report structure.

2. Small, medium and long tasks

The labels describe the scope of the work rather than the quality expected. Small tasks are focused assignments with a compact deliverable; medium tasks require broader analysis or evidence review across more materials; long tasks require the most extensive source synthesis and a fuller legal work product. The labels are relative within this evaluation and do not impose one universal page or token limit.

3. Controlled agent harness

For each run, the harness created an isolated workspace and gave the candidate model the task, matter files and a defined DOCX output contract. Six active tools supported source inventory, bounded source reading, matter-room search, spreadsheet inspection, calculation and workspace notes. A seventh registered tool for external legal research was disabled for these evaluations. The reference answer and rubric were withheld from the candidate.

4. Independent executions

Three separate runs were requested for every candidate model and task. Each run started from the same instructions and matter materials without access to another run's work. The harness retained the submitted DOCX and recorded operational measures including elapsed time, model turns, tool calls, sources read and citation support.

5. LLM judge evaluation

After a candidate submitted its final document, the harness passed that work product—together with the task, expert rubric and reference evidence—to the default judge, Claude Sonnet 5. The judge evaluated every rubric criterion individually, recording whether it was satisfied and grounding the decision in the candidate response and reference materials. The judge assessed the completed output only; it did not participate in or alter the candidate run.

6. Validation and reporting

Deterministic checks confirmed that the required artifact existed and opened successfully. A run entered the results only after both the candidate artifact and criterion-level judgment were complete. The headline percentage is the rubric weight earned divided by the total available weight: high-importance criteria are worth 3 points, medium-importance criteria 2 points and low-importance criteria 1 point. Task results average that weighted score across completed runs; category results group the underlying criterion decisions by rubric theme.

Contract dispute strategy

A multilingual telecommunications dispute requiring a respondent pleading and a client strategy memorandum. Matter languages: English and French.

TASK 001M

Respondent pleading

Draft a structured respondent pleading addressing the contractual history, termination, waiver, quantum, counterclaims, procedure and relief sought.

1
Claude Sonnet 5
83.5%
mean score · 3 completed runs
range 79.3%–87.4%
2
Claude Opus 5
83.1%
mean score · 3 completed runs
range 80.5%–86.2%
3
Gemini 3.6 Flash
75.5%
mean score · 3 completed runs
range 72.4%–78.2%
4
GPT-6 Astra
75.1%
mean score · 3 completed runs
range 73.6%–75.9%
5
Grok 4.6
73.2%
mean score · 3 completed runs
range 70.1%–74.7%
6
GPT-5.6 Sol
69.7%
mean score · 3 completed runs
range 69.0%–71.3%
7
GPT-5.6 Luna
66.3%
mean score · 3 completed runs
range 64.4%–69.0%
Interpretation. Claude Sonnet 5 recorded the highest mean score at 83.5%. The difference between the highest and lowest completed model means was 17.2 points. GPT-5.6 Luna had the lowest mean candidate latency (84.8 seconds).

Agent performance

These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.

ModelScoreLatencyTurnsTool callsSources readCitation support
Claude Sonnet 583.5%323.6 s45.080.748.3100.0%
Claude Opus 583.1%968.6 s120.7156.064.7100.0%
Gemini 3.6 Flash75.5%139.8 s27.326.315.7100.0%
GPT-6 Astra75.1%481.9 s13.063.040.3100.0%
Grok 4.673.2%175.5 s10.332.317.7100.0%
GPT-5.6 Sol69.7%167.1 s8.743.027.398.1%
GPT-5.6 Luna66.3%84.8 s6.023.718.3100.0%

Rubric-category performance

How rubric weighting works. Each high-importance criterion is worth 3 points, each medium-importance criterion 2 points and each low-importance criterion 1 point. A passed criterion earns its full weight and an unmet criterion earns zero. The headline score is the total weight earned divided by the total weight available, averaged across completed runs.

The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.

Claude Sonnet 5
Relief & Remedies
50.0%
Template & Procedural Compliance
66.7%
Document Selection & Sourcing
83.3%
Citations & Evidentiary Grounding
86.7%
Factual Chronology of the Dispute
90.5%
Legal Analysis - Core Issues
93.3%
Governing Legal Framework
100.0%
Parties & Procedural History
100.0%
Claude Opus 5
Template & Procedural Compliance
63.0%
Relief & Remedies
66.7%
Citations & Evidentiary Grounding
80.0%
Document Selection & Sourcing
83.3%
Factual Chronology of the Dispute
90.5%
Legal Analysis - Core Issues
93.3%
Governing Legal Framework
100.0%
Parties & Procedural History
100.0%
Gemini 3.6 Flash
Relief & Remedies
50.0%
Template & Procedural Compliance
59.3%
Factual Chronology of the Dispute
61.9%
Citations & Evidentiary Grounding
80.0%
Document Selection & Sourcing
83.3%
Governing Legal Framework
100.0%
Legal Analysis - Core Issues
100.0%
Parties & Procedural History
100.0%
GPT-6 Astra
Document Selection & Sourcing
50.0%
Template & Procedural Compliance
55.6%
Citations & Evidentiary Grounding
66.7%
Factual Chronology of the Dispute
71.4%
Relief & Remedies
83.3%
Legal Analysis - Core Issues
93.3%
Governing Legal Framework
100.0%
Parties & Procedural History
100.0%
Grok 4.6
Document Selection & Sourcing
50.0%
Relief & Remedies
50.0%
Factual Chronology of the Dispute
61.9%
Citations & Evidentiary Grounding
66.7%
Template & Procedural Compliance
74.1%
Governing Legal Framework
100.0%
Legal Analysis - Core Issues
100.0%
Parties & Procedural History
100.0%
GPT-5.6 Sol
Relief & Remedies
33.3%
Document Selection & Sourcing
50.0%
Template & Procedural Compliance
59.3%
Citations & Evidentiary Grounding
60.0%
Factual Chronology of the Dispute
61.9%
Governing Legal Framework
100.0%
Legal Analysis - Core Issues
100.0%
Parties & Procedural History
100.0%
GPT-5.6 Luna
Document Selection & Sourcing
50.0%
Relief & Remedies
50.0%
Factual Chronology of the Dispute
52.4%
Citations & Evidentiary Grounding
53.3%
Template & Procedural Compliance
59.3%
Legal Analysis - Core Issues
93.3%
Governing Legal Framework
100.0%
Parties & Procedural History
100.0%
Generated work products

Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.

TASK 001L

Client memorandum

Prepare a confidential and privileged client memorandum on the contractual dispute, legal and financial risks, and recommended strategy, using verified facts and contractual wording.

1
Claude Opus 5
88.5%
mean score · 3 completed runs
range 85.7%–91.7%
2
GPT-5.6 Sol
81.7%
mean score · 3 completed runs
range 78.6%–83.3%
3
GPT-6 Astra
81.0%
mean score · 3 completed runs
range 79.8%–82.1%
4
Gemini 3.6 Flash
79.8%
mean score · 3 completed runs
range 73.8%–85.7%
5
Claude Sonnet 5
77.0%
mean score · 3 completed runs
range 66.7%–85.7%
6
GPT-5.6 Luna
73.4%
mean score · 3 completed runs
range 72.6%–73.8%
7
Grok 4.6
67.9%
mean score · 3 completed runs
range 67.9%–67.9%
Interpretation. Claude Opus 5 recorded the highest mean score at 88.5%. The difference between the highest and lowest completed model means was 20.6 points. GPT-5.6 Luna had the lowest mean candidate latency (88.8 seconds).

Agent performance

These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.

ModelScoreLatencyTurnsTool callsSources readCitation support
Claude Opus 588.5%683.9 s96.7129.358.3100.0%
GPT-5.6 Sol81.7%125.3 s7.330.319.0100.0%
GPT-6 Astra81.0%450.6 s11.750.035.0100.0%
Gemini 3.6 Flash79.8%147.1 s36.735.714.0100.0%
Claude Sonnet 577.0%264.0 s41.767.341.7100.0%
GPT-5.6 Luna73.4%88.8 s12.744.012.398.5%
Grok 4.667.9%158.6 s17.343.017.096.8%

Rubric-category performance

How rubric weighting works. Each high-importance criterion is worth 3 points, each medium-importance criterion 2 points and each low-importance criterion 1 point. A passed criterion earns its full weight and an unmet criterion earns zero. The headline score is the total weight earned divided by the total weight available, averaged across completed runs.

The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.

Claude Opus 5
Governing Legal Framework
50.0%
Factual Chronology of the Dispute
66.7%
Template & Procedural Compliance
93.3%
Document Selection & Sourcing
100.0%
Legal Analysis - Core Issues
100.0%
Parties & Procedural History
100.0%
Relief & Remedies
100.0%
GPT-5.6 Sol
Governing Legal Framework
50.0%
Factual Chronology of the Dispute
54.2%
Parties & Procedural History
66.7%
Template & Procedural Compliance
80.0%
Legal Analysis - Core Issues
96.3%
Document Selection & Sourcing
100.0%
Relief & Remedies
100.0%
GPT-6 Astra
Parties & Procedural History
33.3%
Governing Legal Framework
50.0%
Factual Chronology of the Dispute
62.5%
Template & Procedural Compliance
80.0%
Document Selection & Sourcing
83.3%
Legal Analysis - Core Issues
100.0%
Relief & Remedies
100.0%
Gemini 3.6 Flash
Governing Legal Framework
16.7%
Factual Chronology of the Dispute
54.2%
Template & Procedural Compliance
80.0%
Legal Analysis - Core Issues
92.6%
Document Selection & Sourcing
100.0%
Parties & Procedural History
100.0%
Relief & Remedies
100.0%
Claude Sonnet 5
Governing Legal Framework
33.3%
Factual Chronology of the Dispute
54.2%
Template & Procedural Compliance
80.0%
Document Selection & Sourcing
83.3%
Legal Analysis - Core Issues
92.6%
Parties & Procedural History
100.0%
Relief & Remedies
100.0%
GPT-5.6 Luna
Governing Legal Framework
16.7%
Parties & Procedural History
33.3%
Factual Chronology of the Dispute
50.0%
Document Selection & Sourcing
72.2%
Template & Procedural Compliance
80.0%
Legal Analysis - Core Issues
100.0%
Relief & Remedies
100.0%
Grok 4.6
Parties & Procedural History
0.0%
Governing Legal Framework
16.7%
Factual Chronology of the Dispute
33.3%
Template & Procedural Compliance
80.0%
Legal Analysis - Core Issues
85.2%
Document Selection & Sourcing
88.9%
Relief & Remedies
100.0%
Generated work products

Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.

CAIP arbitration defence

A cross-border sale-of-goods arbitration involving force majeure, exhibit assessment and a substantive reply. Matter languages: English, French, Spanish and Chinese.

TASK 002S

Force majeure memorandum

Prepare a five-to-six-page English memorandum on contractual force majeure under French law for a respondent in a CAIP arbitration.

1
Claude Opus 5
80.6%
mean score · 3 completed runs
range 75.9%–83.5%
2
GPT-6 Astra
74.3%
mean score · 3 completed runs
range 72.2%–78.5%
3
GPT-5.6 Sol
73.8%
mean score · 3 completed runs
range 68.4%–79.7%
4
Claude Sonnet 5
71.3%
mean score · 3 completed runs
range 65.8%–81.0%
5
Grok 4.6
68.8%
mean score · 3 completed runs
range 63.3%–78.5%
6
Gemini 3.6 Flash
64.6%
mean score · 3 completed runs
range 59.5%–69.6%
7
GPT-5.6 Luna
46.8%
mean score · 3 completed runs
range 39.2%–54.4%
Interpretation. Claude Opus 5 recorded the highest mean score at 80.6%. The difference between the highest and lowest completed model means was 33.8 points. GPT-5.6 Luna had the lowest mean candidate latency (75.6 seconds).

Agent performance

These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.

ModelScoreLatencyTurnsTool callsSources readCitation support
Claude Opus 580.6%1,159.0 s140.0180.385.0100.0%
GPT-6 Astra74.3%366.8 s13.770.043.093.5%
GPT-5.6 Sol73.8%166.2 s7.046.718.7100.0%
Claude Sonnet 571.3%628.7 s118.7151.740.098.5%
Grok 4.668.8%122.6 s7.723.06.3100.0%
Gemini 3.6 Flash64.6%120.1 s24.323.38.0100.0%
GPT-5.6 Luna46.8%75.6 s5.318.011.792.2%

Rubric-category performance

How rubric weighting works. Each high-importance criterion is worth 3 points, each medium-importance criterion 2 points and each low-importance criterion 1 point. A passed criterion earns its full weight and an unmet criterion earns zero. The headline score is the total weight earned divided by the total weight available, averaged across completed runs.

The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.

Claude Opus 5
Document Selection & Sourcing
0.0%
Legal Authority & Precedent
50.0%
Template & Procedural Compliance
53.3%
Governing Legal Framework
81.0%
Negative Space
88.9%
Legal Analysis - Core Issues
95.8%
Citations & Evidentiary Grounding
100.0%
Factual Chronology of the Dispute
100.0%
Hedging Guards
100.0%
Relief & Remedies
100.0%
GPT-6 Astra
Document Selection & Sourcing
0.0%
Template & Procedural Compliance
40.0%
Legal Authority & Precedent
66.7%
Negative Space
66.7%
Governing Legal Framework
71.4%
Legal Analysis - Core Issues
79.2%
Citations & Evidentiary Grounding
100.0%
Factual Chronology of the Dispute
100.0%
Hedging Guards
100.0%
Relief & Remedies
100.0%
GPT-5.6 Sol
Document Selection & Sourcing
0.0%
Factual Chronology of the Dispute
33.3%
Negative Space
44.4%
Template & Procedural Compliance
46.7%
Legal Authority & Precedent
50.0%
Governing Legal Framework
85.7%
Legal Analysis - Core Issues
87.5%
Citations & Evidentiary Grounding
100.0%
Hedging Guards
100.0%
Relief & Remedies
100.0%
Claude Sonnet 5
Document Selection & Sourcing
0.0%
Citations & Evidentiary Grounding
50.0%
Legal Authority & Precedent
50.0%
Negative Space
55.6%
Governing Legal Framework
66.7%
Template & Procedural Compliance
66.7%
Relief & Remedies
83.3%
Legal Analysis - Core Issues
87.5%
Factual Chronology of the Dispute
100.0%
Hedging Guards
100.0%
Grok 4.6
Document Selection & Sourcing
0.0%
Negative Space
33.3%
Legal Authority & Precedent
50.0%
Factual Chronology of the Dispute
66.7%
Legal Analysis - Core Issues
66.7%
Template & Procedural Compliance
73.3%
Governing Legal Framework
76.2%
Citations & Evidentiary Grounding
100.0%
Hedging Guards
100.0%
Relief & Remedies
100.0%
Gemini 3.6 Flash
Document Selection & Sourcing
0.0%
Hedging Guards
33.3%
Negative Space
33.3%
Legal Authority & Precedent
50.0%
Relief & Remedies
50.0%
Citations & Evidentiary Grounding
66.7%
Governing Legal Framework
66.7%
Template & Procedural Compliance
66.7%
Legal Analysis - Core Issues
87.5%
Factual Chronology of the Dispute
100.0%
GPT-5.6 Luna
Document Selection & Sourcing
0.0%
Factual Chronology of the Dispute
0.0%
Legal Authority & Precedent
16.7%
Citations & Evidentiary Grounding
33.3%
Negative Space
33.3%
Template & Procedural Compliance
46.7%
Relief & Remedies
50.0%
Governing Legal Framework
52.4%
Legal Analysis - Core Issues
58.3%
Hedging Guards
100.0%
Generated work products

Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.

TASK 002M

Exhibit review

Assess the evidential weight of each exhibit MA001–MA017 submitted with the claimant’s case and present the assessment as a usable table.

1
GPT-6 Astra
55.6%
mean score · 3 completed runs
range 54.0%–57.1%
2
Claude Opus 5
54.5%
mean score · 3 completed runs
range 46.0%–63.5%
3
Gemini 3.6 Flash
50.8%
mean score · 3 completed runs
range 47.6%–55.6%
4
GPT-5.6 Sol
40.2%
mean score · 3 completed runs
range 31.7%–46.0%
5
Grok 4.6
33.9%
mean score · 3 completed runs
range 31.7%–36.5%
6
Claude Sonnet 5
30.7%
mean score · 3 completed runs
range 30.2%–31.7%
7
GPT-5.6 Luna
18.0%
mean score · 3 completed runs
range 9.5%–25.4%
Interpretation. GPT-6 Astra recorded the highest mean score at 55.6%. The difference between the highest and lowest completed model means was 37.6 points. GPT-5.6 Luna had the lowest mean candidate latency (75.4 seconds).

Agent performance

These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.

ModelScoreLatencyTurnsTool callsSources readCitation support
GPT-6 Astra55.6%394.7 s14.068.337.0100.0%
Claude Opus 554.5%753.5 s108.7140.061.7100.0%
Gemini 3.6 Flash50.8%137.7 s14.713.710.3100.0%
GPT-5.6 Sol40.2%148.9 s5.319.711.0100.0%
Grok 4.633.9%92.2 s6.315.37.0100.0%
Claude Sonnet 530.7%519.7 s130.0182.041.0100.0%
GPT-5.6 Luna18.0%75.4 s3.39.08.0100.0%

Rubric-category performance

How rubric weighting works. Each high-importance criterion is worth 3 points, each medium-importance criterion 2 points and each low-importance criterion 1 point. A passed criterion earns its full weight and an unmet criterion earns zero. The headline score is the total weight earned divided by the total weight available, averaged across completed runs.

The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.

GPT-6 Astra
Hedging Guards
0.0%
Legal Analysis - Core Issues
16.7%
Factual Chronology of the Dispute
33.3%
Template & Procedural Compliance
33.3%
Legal Authority & Precedent
50.0%
Negative Space
50.0%
Governing Legal Framework
91.7%
Citations & Evidentiary Grounding
100.0%
Document Selection & Sourcing
100.0%
Parties & Procedural History
100.0%
Relief & Remedies
100.0%
Claude Opus 5
Hedging Guards
0.0%
Negative Space
16.7%
Legal Analysis - Core Issues
22.2%
Template & Procedural Compliance
22.2%
Parties & Procedural History
33.3%
Legal Authority & Precedent
66.7%
Document Selection & Sourcing
83.3%
Governing Legal Framework
83.3%
Factual Chronology of the Dispute
88.9%
Citations & Evidentiary Grounding
100.0%
Relief & Remedies
100.0%
Gemini 3.6 Flash
Hedging Guards
0.0%
Negative Space
0.0%
Governing Legal Framework
25.0%
Legal Authority & Precedent
33.3%
Legal Analysis - Core Issues
38.9%
Template & Procedural Compliance
66.7%
Factual Chronology of the Dispute
88.9%
Citations & Evidentiary Grounding
100.0%
Document Selection & Sourcing
100.0%
Parties & Procedural History
100.0%
Relief & Remedies
100.0%
GPT-5.6 Sol
Hedging Guards
0.0%
Parties & Procedural History
0.0%
Legal Analysis - Core Issues
16.7%
Factual Chronology of the Dispute
22.2%
Template & Procedural Compliance
22.2%
Legal Authority & Precedent
33.3%
Negative Space
50.0%
Governing Legal Framework
58.3%
Document Selection & Sourcing
66.7%
Citations & Evidentiary Grounding
100.0%
Relief & Remedies
100.0%
Grok 4.6
Hedging Guards
0.0%
Legal Authority & Precedent
0.0%
Governing Legal Framework
16.7%
Legal Analysis - Core Issues
16.7%
Negative Space
33.3%
Template & Procedural Compliance
33.3%
Factual Chronology of the Dispute
44.4%
Document Selection & Sourcing
66.7%
Citations & Evidentiary Grounding
100.0%
Parties & Procedural History
100.0%
Relief & Remedies
100.0%
Claude Sonnet 5
Hedging Guards
0.0%
Legal Authority & Precedent
0.0%
Negative Space
0.0%
Template & Procedural Compliance
0.0%
Governing Legal Framework
16.7%
Legal Analysis - Core Issues
16.7%
Factual Chronology of the Dispute
55.6%
Parties & Procedural History
66.7%
Document Selection & Sourcing
83.3%
Citations & Evidentiary Grounding
100.0%
Relief & Remedies
100.0%
GPT-5.6 Luna
Factual Chronology of the Dispute
0.0%
Governing Legal Framework
0.0%
Hedging Guards
0.0%
Legal Authority & Precedent
0.0%
Negative Space
0.0%
Legal Analysis - Core Issues
11.1%
Template & Procedural Compliance
11.1%
Citations & Evidentiary Grounding
33.3%
Parties & Procedural History
66.7%
Document Selection & Sourcing
83.3%
Relief & Remedies
100.0%
Generated work products

Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.

TASK 002L

Arbitration reply

Draft an approximately 20-page English reply using the supplied legal authorities RL-1–RL-23 and client factual materials R-1–R-24.

1
Claude Opus 5
83.3%
mean score · 3 completed runs
range 83.3%–83.3%
2
Grok 4.6
79.0%
mean score · 3 completed runs
range 77.8%–79.6%
3
Gemini 3.6 Flash
76.5%
mean score · 3 completed runs
range 64.8%–85.2%
4
GPT-5.6 Sol
75.3%
mean score · 3 completed runs
range 68.5%–83.3%
5
GPT-6 Astra
69.8%
mean score · 3 completed runs
range 68.5%–72.2%
6
Claude Sonnet 5
67.9%
mean score · 3 completed runs
range 57.4%–79.6%
7
GPT-5.6 Luna
63.6%
mean score · 3 completed runs
range 57.4%–70.4%
Interpretation. Claude Opus 5 recorded the highest mean score at 83.3%. The difference between the highest and lowest completed model means was 19.8 points. GPT-5.6 Luna had the lowest mean candidate latency (145.3 seconds).

Agent performance

These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.

ModelScoreLatencyTurnsTool callsSources readCitation support
Claude Opus 583.3%1,291.4 s129.7182.772.7100.0%
Grok 4.679.0%192.7 s8.735.028.3100.0%
Gemini 3.6 Flash76.5%252.3 s50.349.345.0100.0%
GPT-5.6 Sol75.3%299.1 s7.072.060.0100.0%
GPT-6 Astra69.8%734.9 s19.7130.764.0100.0%
Claude Sonnet 567.9%422.4 s50.7107.359.7100.0%
GPT-5.6 Luna63.6%145.3 s6.046.743.097.8%

Rubric-category performance

How rubric weighting works. Each high-importance criterion is worth 3 points, each medium-importance criterion 2 points and each low-importance criterion 1 point. A passed criterion earns its full weight and an unmet criterion earns zero. The headline score is the total weight earned divided by the total weight available, averaged across completed runs.

The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.

Claude Opus 5
Governing Legal Framework
50.0%
Negative Space
50.0%
Relief & Remedies
66.7%
Template & Procedural Compliance
80.0%
Citations & Evidentiary Grounding
100.0%
Document Selection & Sourcing
100.0%
Factual Chronology of the Dispute
100.0%
Hedging Guards
100.0%
Legal Analysis - Core Issues
100.0%
Parties & Procedural History
100.0%
Grok 4.6
Governing Legal Framework
50.0%
Negative Space
50.0%
Factual Chronology of the Dispute
66.7%
Relief & Remedies
66.7%
Template & Procedural Compliance
80.0%
Legal Analysis - Core Issues
93.3%
Citations & Evidentiary Grounding
100.0%
Document Selection & Sourcing
100.0%
Hedging Guards
100.0%
Parties & Procedural History
100.0%
Gemini 3.6 Flash
Citations & Evidentiary Grounding
33.3%
Governing Legal Framework
50.0%
Relief & Remedies
55.6%
Hedging Guards
66.7%
Negative Space
66.7%
Legal Analysis - Core Issues
86.7%
Template & Procedural Compliance
93.3%
Document Selection & Sourcing
100.0%
Factual Chronology of the Dispute
100.0%
Parties & Procedural History
100.0%
GPT-5.6 Sol
Negative Space
33.3%
Parties & Procedural History
33.3%
Governing Legal Framework
50.0%
Relief & Remedies
66.7%
Template & Procedural Compliance
80.0%
Legal Analysis - Core Issues
86.7%
Citations & Evidentiary Grounding
100.0%
Document Selection & Sourcing
100.0%
Factual Chronology of the Dispute
100.0%
Hedging Guards
100.0%
GPT-6 Astra
Citations & Evidentiary Grounding
0.0%
Parties & Procedural History
33.3%
Factual Chronology of the Dispute
50.0%
Governing Legal Framework
50.0%
Negative Space
50.0%
Template & Procedural Compliance
60.0%
Relief & Remedies
66.7%
Document Selection & Sourcing
100.0%
Hedging Guards
100.0%
Legal Analysis - Core Issues
100.0%
Claude Sonnet 5
Parties & Procedural History
0.0%
Governing Legal Framework
33.3%
Negative Space
33.3%
Citations & Evidentiary Grounding
66.7%
Relief & Remedies
66.7%
Template & Procedural Compliance
73.3%
Legal Analysis - Core Issues
80.0%
Document Selection & Sourcing
100.0%
Factual Chronology of the Dispute
100.0%
Hedging Guards
100.0%
GPT-5.6 Luna
Citations & Evidentiary Grounding
0.0%
Negative Space
16.7%
Parties & Procedural History
33.3%
Governing Legal Framework
50.0%
Template & Procedural Compliance
60.0%
Factual Chronology of the Dispute
66.7%
Relief & Remedies
66.7%
Legal Analysis - Core Issues
86.7%
Document Selection & Sourcing
100.0%
Hedging Guards
100.0%
Generated work products

Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.

Arbitration clause drafting

A clause-drafting exercise for an international telecommunications agreement, applying client instructions to governing-law and dispute-resolution provisions. Matter languages: English and French.

Input documents used for the reported runs
  • Client instructions
  • Pseudonymized roaming framework agreement
  • ICC Arbitration Rules 2026 / Mediation Rules 2014
  • SCC Arbitration Rules 2023
TASK 003S

Clause drafting

Revise governing-law and arbitration clauses in a roaming framework agreement, including French law, Paris arbitration, SCC procedure, a sole arbitrator and emergency relief.

1
Grok 4.6
91.3%
mean score · 3 completed runs
range 90.5%–92.9%
2
Claude Opus 5
88.9%
mean score · 3 completed runs
range 85.7%–90.5%
3
GPT-5.6 Sol
84.9%
mean score · 3 completed runs
range 78.6%–90.5%
4
Claude Sonnet 5
84.1%
mean score · 3 completed runs
range 83.3%–85.7%
5
GPT-6 Astra
84.1%
mean score · 3 completed runs
range 83.3%–85.7%
6
Gemini 3.6 Flash
81.0%
mean score · 3 completed runs
range 78.6%–85.7%
7
GPT-5.6 Luna
80.2%
mean score · 3 completed runs
range 73.8%–83.3%
Interpretation. Grok 4.6 recorded the highest mean score at 91.3%. The difference between the highest and lowest completed model means was 11.1 points. GPT-5.6 Luna had the lowest mean candidate latency (44.2 seconds).

Agent performance

These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.

ModelScoreLatencyTurnsTool callsSources readCitation support
Grok 4.691.3%57.6 s11.317.33.0100.0%
Claude Opus 588.9%406.2 s44.348.713.0100.0%
GPT-5.6 Sol84.9%72.5 s5.09.73.3100.0%
Claude Sonnet 584.1%499.2 s176.7182.03.0100.0%
GPT-6 Astra84.1%259.7 s7.716.05.3100.0%
Gemini 3.6 Flash81.0%49.5 s12.011.03.0100.0%
GPT-5.6 Luna80.2%44.2 s6.312.04.3100.0%

Rubric-category performance

How rubric weighting works. Each high-importance criterion is worth 3 points, each medium-importance criterion 2 points and each low-importance criterion 1 point. A passed criterion earns its full weight and an unmet criterion earns zero. The headline score is the total weight earned divided by the total weight available, averaged across completed runs.

The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.

Grok 4.6
Template & Procedural Compliance
66.7%
Legal Analysis - Core Issues
83.3%
Document Selection & Sourcing
91.7%
Governing Legal Framework
100.0%
Claude Opus 5
Document Selection & Sourcing
83.3%
Legal Analysis - Core Issues
88.9%
Governing Legal Framework
100.0%
Template & Procedural Compliance
100.0%
GPT-5.6 Sol
Template & Procedural Compliance
66.7%
Legal Analysis - Core Issues
77.8%
Document Selection & Sourcing
83.3%
Governing Legal Framework
100.0%
Claude Sonnet 5
Template & Procedural Compliance
66.7%
Document Selection & Sourcing
79.2%
Legal Analysis - Core Issues
83.3%
Governing Legal Framework
100.0%
GPT-6 Astra
Template & Procedural Compliance
66.7%
Document Selection & Sourcing
79.2%
Legal Analysis - Core Issues
83.3%
Governing Legal Framework
100.0%
Gemini 3.6 Flash
Legal Analysis - Core Issues
66.7%
Document Selection & Sourcing
79.2%
Governing Legal Framework
100.0%
Template & Procedural Compliance
100.0%
GPT-5.6 Luna
Template & Procedural Compliance
66.7%
Document Selection & Sourcing
75.0%
Legal Analysis - Core Issues
77.8%
Governing Legal Framework
100.0%
Generated work products

Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.

CHANGE LOG

What changed between versions

Each release is preserved as a dated evaluation snapshot. The unversioned report URL always opens the current release, while the archived links preserve the results and methodology shown at the time.

Version 1.1Current
  • Expanded the evaluation from the preliminary single-task release to Cases 001–003 and their current task set.
  • Added GPT-6 Astra and Gemini 3.6 Flash to the evaluated model group.
  • Corrected the preliminary low results after completing the task and source configuration, then reran the evaluation under the controlled harness.
  • Added the cross-case executive summary, model-level operational metrics, rubric-category analysis and downloadable work products.
Version 1.0Superseded
  • Initial preliminary single-task report.
  • Preserved for audit history, including the unusually low scores later determined to be incorrect.
  • Open the archived v1.0 report →