Claude Opus 5 had the highest overall equal-weight mean at 79.8%. The next-highest model was GPT-6 Astra at 73.3%.
AGENT EVALUATION · INTERNATIONAL ARBITRATION
This report presents the overall evaluation of agentic legal performance on international-arbitration work. The benchmark uses structured tasks, controlled matter rooms, reference answers and expert rubrics to assess both answer quality and agentic execution, and will expand as new tasks and models are evaluated.
What the overall evaluation indicates
Opus achieved the strongest quality but used 106.7 mean model turns and 139.5 mean tool calls. Among models averaging no more than 20 turns per task, GPT-6 Astra provided the strongest quality-efficiency balance: 73.3% with 13.3 turns and 66.3 tool calls.
Scores depended strongly on the task. Across the fully covered models, 001L and 002L were comparatively strong, while 002M produced the lowest cross-model performance. Model ranking should therefore be read task by task, not as one universal capability score.
Overall model comparison
This table puts overall quality beside the amount of agent activity required to produce it. The score is the equal-weight mean of the task-level weighted scores currently included in the evaluation; turns and tool calls are averaged across those same tasks.
| Model | Overall mean | Mean turns | Mean tool calls |
|---|---|---|---|
| Claude Opus 5 | 79.8% | 106.7 | 139.5 |
| GPT-6 Astra | 73.3% | 13.3 | 66.3 |
| Gemini 3.6 Flash | 71.3% | 27.6 | 26.6 |
| GPT-5.6 Sol | 71.0% | 6.7 | 36.9 |
| Claude Sonnet 5 | 69.1% | 93.8 | 128.5 |
| Grok 4.6 | 69.0% | 10.3 | 27.7 |
| GPT-5.6 Luna | 58.0% | 6.6 | 25.6 |
- This is an independently implemented agentic legal evaluation. Results are specific to this harness, task package and judge configuration.
- The default judge was Claude Sonnet 5. Its assessment of Sonnet candidate outputs creates a potential same-family evaluation effect.
- Three requested runs per task provide an initial consistency signal, not a statistically robust estimate. Reported scores use completed, fully judged runs only.
- Candidate models used the same task materials, agent tools and output contract, but provider-specific reasoning and runtime settings were not identical.
- Rubric categories summarize criterion outcomes; they do not replace expert review of the generated legal work products.
How the agent harness and evaluation worked
The tasks were derived from real-world international arbitration matters and organized into controlled matter rooms. Client-facing case descriptions and downloadable work products in this report are anonymized.
Under this agentic legal evaluation structure, each task combined a client-style instruction, a controlled matter room, a reference answer and an expert rubric. The current release covers deliverables across Cases 001–003; later versions can extend the task set without changing the report structure.
The labels describe the scope of the work rather than the quality expected. Small tasks are focused assignments with a compact deliverable; medium tasks require broader analysis or evidence review across more materials; long tasks require the most extensive source synthesis and a fuller legal work product. The labels are relative within this evaluation and do not impose one universal page or token limit.
For each run, the harness created an isolated workspace and gave the candidate model the task, matter files and a defined DOCX output contract. Six active tools supported source inventory, bounded source reading, matter-room search, spreadsheet inspection, calculation and workspace notes. A seventh registered tool for external legal research was disabled for these evaluations. The reference answer and rubric were withheld from the candidate.
Three separate runs were requested for every candidate model and task. Each run started from the same instructions and matter materials without access to another run's work. The harness retained the submitted DOCX and recorded operational measures including elapsed time, model turns, tool calls, sources read and citation support.
After a candidate submitted its final document, the harness passed that work product—together with the task, expert rubric and reference evidence—to the default judge, Claude Sonnet 5. The judge evaluated every rubric criterion individually, recording whether it was satisfied and grounding the decision in the candidate response and reference materials. The judge assessed the completed output only; it did not participate in or alter the candidate run.
Deterministic checks confirmed that the required artifact existed and opened successfully. A run entered the results only after both the candidate artifact and criterion-level judgment were complete. The headline percentage is the rubric weight earned divided by the total available weight: high-importance criteria are worth 3 points, medium-importance criteria 2 points and low-importance criteria 1 point. Task results average that weighted score across completed runs; category results group the underlying criterion decisions by rubric theme.
Contract dispute strategy
A multilingual telecommunications dispute requiring a respondent pleading and a client strategy memorandum. Matter languages: English and French.
Respondent pleading
Draft a structured respondent pleading addressing the contractual history, termination, waiver, quantum, counterclaims, procedure and relief sought.
Agent performance
These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.
| Model | Score | Latency | Turns | Tool calls | Sources read | Citation support |
|---|---|---|---|---|---|---|
| Claude Sonnet 5 | 83.5% | 323.6 s | 45.0 | 80.7 | 48.3 | 100.0% |
| Claude Opus 5 | 83.1% | 968.6 s | 120.7 | 156.0 | 64.7 | 100.0% |
| Gemini 3.6 Flash | 75.5% | 139.8 s | 27.3 | 26.3 | 15.7 | 100.0% |
| GPT-6 Astra | 75.1% | 481.9 s | 13.0 | 63.0 | 40.3 | 100.0% |
| Grok 4.6 | 73.2% | 175.5 s | 10.3 | 32.3 | 17.7 | 100.0% |
| GPT-5.6 Sol | 69.7% | 167.1 s | 8.7 | 43.0 | 27.3 | 98.1% |
| GPT-5.6 Luna | 66.3% | 84.8 s | 6.0 | 23.7 | 18.3 | 100.0% |
Rubric-category performance
The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.
Claude Sonnet 5
Claude Opus 5
Gemini 3.6 Flash
GPT-6 Astra
Grok 4.6
GPT-5.6 Sol
GPT-5.6 Luna
Generated work products
Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.
Client memorandum
Prepare a confidential and privileged client memorandum on the contractual dispute, legal and financial risks, and recommended strategy, using verified facts and contractual wording.
Agent performance
These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.
| Model | Score | Latency | Turns | Tool calls | Sources read | Citation support |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 88.5% | 683.9 s | 96.7 | 129.3 | 58.3 | 100.0% |
| GPT-5.6 Sol | 81.7% | 125.3 s | 7.3 | 30.3 | 19.0 | 100.0% |
| GPT-6 Astra | 81.0% | 450.6 s | 11.7 | 50.0 | 35.0 | 100.0% |
| Gemini 3.6 Flash | 79.8% | 147.1 s | 36.7 | 35.7 | 14.0 | 100.0% |
| Claude Sonnet 5 | 77.0% | 264.0 s | 41.7 | 67.3 | 41.7 | 100.0% |
| GPT-5.6 Luna | 73.4% | 88.8 s | 12.7 | 44.0 | 12.3 | 98.5% |
| Grok 4.6 | 67.9% | 158.6 s | 17.3 | 43.0 | 17.0 | 96.8% |
Rubric-category performance
The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.
Claude Opus 5
GPT-5.6 Sol
GPT-6 Astra
Gemini 3.6 Flash
Claude Sonnet 5
GPT-5.6 Luna
Grok 4.6
Generated work products
Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.
CAIP arbitration defence
A cross-border sale-of-goods arbitration involving force majeure, exhibit assessment and a substantive reply. Matter languages: English, French, Spanish and Chinese.
Force majeure memorandum
Prepare a five-to-six-page English memorandum on contractual force majeure under French law for a respondent in a CAIP arbitration.
Agent performance
These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.
| Model | Score | Latency | Turns | Tool calls | Sources read | Citation support |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 80.6% | 1,159.0 s | 140.0 | 180.3 | 85.0 | 100.0% |
| GPT-6 Astra | 74.3% | 366.8 s | 13.7 | 70.0 | 43.0 | 93.5% |
| GPT-5.6 Sol | 73.8% | 166.2 s | 7.0 | 46.7 | 18.7 | 100.0% |
| Claude Sonnet 5 | 71.3% | 628.7 s | 118.7 | 151.7 | 40.0 | 98.5% |
| Grok 4.6 | 68.8% | 122.6 s | 7.7 | 23.0 | 6.3 | 100.0% |
| Gemini 3.6 Flash | 64.6% | 120.1 s | 24.3 | 23.3 | 8.0 | 100.0% |
| GPT-5.6 Luna | 46.8% | 75.6 s | 5.3 | 18.0 | 11.7 | 92.2% |
Rubric-category performance
The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.
Claude Opus 5
GPT-6 Astra
GPT-5.6 Sol
Claude Sonnet 5
Grok 4.6
Gemini 3.6 Flash
GPT-5.6 Luna
Generated work products
Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.
Exhibit review
Assess the evidential weight of each exhibit MA001–MA017 submitted with the claimant’s case and present the assessment as a usable table.
Agent performance
These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.
| Model | Score | Latency | Turns | Tool calls | Sources read | Citation support |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 55.6% | 394.7 s | 14.0 | 68.3 | 37.0 | 100.0% |
| Claude Opus 5 | 54.5% | 753.5 s | 108.7 | 140.0 | 61.7 | 100.0% |
| Gemini 3.6 Flash | 50.8% | 137.7 s | 14.7 | 13.7 | 10.3 | 100.0% |
| GPT-5.6 Sol | 40.2% | 148.9 s | 5.3 | 19.7 | 11.0 | 100.0% |
| Grok 4.6 | 33.9% | 92.2 s | 6.3 | 15.3 | 7.0 | 100.0% |
| Claude Sonnet 5 | 30.7% | 519.7 s | 130.0 | 182.0 | 41.0 | 100.0% |
| GPT-5.6 Luna | 18.0% | 75.4 s | 3.3 | 9.0 | 8.0 | 100.0% |
Rubric-category performance
The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.
GPT-6 Astra
Claude Opus 5
Gemini 3.6 Flash
GPT-5.6 Sol
Grok 4.6
Claude Sonnet 5
GPT-5.6 Luna
Generated work products
Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.
Arbitration reply
Draft an approximately 20-page English reply using the supplied legal authorities RL-1–RL-23 and client factual materials R-1–R-24.
Agent performance
These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.
| Model | Score | Latency | Turns | Tool calls | Sources read | Citation support |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 83.3% | 1,291.4 s | 129.7 | 182.7 | 72.7 | 100.0% |
| Grok 4.6 | 79.0% | 192.7 s | 8.7 | 35.0 | 28.3 | 100.0% |
| Gemini 3.6 Flash | 76.5% | 252.3 s | 50.3 | 49.3 | 45.0 | 100.0% |
| GPT-5.6 Sol | 75.3% | 299.1 s | 7.0 | 72.0 | 60.0 | 100.0% |
| GPT-6 Astra | 69.8% | 734.9 s | 19.7 | 130.7 | 64.0 | 100.0% |
| Claude Sonnet 5 | 67.9% | 422.4 s | 50.7 | 107.3 | 59.7 | 100.0% |
| GPT-5.6 Luna | 63.6% | 145.3 s | 6.0 | 46.7 | 43.0 | 97.8% |
Rubric-category performance
The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.
Claude Opus 5
Grok 4.6
Gemini 3.6 Flash
GPT-5.6 Sol
GPT-6 Astra
Claude Sonnet 5
GPT-5.6 Luna
Generated work products
Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.
Arbitration clause drafting
A clause-drafting exercise for an international telecommunications agreement, applying client instructions to governing-law and dispute-resolution provisions. Matter languages: English and French.
- Client instructions
- Pseudonymized roaming framework agreement
- ICC Arbitration Rules 2026 / Mediation Rules 2014
- SCC Arbitration Rules 2023
Clause drafting
Revise governing-law and arbitration clauses in a roaming framework agreement, including French law, Paris arbitration, SCC procedure, a sole arbitrator and emergency relief.
Agent performance
These operational measures explain how the agent reached its answer. Latency is elapsed candidate-run time; turns count model-response cycles; tool calls count actions taken in the matter-room workspace; sources read counts distinct source documents accessed; and citation support measures how often the answer's cited propositions were supported by the cited material.
| Model | Score | Latency | Turns | Tool calls | Sources read | Citation support |
|---|---|---|---|---|---|---|
| Grok 4.6 | 91.3% | 57.6 s | 11.3 | 17.3 | 3.0 | 100.0% |
| Claude Opus 5 | 88.9% | 406.2 s | 44.3 | 48.7 | 13.0 | 100.0% |
| GPT-5.6 Sol | 84.9% | 72.5 s | 5.0 | 9.7 | 3.3 | 100.0% |
| Claude Sonnet 5 | 84.1% | 499.2 s | 176.7 | 182.0 | 3.0 | 100.0% |
| GPT-6 Astra | 84.1% | 259.7 s | 7.7 | 16.0 | 5.3 | 100.0% |
| Gemini 3.6 Flash | 81.0% | 49.5 s | 12.0 | 11.0 | 3.0 | 100.0% |
| GPT-5.6 Luna | 80.2% | 44.2 s | 6.3 | 12.0 | 4.3 | 100.0% |
Rubric-category performance
The category bars below show the criterion pass rate within each rubric category; they are diagnostic breakdowns and are not separately reweighted. Models appear in the same weighted-score order shown above. Categories are ordered from lowest to highest pass rate within each model.
Grok 4.6
Claude Opus 5
GPT-5.6 Sol
Claude Sonnet 5
GPT-6 Astra
Gemini 3.6 Flash
GPT-5.6 Luna
Generated work products
Download the evaluated DOCX submitted by each completed run. Evaluation traces and internal JSON files are intentionally omitted.
What changed between versions
Each release is preserved as a dated evaluation snapshot. The unversioned report URL always opens the current release, while the archived links preserve the results and methodology shown at the time.
- Expanded the evaluation from the preliminary single-task release to Cases 001–003 and their current task set.
- Added GPT-6 Astra and Gemini 3.6 Flash to the evaluated model group.
- Corrected the preliminary low results after completing the task and source configuration, then reran the evaluation under the controlled harness.
- Added the cross-case executive summary, model-level operational metrics, rubric-category analysis and downloadable work products.
- Initial preliminary single-task report.
- Preserved for audit history, including the unusually low scores later determined to be incorrect.
- Open the archived v1.0 report →