Checking the judges
Can our judges identify agreed terms and interpret claims and supporting text? This pilot checks their answers against annotations from external datasets. Each source tests a different skill, so results stay separate.
Exploratory results · 200 cases
Some cases remain unresolved
Results are shown for each judge and task. Failed and pending cases remain visible in their denominators.
Sample coverage
200 source records · 2 datasets
Easy · 68 Medium · 66 Hard · 66
100 cases
100 cases
Only CaSiNo records are conversations. ContractNLI provides documents and propositions.
Download aggregate (JSON) · Dataset subsets on Hugging Face
Earlier reading, kept as published: External-source judge pilot: first judge readings.
Results by judge and task
“Correct” is strict: the label agrees with the source annotation and every quoted citation verifies. “Label agreement” counts a matching label whether or not its citation verified. Failed responses are not counted as correct or incorrect.
Sol
200 cases completed · 0 failed · 0 pending
Scroll the table to see all counts.
| Task | Correct | Label agreement | Incorrect | Failed | Pending |
|---|---|---|---|---|---|
| CaSiNoRecorded deal acceptance100 cases | 75 | 75 | 25 | 0 | 0 |
| CaSiNoExact allocation100 cases | 69 | 69 | 31 | 0 | 0 |
| ContractNLIContract entailment100 cases | 77 | 77 | 23 | 0 | 0 |
By source answer
The same counts split by what the source annotation records for each case. For CaSiNo, “true” and “deal” are recorded deals; “false” and “no_deal” are recorded walk-aways.
- CaSiNo · Recorded deal acceptance · false: 3 correct, 3 incorrect, 0 failed, 0 pending.
- CaSiNo · Recorded deal acceptance · true: 72 correct, 22 incorrect, 0 failed, 0 pending.
- CaSiNo · Exact allocation · deal: 66 correct, 28 incorrect, 0 failed, 0 pending.
- CaSiNo · Exact allocation · no_deal: 3 correct, 3 incorrect, 0 failed, 0 pending.
- ContractNLI · Contract entailment · Contradiction: 21 correct, 12 incorrect, 0 failed, 0 pending.
- ContractNLI · Contract entailment · Entailment: 30 correct, 4 incorrect, 0 failed, 0 pending.
- ContractNLI · Contract entailment · NotMentioned: 26 correct, 7 incorrect, 0 failed, 0 pending.
Evidence-reference checks
These counts check each response’s references to source text, once per case. They do not measure whether the cited text supports each answer.
- CaSiNo: 100 valid, 0 invalid.
- ContractNLI: 100 valid, 0 invalid.
Gemini
198 cases completed · 2 failed · 0 pending
Scroll the table to see all counts.
| Task | Correct | Label agreement | Incorrect | Failed | Pending |
|---|---|---|---|---|---|
| CaSiNoRecorded deal acceptance100 cases | 71 | 71 | 29 | 0 | 0 |
| CaSiNoExact allocation100 cases | 67 | 67 | 33 | 0 | 0 |
| ContractNLIContract entailment100 cases | 82 | 84 | 16 | 2 | 0 |
By source answer
The same counts split by what the source annotation records for each case. For CaSiNo, “true” and “deal” are recorded deals; “false” and “no_deal” are recorded walk-aways.
- CaSiNo · Recorded deal acceptance · false: 4 correct, 2 incorrect, 0 failed, 0 pending.
- CaSiNo · Recorded deal acceptance · true: 67 correct, 27 incorrect, 0 failed, 0 pending.
- CaSiNo · Exact allocation · deal: 63 correct, 31 incorrect, 0 failed, 0 pending.
- CaSiNo · Exact allocation · no_deal: 4 correct, 2 incorrect, 0 failed, 0 pending.
- ContractNLI · Contract entailment · Contradiction: 24 correct, 9 incorrect, 0 failed, 0 pending.
- ContractNLI · Contract entailment · Entailment: 29 correct, 3 incorrect, 2 failed, 0 pending.
- ContractNLI · Contract entailment · NotMentioned: 29 correct, 4 incorrect, 0 failed, 0 pending.
Evidence-reference checks
These counts check each response’s references to source text, once per case. They do not measure whether the cited text supports each answer.
- CaSiNo: 100 valid, 0 invalid.
- ContractNLI: 98 valid, 2 invalid.
Grok
196 cases completed · 4 failed · 0 pending
Scroll the table to see all counts.
| Task | Correct | Label agreement | Incorrect | Failed | Pending |
|---|---|---|---|---|---|
| CaSiNoRecorded deal acceptance100 cases | 69 | 71 | 28 | 3 | 0 |
| CaSiNoExact allocation100 cases | 64 | 66 | 33 | 3 | 0 |
| ContractNLIContract entailment100 cases | 77 | 78 | 22 | 1 | 0 |
By source answer
The same counts split by what the source annotation records for each case. For CaSiNo, “true” and “deal” are recorded deals; “false” and “no_deal” are recorded walk-aways.
- CaSiNo · Recorded deal acceptance · false: 4 correct, 2 incorrect, 0 failed, 0 pending.
- CaSiNo · Recorded deal acceptance · true: 65 correct, 26 incorrect, 3 failed, 0 pending.
- CaSiNo · Exact allocation · deal: 60 correct, 31 incorrect, 3 failed, 0 pending.
- CaSiNo · Exact allocation · no_deal: 4 correct, 2 incorrect, 0 failed, 0 pending.
- ContractNLI · Contract entailment · Contradiction: 21 correct, 11 incorrect, 1 failed, 0 pending.
- ContractNLI · Contract entailment · Entailment: 29 correct, 5 incorrect, 0 failed, 0 pending.
- ContractNLI · Contract entailment · NotMentioned: 27 correct, 6 incorrect, 0 failed, 0 pending.
Evidence-reference checks
These counts check each response’s references to source text, once per case. They do not measure whether the cited text supports each answer.
- CaSiNo: 97 valid, 3 invalid.
- ContractNLI: 99 valid, 1 invalid.
Shared results across judges
Over the cases every judge labeled (sol, gemini, grok), counted at the label level. Cases where no judge matched the source annotation are pending human review of the source material; they are not evidence for either side.
| Task | Labeled by all | All agree with source | None agree with source | Of which same label |
|---|---|---|---|---|
| CaSiNoRecorded deal acceptance | 100 | 68 | 24 | 24 |
| CaSiNoExact allocation | 100 | 64 | 29 | 28 |
| ContractNLIContract entailment | 100 | 73 | 14 | 13 |
None agree, by source answer
- CaSiNo · Recorded deal acceptance · false: 2.
- CaSiNo · Recorded deal acceptance · true: 22.
- CaSiNo · Exact allocation · deal: 27.
- CaSiNo · Exact allocation · no_deal: 2.
- ContractNLI · Contract entailment · Contradiction: 9.
- ContractNLI · Contract entailment · Entailment: 1.
- ContractNLI · Contract entailment · NotMentioned: 4.
Two readings
The first reading (2026-09-05) showed the judges the recorded deal actions, so the CaSiNo acceptance task was extraction, and three defects in our instrument shaped the rest: a contradictory allocation instruction, mixed task schemas, and a circuit breaker that treated wrong-shape answers as outages. The second reading (2026-09-06) withholds every deal action, sends one task schema per request, records wrong-shape answers as judge failures, and interleaves dispatch across sources and difficulty bands. Several changes landed together, so no difference between readings is attributed to one of them. Both readings stay published; the earlier one is linked above.
“Correct” is strict: the label matches the source annotation and every quoted citation verifies. Where a citation failed but the label matched, the label agreement column keeps that count separately. Neither number replaces the other, and no ordering of judges is established.
What this check does not test
The prompts ask for an outcome: whether a deal was accepted and on what split, or whether a contract supports a statement. They do not ask what our production judges are asked about a negotiation, namely what became clearer, which terms moved toward agreement, and what remains unresolved. These results therefore say nothing about whether the judges recognize progress in a human negotiation. Cases where every judge disagreed with the source annotation are counted in the shared-results table; they are pending human review, not evidence against either side.
What this check establishes
The sample spans easier and harder cases using source-based selection rules. Difficulty is a heuristic, not a human rating or an observed judge result. Selection happens before judges run.
Source annotations provide reference answers. They do not establish whether a person is ready to sign, has authorized a concession, or has completed an agreement. ContractNLI tests statements about contract text, not legal advice or enforceability.
This is an exploratory sample, not a representative estimate of product performance. These results do not change the mediator leaderboard .