Can our judges identify agreed terms and interpret claims and supporting text? This pilot checks their answers against annotations from external datasets. Each source tests a different skill, so results stay separate.

Exploratory results · 200 cases

Some cases remain unresolved

Results are shown for each judge and task. Failed and pending cases remain visible in their denominators.

Sample coverage

200 source records · 2 datasets

Easy · 68 Medium · 66 Hard · 66

CaSiNo

Human negotiations; agreed allocations.

CC-BY-4.0

100 cases

ContractNLI

Contract documents and propositions.

CC-BY-4.0

100 cases

Only CaSiNo records are conversations. ContractNLI provides documents and propositions.

Download aggregate (JSON) · Dataset subsets on Hugging Face

Earlier reading, kept as published: External-source judge pilot: first judge readings.

Results by judge and task

“Correct” is strict: the label agrees with the source annotation and every quoted citation verifies. “Label agreement” counts a matching label whether or not its citation verified. Failed responses are not counted as correct or incorrect.

Sol

200 cases completed · 0 failed · 0 pending

Scroll the table to see all counts.

Sol results against each source annotation, including every task denominator
TaskCorrectLabel agreementIncorrectFailedPending
CaSiNoRecorded deal acceptance100 cases75752500
CaSiNoExact allocation100 cases69693100
ContractNLIContract entailment100 cases77772300
By source answer

The same counts split by what the source annotation records for each case. For CaSiNo, “true” and “deal” are recorded deals; “false” and “no_deal” are recorded walk-aways.

  • CaSiNo · Recorded deal acceptance · false: 3 correct, 3 incorrect, 0 failed, 0 pending.
  • CaSiNo · Recorded deal acceptance · true: 72 correct, 22 incorrect, 0 failed, 0 pending.
  • CaSiNo · Exact allocation · deal: 66 correct, 28 incorrect, 0 failed, 0 pending.
  • CaSiNo · Exact allocation · no_deal: 3 correct, 3 incorrect, 0 failed, 0 pending.
  • ContractNLI · Contract entailment · Contradiction: 21 correct, 12 incorrect, 0 failed, 0 pending.
  • ContractNLI · Contract entailment · Entailment: 30 correct, 4 incorrect, 0 failed, 0 pending.
  • ContractNLI · Contract entailment · NotMentioned: 26 correct, 7 incorrect, 0 failed, 0 pending.
Evidence-reference checks

These counts check each response’s references to source text, once per case. They do not measure whether the cited text supports each answer.

  • CaSiNo: 100 valid, 0 invalid.
  • ContractNLI: 100 valid, 0 invalid.

Gemini

198 cases completed · 2 failed · 0 pending

Scroll the table to see all counts.

Gemini results against each source annotation, including every task denominator
TaskCorrectLabel agreementIncorrectFailedPending
CaSiNoRecorded deal acceptance100 cases71712900
CaSiNoExact allocation100 cases67673300
ContractNLIContract entailment100 cases82841620
By source answer

The same counts split by what the source annotation records for each case. For CaSiNo, “true” and “deal” are recorded deals; “false” and “no_deal” are recorded walk-aways.

  • CaSiNo · Recorded deal acceptance · false: 4 correct, 2 incorrect, 0 failed, 0 pending.
  • CaSiNo · Recorded deal acceptance · true: 67 correct, 27 incorrect, 0 failed, 0 pending.
  • CaSiNo · Exact allocation · deal: 63 correct, 31 incorrect, 0 failed, 0 pending.
  • CaSiNo · Exact allocation · no_deal: 4 correct, 2 incorrect, 0 failed, 0 pending.
  • ContractNLI · Contract entailment · Contradiction: 24 correct, 9 incorrect, 0 failed, 0 pending.
  • ContractNLI · Contract entailment · Entailment: 29 correct, 3 incorrect, 2 failed, 0 pending.
  • ContractNLI · Contract entailment · NotMentioned: 29 correct, 4 incorrect, 0 failed, 0 pending.
Evidence-reference checks

These counts check each response’s references to source text, once per case. They do not measure whether the cited text supports each answer.

  • CaSiNo: 100 valid, 0 invalid.
  • ContractNLI: 98 valid, 2 invalid.

Grok

196 cases completed · 4 failed · 0 pending

Scroll the table to see all counts.

Grok results against each source annotation, including every task denominator
TaskCorrectLabel agreementIncorrectFailedPending
CaSiNoRecorded deal acceptance100 cases69712830
CaSiNoExact allocation100 cases64663330
ContractNLIContract entailment100 cases77782210
By source answer

The same counts split by what the source annotation records for each case. For CaSiNo, “true” and “deal” are recorded deals; “false” and “no_deal” are recorded walk-aways.

  • CaSiNo · Recorded deal acceptance · false: 4 correct, 2 incorrect, 0 failed, 0 pending.
  • CaSiNo · Recorded deal acceptance · true: 65 correct, 26 incorrect, 3 failed, 0 pending.
  • CaSiNo · Exact allocation · deal: 60 correct, 31 incorrect, 3 failed, 0 pending.
  • CaSiNo · Exact allocation · no_deal: 4 correct, 2 incorrect, 0 failed, 0 pending.
  • ContractNLI · Contract entailment · Contradiction: 21 correct, 11 incorrect, 1 failed, 0 pending.
  • ContractNLI · Contract entailment · Entailment: 29 correct, 5 incorrect, 0 failed, 0 pending.
  • ContractNLI · Contract entailment · NotMentioned: 27 correct, 6 incorrect, 0 failed, 0 pending.
Evidence-reference checks

These counts check each response’s references to source text, once per case. They do not measure whether the cited text supports each answer.

  • CaSiNo: 97 valid, 3 invalid.
  • ContractNLI: 99 valid, 1 invalid.

Shared results across judges

Over the cases every judge labeled (sol, gemini, grok), counted at the label level. Cases where no judge matched the source annotation are pending human review of the source material; they are not evidence for either side.

Cross-judge agreement with the source annotation, per task
TaskLabeled by allAll agree with sourceNone agree with sourceOf which same label
CaSiNoRecorded deal acceptance100682424
CaSiNoExact allocation100642928
ContractNLIContract entailment100731413
None agree, by source answer
  • CaSiNo · Recorded deal acceptance · false: 2.
  • CaSiNo · Recorded deal acceptance · true: 22.
  • CaSiNo · Exact allocation · deal: 27.
  • CaSiNo · Exact allocation · no_deal: 2.
  • ContractNLI · Contract entailment · Contradiction: 9.
  • ContractNLI · Contract entailment · Entailment: 1.
  • ContractNLI · Contract entailment · NotMentioned: 4.

Two readings

The first reading (2026-09-05) showed the judges the recorded deal actions, so the CaSiNo acceptance task was extraction, and three defects in our instrument shaped the rest: a contradictory allocation instruction, mixed task schemas, and a circuit breaker that treated wrong-shape answers as outages. The second reading (2026-09-06) withholds every deal action, sends one task schema per request, records wrong-shape answers as judge failures, and interleaves dispatch across sources and difficulty bands. Several changes landed together, so no difference between readings is attributed to one of them. Both readings stay published; the earlier one is linked above.

“Correct” is strict: the label matches the source annotation and every quoted citation verifies. Where a citation failed but the label matched, the label agreement column keeps that count separately. Neither number replaces the other, and no ordering of judges is established.

What this check does not test

The prompts ask for an outcome: whether a deal was accepted and on what split, or whether a contract supports a statement. They do not ask what our production judges are asked about a negotiation, namely what became clearer, which terms moved toward agreement, and what remains unresolved. These results therefore say nothing about whether the judges recognize progress in a human negotiation. Cases where every judge disagreed with the source annotation are counted in the shared-results table; they are pending human review, not evidence against either side.

What this check establishes

The sample spans easier and harder cases using source-based selection rules. Difficulty is a heuristic, not a human rating or an observed judge result. Selection happens before judges run.

Source annotations provide reference answers. They do not establish whether a person is ready to sign, has authorized a concession, or has completed an agreement. ContractNLI tests statements about contract text, not legal advice or enforceability.

This is an exploratory sample, not a representative estimate of product performance. These results do not change the mediator leaderboard .