MediationBench tests AI mediators in 62 authored disputes. Language models play both parties, so the results compare model behavior in a controlled setting—not outcomes with people.

This page shows every published run, explains the current leaderboard, records the original controlled study, and collects the methods, limitations, corrections, data, and working paper.

Every public measurement, two linked views

The graph below holds every published run, from the current provisional test and the earlier studies, as a mediator-by-participant matrix and a score distribution. The same filters narrow both views, and selecting a dot or a cell opens one run’s detail. The working paper explains how these measurements were designed and what they can and cannot show.

Read the paper (PDF)

Working draft · 5 September 2026 · All revisions

Checking the judges is a separate exploratory evaluation using source annotations from external datasets.

The explorer reads the latest public export, so it may show a completed run before the leaderboard updates; its date and coverage come from that export.

Loading the public run snapshot.

Public snapshot

Loading results…

The latest benchmark scores are loading.

Evaluated systems

Baseline, simple mediation, and skilled mediation are separate tests.

No-mediator baseline

The participants continue the seeded conflict without an intervening mediator. Moderator Quality is fixed at zero, so baseline rows are reference points, not mediator products.

Reflective-listening control

Published in the historical export as HAI Simple Echo. It is a deliberately minimal active control that acknowledges and reflects what each side said. Historical rows are not necessarily backbone-matched; the v2.5 factorial uses model-matched reflective and skilled policies.

HAI Skilled Mediator

An HAI reference mediator that asks about hidden constraints, names tradeoffs, and seeks specific commitments from both parties.

How to use the explorer

  • Compare like with like. Use the same test version, simulated parties, prompt, and scoring model.
  • Use the reflective control for the main policy comparison. It shows what a mediation strategy adds beyond acknowledgment and reflection.
  • Treat the no-mediator result as an earlier reference. The current score is not presented as a difference from that row.
  • Keep stopping and scoring separate. The mediator decides when to end within the test limits. A final judge scores the completed transcript and checks whether the ending was appropriate.
  • A re-score is not another run. It applies a different judge to the same saved conversation.

How leaderboard scores are calculated

Leaderboard runs use one fixed setup: 62 authored disputes, Gemma 4 31B playing both parties, and a simple goal prompt for the mediator being tested.

When a conversation ends

The mediator is asked after the opening and after each party response whether to speak, wait, or end the conversation. The test—not a second model—enforces the limits:

  • An end can take effect only after each party has responded in three complete rounds after the opening.
  • The conversation stops after 25 complete rounds. The mediator gets one last chance to explain the ending at that point.
  • The mediator is not told these limits during the conversation.

The mediator may end because the parties agreed, because they reached another useful resolution, or because no responsible next step remains. Its claim does not establish what happened. The final judge reads the completed transcript and classifies the ending independently.

Provisional scoring

One run sends a mediator through all 62 conflicts. Each finished conversation receives a 0–100 score from GPT-5.6 Sol. A row scored by Sol alone is the average of Sol’s 62 scores. It also shows how many endings Sol classified as appropriate: a supported agreement, a useful resolution without an agreement, or a reasonable decision to stop without resolution.

Fable, Gemini, or Grok may also score all 62 saved conversations. Each extra judge is listed under the row with its own score, and once a judge has scored all 62 the row score becomes the equal-weight mean of every complete judge. One or two judges keep the row provisional; three or more do not. Rows are ordered by score for reading, and a small gap between rows is not a measured difference. A partial pass never delays the row. If a judge and mediator come from the same provider lab, the judge line says so; the judge still counts.

How scores are combined

The five score parts are cooperation, depth of resolution, useful information revealed, balance of commitments, and moderator quality. Their weights are 25%, 25%, 20%, 15%, and 15%. The ending classification is reported beside the score and does not multiply it.

The judge must point to transcript evidence for a resolved issue or accepted agreement. Silence, an unanswered offer, or vague approval does not count as agreement.

What provisional means

A provisional result is one complete run scored in full by one or two judge models. It is a useful demonstration, but it is not enough to show that a small score difference is meaningful. We therefore order provisional rows by score without calling any row qualified, and do not show a confidence range or a difference from the old baseline.

When a system is ranked

The current launch orders rows by score but calls no row qualified and makes no claim that one system beats another. Before adding a qualified tier, we will publish one fixed rule for repeated runs, a three-judge panel, model-lab conflicts, combining scores, and deciding when results are different enough to order. We will not invent those rules after seeing a result.

What the score does—and does not—mean

  • The rubric has not been validated against human mediators or participants. A high score is not proof of good mediation.
  • The 62 conflicts are a fixed synthetic test, not a sample of human disputes.
  • One run and one judge provide early evidence, not a stable ranking.
  • The benchmark tests unguided stopping. The product may guide a mediator with another model, so the benchmark and product are not the same setup.
  • A held-out set of 38 conflicts remains private and unused by leaderboard runs.

Earlier methods

Earlier published results used different simulated parties or scoring rules. In the earlier setup, one judge scored the conversation while it ran and could end it, and final judges read all 62 conflicts. Those results remain visible under Earlier results, but they are not combined with or ranked against the current provisional results. The old no-mediator run is a historical reference, not a baseline difference for the new score.

Original controlled study: what it found

MediationBench ran 62 authored two-party disputes with language models playing both sides. It compared two sets of instructions on the same mediator model: a minimal reflective-listening control and a skilled mediation policy. Each was tested with and without both parties’ private background information. One fixed judge model scored every conversation from 0 to 100.

  • What held. With Qwen as the mediator model, an open-weight model (Stheno) playing both parties, and only public context given to the mediator, the skilled instructions scored 13.17 points above the reflective-listening control. The 90% range was 9.67 to 17.55 points. This comparison was repeated twice under every condition.
  • What the information comparison showed. In that same setup, the skilled instructions scored above the control both with private briefs (+7.27) and without them (+13.17). The study did not show that the two information conditions are equivalent.
  • What did not repeat everywhere. The Kimi comparison, mostly single runs, pointed in the same direction. The gpt-oss test did not pass either prewritten rule, and the public-context result did not hold when Gemma played the parties.

How certain is this? The range shows how the result changes when the 62 observed conflicts receive different weight. It does not cover new model responses, a different judge, provider or model updates, or human disputes. Only the Qwen-mediator comparison repeats every condition.

Follow-up check using the same test version

Two new skilled-policy runs with private briefs, plus the analysis plan, were registered before either result was known. We then recalculated the Qwen comparison using two runs from the same test version in every condition.

The skilled-versus-control difference with private briefs was +9.93 points (90% range 2.65 to 13.41). The interaction (E3: the skilled-versus-control difference under public context minus that under private briefs) was +3.23 points (-1.13 to 11.93), so its direction remains unclear. The skilled policy’s public-context score minus its private-brief score was -2.66 points (-6.40 to 2.04). The public-context comparison remained +13.17 by construction because that condition did not change. This is a post-outcome sensitivity, not a replication. The original analysis remains primary. Download the aggregate-only sensitivity artifact .

Prospectively frozen study · snapshot 2026-07-24

The observed pattern varied across tested backbones.

MediationBench crosses a skilled policy with a minimal reflective-listening control under two information conditions: a full bilateral brief of researcher-authored private facts, or public context only. A backbone is the mediator model running both policies; the disputants and judge are held fixed.

Factorial readout

Model-matched information-study estimands

E1 compares the two policies under public context only; E2 compares them with the full brief; E4 compares the two information conditions within one policy. E3 = E1 − E2 = E4(skilled) − E4(reflective), up to rounding at two decimals.

Qwen is the prospectively frozen replicated core with two runs per arm. Kimi and gpt-oss are conditional checks with one new run per condition. The paper carries this table as Table 7, and these are the primary mixed-suite estimates. The suite-matched sensitivity above moves the Qwen E2, E3, and skilled-E4 values shown here; E1 and the reflective E4 are unchanged by construction, because no new data entered them.

E1 · Public-context skillQwen3-235BKimi K2.5gpt-oss-120bE2 · Full-brief skillQwen3-235BKimi K2.5gpt-oss-120bE3 · Interaction (rule: lower bound > -4)Qwen3-235BKimi K2.5gpt-oss-120bE4 · Public minus full, reflectiveQwen3-235BKimi K2.5gpt-oss-120bE4 · Public minus full, skilledQwen3-235BKimi K2.5gpt-oss-120b-20-10-401020
Points are composite-point differences; whiskers are 90% scenario-reweighting intervals. Filled dots: the replicated Qwen core (two runs per arm). Open dots: one run per condition. End caps mark fully contemporaneous contrasts; capless rows reuse a historical, protocol-matched full-brief skilled run. Dashed rule: the frozen -4-point E3 margin. Exact values are in the table below.
BackboneEvidencePublic-context skill (E1)Full-brief skill (E2)Interaction (E3)Public minus full (E4)
Qwen3-235BTwo runs per condition — the replicated frozen core Reps skilled-full / skilled-public / reflective-full / reflective-public: 2 / 2 / 2 / 2+13.17 [9.67, 17.55]+7.27 [1.31, 10.52]+5.90 [1.34, 13.81] Met the prospectively frozen four-point non-inferiority ruleReflective: -5.89 [-12.89, -2.52] Skilled: +0.01 [-4.15, 4.28]
Kimi K2.5One new run per condition; conditional Reps skilled-full / skilled-public / reflective-full / reflective-public: 2 / 1 / 1 / 1+9.66 [2.32, 16.21]+1.68 [-2.06, 5.85]+7.98 [-0.32, 15.47] Met the inherited frozen four-point rule; exploratoryReflective: -6.45 [-11.99, 0.60] Skilled: +1.53 [-1.64, 5.75]
gpt-oss-120bOne run per condition; conditional Reps skilled-full / skilled-public / reflective-full / reflective-public: 1 / 1 / 1 / 1-1.02 [-4.51, 5.01]+4.74 [2.19, 15.55]-5.75 [-16.75, 0.38] Did not meet the frozen non-inferiority ruleReflective: +4.12 [0.23, 15.22] Skilled: -1.64 [-4.59, 3.38]

Participant-axis transport check · snapshot 2026-07-26

Participant transport: the same 2×2 on Gemma participants

The frozen 2×2 was replayed with gemma-4-31b participants (cerebras) against the Qwen/Qwen3-235B-A22B-Instruct-2507 mediator backbone. Axis varied: participant configuration — mediator backbone, both policy prompts, information conditions, judge, scenario set, and suite held fixed. Policy x information 2x2 plus a scenario-matched no-mediator baseline; one run per arm (1/1/1/1); exploratory post-launch, pre-outcome amendment frozen 2026-07-25T04:19Z before any arm had evaluated scenarios.

Composite scores: no-mediator baseline 15; reflective 34 / 35; skilled 41 / 38 (full brief / public context). All four mediated arms scored above the baseline with 90% intervals excluding zero; per-run identifiers and baseline deltas are in the aggregate download.

Gemma participant-transport E1 to E4 contrasts with 90 percent intervals
Contrast (this participant column)Estimate · 90% interval
E1 · Public-context skill+5.70 [-3.79, 12.90] Failed the prospectively frozen positivity rule
E2 · Full-brief skill+11.39 [5.77, 16.92]
E3 · Interaction (rule: lower bound > -4)-5.68 [-15.95, 2.30] Failed the prospectively frozen four-point non-inferiority rule
E4 · Public minus full, reflective+0.73 [-3.66, 6.34]
E4 · Public minus full, skilled-4.95 [-13.69, 1.94]

Single run per arm, exploratory. Both prospectively frozen reads failed on this column: the public-context skill increment did not separate, and the interaction failed the -4-point non-inferiority margin. This is a participant-axis transport check, never a fourth row of the E1-E4 table above; the paper carries it as Table 8. First failed frozen-rule outcome on the participant axis (gpt-oss failed both rules on the mediator-backbone axis). Mediation still cleared the no-mediator baseline in all four arms. One run per arm; this does not establish that the private brief is necessary. This is a participant column, not a mediator-backbone row.

These are synthetic-dispute results, not evidence of effectiveness with human parties. The whiskers reweight the 62 observed scenarios; they do not estimate variability from fresh generations, different judges, provider or model updates, or human disputes. Only Qwen repeats every condition. Fine model rankings remain judge- and run-sensitive; the conditionality and historical-reuse boundaries are stated once in What is not established.

Analysis provenance
Design registry
2.5.0
Suite provenance
The v2.5 information-study registry combines contemporaneous v2.5 cells with protocol-matched historical full-brief skilled arms that retain their original suite stamps.
Snapshot
2026-07-24
Joint scenarios
62
Scale
62 scenarios × 17 study runs ≈ 1,050 scored conversations
Estimand and intervals
Each point estimate is a difference between arm medians of per-scenario replicate means across 62 scenarios jointly completed in all four arms. Each 90% percentile interval comes from the four-arm scenario-co-resampled bootstrap (B=10000; seed=20260706): scenarios are resampled once per draw across all arms, each arm's observed per-scenario replicate mean is held fixed, and every draw recomputes the point-estimate statistic. An earlier implementation independently resampled replicate identity within scenarios, creating hybrid pseudo-runs; this correction changes intervals only. All intervals are conditional on the observed generation runs and operational judge; they are not run-sampling or human-efficacy intervals.
Participant model
Sao10K/L3-8B-Stheno-v3.2
Operational judge
Qwen/Qwen3-235B-A22B-Instruct-2507
Methodology
seeded_25_v1
Sampling
unseeded
Temperatures
judge 0.3 / mediator 0.7 / participant 0.8
Decision rules
Two decision rules were frozen 2026-07-15, before any public-context run existed: E1 positivity and E3 non-inferiority. E1 passes when its 90% interval lower bound is greater than 0 composite points. E3 passes when its 90% interval lower bound is greater than -4.0 composite points. This is a non-inferiority rule, not an equivalence test; an interval spanning zero does not establish equivalence.
Source commit
abaff7086941f2d4a0da129e38b7019d88933db3 (internal HAI repository, not publicly resolvable)
Analysis artifact
docs/research/analysis/compute_claims_2026-07-28.json (internal HAI repository, not publicly resolvable; all 17 study run IDs, the five participant-transport run IDs, and the GPT-5.5 extension-run ID resolve in the public export, alongside the cross-family rejudge series)
Artifact SHA-256
c1c53164a1441186e7369b8a1a3592a3dcf55021510322a68f0ef1d0c7e266b9
Download SHA-256
ba6f96938f767d0609469d551a65cbd9d8c73d73df1e6022c661de4f9d83d995
Generation completion
All six newly frozen generation cells completed 62 of 62 scenarios after repair, with zero remaining failures. Every regeneration was triggered by completion status alone - a failed scenario or a zero-token generation - and never by the score a run produced (attested by the study author, 2026-07-26).

How the study is organized

The study table labels four comparisons E1 through E4 and defines each in its caption. A negative E4 means the mediator scored higher when it received the private briefs. A condition is one combination of instructions and information. A run covers every conflict in one condition. The judge is the model that scored every conversation in this original study.

Extensions and stress tests

  • With gpt-oss as the mediator, the public-context range included zero and the information comparison missed the prewritten -4-point limit. It passed neither rule.
  • When Gemma played the parties, the public-context result did not hold under the prewritten rules. The interaction was -5.68 points. The participant-results table above shows all five runs and comparisons.
  • Two judges from other model families later scored those same transcripts and also found a negative interaction. The limits explain why this supports the direction, not the size, of the finding.
  • A preplanned Sonnet check tested whether showing a judge the parties’ private briefs changed its scores. All 558 expected run-conflict pairs were scored, and all 22 reported 90% ranges included zero. The direction remains unresolved; this is not evidence that the two conditions are equivalent. The briefs-visible and briefs-withheld scores came from different serving periods, and the response-model revision was not retained. Download the aggregate artifact .
  • A registered post-freeze GPT-5.5 skilled/full-brief extension scored 50 on the Gemma participant setup—one run, without a matching reflective-listening control and outside the original 2 × 2 study. Selected follow-up results can look unusually strong by chance.

The information experiment

The original study compared two information conditions:

  • Private briefs: the public conflict, shared facts, conversation so far, and both parties’ researcher-written private facts, word for word.
  • Public context only: the public conflict, shared facts, and conversation so far. The mediator must discover any relevant private constraint during the exchange.

Every mediator also received researcher-written focus areas. In 3 of the 100 authored conflicts, a focus area substantially overlaps an expected revelation. Those public-context cases therefore contain a directional hint, not an answer key.

Across Qwen, Kimi, and gpt-oss, the skilled policy’s public-context score minus its private-brief score was close to zero (+0.01, +1.53, and -1.64), and every 90% range included zero. This does not establish equivalence: no equivalence limit was set, and the ranges still allow meaningful differences. For Qwen, the same-version follow-up changed the estimate from +0.01 to -2.66.

Methodology

In the E1–E4 table, each model receives both the skilled and reflective-listening instructions. The simulated parties and judge stay fixed, so each comparison is between two instruction sets on the same model. Exact model IDs, test version, temperatures, and resampling details appear in the table’s provenance panel. Model generation was not seeded.

Runs are regenerated only when a conflict fails or a model returns no text, never because of the resulting score. The study author recorded this rule on 26 July 2026.

How synthetic disputes are generated

Each authored conflict has exactly two parties. Public material establishes the dispute and opening exchange. Private material gives each party a backstory, position, hidden interests, and a walk-away alternative. Every matched condition uses the same conflict.

The evaluated set contains 62 English text conflicts from a 100-conflict collection that also informed benchmark development, so it is not a clean holdout. The remaining 38 are private and unused. This controlled setup gives each party known private information and allows matched comparisons. It does not prove that a model acted from every assigned fact or make the simulated parties representative of people.

Construct and scoring

MediationBench measures conversation process, not only whether a settlement appears at the end: whether the dispute moves from zero-sum toward cooperative, the movement the philosophy sets out as the construct. Scores use a 0–100 MediationBench composite. Every E1–E4 estimate and the -4-point limit is a difference on that scale. Its expert-chosen weights are:

  • Cooperative Dimensions (cooperative movement between the parties), 25%;
  • Resolution Depth (how fully the dispute is resolved), 25%;
  • Hidden Revelations (discovery of private interests), 20%;
  • Commitment Symmetry (reciprocal commitment), 15%;
  • Moderator Quality (mediator process quality), 15%.

A high composite is not necessarily good mediation. This score does not assess voluntary consent, procedural fairness, power imbalance, safe non-agreement, legality, feasibility, or durability; a Hidden Revelations score does not establish that disclosure was safe or authorized.

The test setup has known effects. Mediated conversations may stop earlier than no-mediator conversations. Mediator messages do not use up party turns. Moderator Quality is zero without a mediator and high for almost every mediated run, so a mediator-versus-reference difference partly reflects that built-in gap. The original analysis treats the ending rules and score design as part of what was tested.

Two decision rules were frozen 2026-07-15, before any public-context run existed: E1 positivity, whose interval lower bound must exceed zero, and E3 non-inferiority, whose lower bound must exceed -4 points — the skill increment under public context only must not fall more than 4 composite points below the increment measured with the full bilateral brief. An interval crossing zero does not establish equivalence.

What is not established

  • All parties are language models, not people. Every range is limited to the observed runs and judge. It does not cover fresh model responses, different serving systems, or human disputes.
  • The Qwen mediator is the same model checkpoint used as the judge, so the replicated core is self-judged: a constant self-preference offset cancels in the within-backbone contrast, a style-varying one would not, and cross-backbone readings including the gpt-oss non-replication are unprotected. The export flags degraded judge reliability on two of five categories (κ 0.42 and 0.48, n=40), 45% of the composite weight.
  • Two other model families later scored the fixed transcripts and also found a negative Gemma-party interaction (-8.50 and -4.50 versus -5.68 from the original judge), corroborating its direction, not its magnitude; only the gpt-oss interval excludes zero (crossed-judge readout ). Scale and range width still depends on the judge. Later scoring cannot change a conversation or an ending decision already influenced by the original judge.
  • Qwen is the only replicated information core; Kimi (one run in three of its four conditions) and gpt-oss (one run in each) estimate scenario, not run-to-run generation, variation. For every tested backbone, E2, E3, and the skilled-policy information contrast reuse an earlier same-setup private-brief skilled run, so the factorial is not wholly contemporaneous; the cleanest contemporaneous readings are the public-context skill increments and the reflective-policy information effects.
  • The estimands and rules were frozen in an internal registry, not a conventional public preregistration, and the -4-point E3 margin rationale used already-observed full-brief results; no primary estimand, margin, or rule changed after the public-context outcomes.
  • No human criterion study establishes that the composite corresponds to professional mediator or participant judgments, so the study supports configuration-specific policy comparisons, not a universal model-family ranking or a claim of safe deployment.
  • HAI.AI designed and operates the benchmark, develops mediation policies and the agreement factory at hai.ai, and publishes this site. These results are not an independent third-party evaluation.

For model labs

MediationBench publishes aggregate results and its E1–E4 data. The run harness, judge prompts, scenario-level records, transcripts, and held-out scenarios are not, so third parties cannot independently recompute E1-E4 or their intervals. Model labs can arrange a new evaluation by email to:

  • estimate your mediation prompt’s score contrast against our reflective control;
  • measure what a bilateral private-fact brief changes;
  • inspect failures by scenario type, difficulty, and participant resistance;
  • re-score fixed transcripts without regenerating the conversation.

Useful feedback includes missing conflict domains, more credible participant manipulations, human-review design, and the metadata needed for independent reproduction. To discuss an evaluation or methodology contribution, contact hello@hai.io .

Publication boundary

The public site publishes aggregate, versioned evidence. It does not publish the held-out set, raw private briefs, unreleased prompts, private transcripts, canary material, or personally identifying information.

A sample of the benchmark data is published on Hugging Face: mediationbench-sample .

Publication does not depend on whether a result looks good: we publish every registered clean run and every registered re-score, including null or negative results. Re-scoring an existing transcript with another judge does not count as a new run.

Full-file SHA-256 checksums for every circulated dated JSON and PDF are in the public checksum manifest .

Every run published here is HAI’s own: HAI commissioned it, on HAI’s scenarios, under the conditions described above. No outside party has submitted a system to MediationBench. Evaluation of an outside system is arranged directly with HAI by email, and what may be published from such an evaluation is agreed at the time it is arranged rather than by any standing rule on this page.

Terms and data use

Participation is optional. HAI does not sell personal data and does not publish personally identifying information. Use of hai.ai is governed by the HAI.AI Terms of Service .

Corrections and responses

Operators of evaluated systems may request a factual correction or submit a response for publication by contacting hello@hai.io .

For a safe public fixture excerpt, read the sample conversation . To propose a synthetic dispute, see Contribute a Scenario .

Paper

The paper gives the full study design, analysis, and limits. The current revision (5 September 2026) republishes the merged paper with Cooperative AI context and the prepared 200-case CaSiNo and ContractNLI judge pilot. Every numeric claim is unchanged from the 4 September 2026 revision. It remains a working draft, so numbers may change before the preprint is complete. This page and the dated aggregate data files are the citable web record.

Paper revision 2026-09-05 (5 September 2026)

Download the current working draft (PDF)

Draft revisions

License

Two licenses apply. The aggregate results — the E1-E4 table, the current versioned aggregate JSON , the Qwen suite-matched sensitivity , the Sonnet evaluator-context readout , and the crossed-judge readout — are licensed under CC BY 4.0 : reuse the numbers anywhere, including commercially, with attribution. The narrative text and the figures on this page remain licensed under CC BY-NC-SA 4.0 .

Neither license covers the hidden evaluation set, private transcripts, raw prompts, or scoring internals.

How to cite: Jonathan Hendler. MediationBench. Human Assisted Intelligence, PBC. Frozen study snapshot 2026-07-24. https://mediationbench.com/leaderboard/