MediationBench is a synthetic test harness for auditing where AI mediation-policy claims fail to generalize. This page carries the provisional standings and the public run explorer; the findings, methods, and claim boundaries live in the research record and the working-draft paper (PDF) .

Standings at a glance

Current standings · preview · updated 2026-07-31

Mediator systems, HAI Score composite (0–100)

Provisional — not yet leaderboard-eligible

  1. 1GLM-5.2 Z.ai 52
  2. 2Sonnet 5 Anthropic 51
  3. 3gpt-oss-120b OpenAI 49
  4. 4Qwen3-235B Alibaba 46 runs 48 · 44
    with Gemma 4 31B participants (suite 2.5) 39.5 runs 41 · 38
  5. 5Kimi K2.5 Moonshot AI 45 runs 46 · 44

Reference rows · not ranked

  • No-mediator baseline Stheno 8B pair, no mediator 17 runs 18 · 16
  • Reflective control Sonnet 4.5, reflective-listening policy 40.5 runs 41 · 40

Rows are single-run, single-judge composites from suite-2.0 generation runs with the Stheno 8B participant pair: each mark is one generation run scored once by the operational judge on the HAI Score composite (0–100) — not panel medians. Where a system has two runs the bar shows the mean and the ticks mark the individual runs; the labeled Qwen sub-row swaps in the suite-2.5 Gemma 4 31B participant pair. Ranked status requires 5 runs per cell under the frozen leaderboard rules, which publish at leaderboard launch. The full study record is on the research page. Reference rows are not ranked.

Standings are configuration-specific score contrasts inside a synthetic scoring harness. Treat small gaps as ties: re-runs and alternative evaluators changed some orderings, so this is not a stable model ranking.

Public run explorer

The explorer shows the currently published rolling run snapshot, separate from the frozen E1-E4 readout on the research page . A snapshot may lag a newly completed study cell until the publication step creates the next version; the provenance strip below carries the export’s own counts and dates, which are authoritative for current coverage. The public export does not yet stamp an information field, so the explorer labels this study’s runs from the frozen E1-E4 registry and reports other rows as not recorded.

Publication is outcome-independent by rule: every registered clean generation arm, and every registered fixed-transcript rejudge series, is queued for publication, including public-context-only arms and null or negative outcomes. A rejudge re-scores an existing transcript and is never counted as a new generation run or generation replicate.

Loading the public run snapshot.

Public snapshot

Loading results…

The latest aggregate HAI Score snapshot is loading.

Evaluated systems

Baseline, simple mediation, and skilled mediation are separate tests.

No-mediator baseline

The participants continue the seeded conflict without an intervening mediator. Moderator Quality is fixed at zero, so baseline rows are reference points, not mediator products.

Reflective-listening control

Published in the historical export as HAI Simple Echo. It is a deliberately minimal active control that acknowledges and reflects what each side said. Historical rows are not necessarily backbone-matched; the v2.5 factorial uses model-matched reflective and skilled policies.

HAI Skilled Mediator

The higher-scoring HAI reference mediator in this harness. It probes hidden constraints, names tradeoffs, and tries to move the parties toward specific reciprocal commitments.

How to read the run browser

  • Version first. Compare runs inside the same suite, fixture version, and information condition.
  • The active control is primary. Skilled minus reflective control asks what mediation strategy adds beyond acknowledgment and reflection; reflective control minus no mediator asks what basic facilitation contributes.
  • Baseline delta is practical, not statistical. It is a descriptive difference from a matching no-mediator row, and baseline carries no mediator-quality score. Treat small gaps as ties: re-runs and alternative evaluators changed some orderings, so this is not a stable model ranking.
  • Three judge families on the same transcripts. Every suite-2.5.0 study run is scored by the operational Qwen judge and re-scored by Gemma-4-31B and gpt-oss-120b judges, with deltas against the baseline re-scored by the same judge. The default view pools every judge and measurement class, so select a single judge and the generation class before comparing runs.
  • Evaluator context is part of the instrument. The preplanned fixed-transcript Sonnet context-withheld re-score is complete — 558 of 558 expected run-scenario pairs — and all 22 reported 90% intervals included zero. That leaves the context effect directionally unresolved; it is not evidence of equivalence. Two caveats: visible and withheld scores came from different serving epochs, and the provider’s response-model revision was not retained, so this is not a clean isolation of evaluator context alone. The aggregate-only context readout carries the paired analysis.

Every measurement belongs to a named configuration, and a score is not a certification or permanent judgment of a named model. The current versioned E1-E4 aggregate JSON downloads separately from the rolling explorer snapshot; it carries the aggregate estimates and provenance, but does not contain the scenario-level records needed to reproduce the intervals. This working-draft revision r3 leaves r2 byte-for-byte unchanged and adds the committed HAI analysis revision abaff7086941f2d4a0da129e38b7019d88933db3 alongside the producing artifact and full SHA-256. Superseded files remain available unchanged: 2026-07-26 , 2026-07-26-r2 , 2026-07-24 , and 2026-07-24-r2 .

The full research record

The complete study record lives on the research page : the prospectively frozen 2 × 2 study and its E1-E4 estimands, the suite-matched sensitivity, participant transport, study design, methodology, construct and scoring, what is not established, the publication boundary, corrections, and licensing. For the complete methods and claim boundaries, read the working-draft paper (PDF) ; dated paper revisions are listed on the research page .


Run a system through MediationBench via hai.ai ; to arrange an evaluation, email hello@hai.io . This site publishes results and does not accept evaluation-run submissions directly.

Frequently asked questions

What is MediationBench?
MediationBench is a controlled synthetic-conflict testbed for evaluating AI mediation policies. It compares the same authored two-party dispute and opening exchange across no-mediator, reflective-control, and skilled-mediator conditions while recording the participant, mediator, information, and judge configuration.
How certain are these results?
The reported 90 percent intervals reweight the 62 observed scenarios while holding the observed generation runs and operational judge fixed. They do not estimate variability from fresh generations, a different judge, provider or model updates, or human disputes. Only the Qwen information study repeats every condition; Kimi and gpt-oss are single-run checks.
Does MediationBench prove that AI mediation works with people?
No. Every dispute in this study is synthetic and every party is played by a language model. The results test controlled model behavior, not safety or effectiveness with human parties.