This page is the durable research record behind the leaderboard page : the prospectively frozen 2 × 2 study and its E1-E4 estimands, the suite-matched sensitivity, participant transport, study design, methodology, construct and scoring, what is not established, and the working-draft paper. Every dispute in this study is synthetic and every party is played by a language model; the results test controlled model behavior, not safety or effectiveness with human parties.

What the study found

MediationBench ran 62 authored two-party disputes with both sides played by language models. Two mediation policies — a minimal reflective-listening control and a skilled policy — run as different instructions on the same backbone model, under two information conditions. A fixed judge model scores each conversation on a 0–100 composite.

  • What held. In the replicated Qwen public-context comparison, the skilled policy scored +13.17 points on the 0–100 composite above the model-matched Reflective Listening control (90% scenario-reweighting interval [9.67, 17.55]).
  • What the information comparison showed. In that Qwen configuration, the skilled-minus-reflective difference remained positive with a full bilateral brief (+7.27) and with public context only (+13.17). The study did not establish that the two information conditions are equivalent.
  • What did not generalize. A single-run Kimi check pointed in the same direction, while gpt-oss met neither frozen rule and the public-context result did not transport to the second participant configuration.

How certain is this? The interval reweights the 62 observed scenarios. It does not estimate variability from fresh generations, a different judge, provider or model updates, or human disputes. Only the Qwen comparison repeats every condition.

Suite-matched sensitivity: the full-brief side shifted, the frozen rules still passed

Two new full-brief Skilled runs and the estimator were registered before either outcome, then the Qwen comparison was recalculated with two suite-2.5 runs in every arm. Only the Skilled full-brief arm changed, so the public-context difference (+13.17) and the reflective information contrast carry over unchanged by construction. The full-brief difference rose to +9.93 ([2.65, 13.41]); the interaction fell from +5.90 to +3.23, and its interval ([-1.13, 11.93]) now includes zero — the frozen -4-point rule still passes, but the interaction’s direction is no longer interval-resolved. The skilled public-minus-full contrast moved from +0.01 to -2.66 ([-6.40, 2.04]). This is a post-outcome sensitivity, not a replication, and the original mixed-suite analysis remains primary. Download the aggregate-only sensitivity artifact .

Prospectively frozen study · snapshot 2026-07-24

The observed pattern varied across tested backbones.

MediationBench crosses a skilled policy with a minimal reflective-listening control under two information conditions: a full bilateral brief of researcher-authored private facts, or public context only. A backbone is the mediator model running both policies; the disputants and judge are held fixed.

Factorial readout

Model-matched information-study estimands

E1 compares the two policies under public context only; E2 compares them with the full brief; E4 compares the two information conditions within one policy. E3 = E1 − E2 = E4(skilled) − E4(reflective), up to rounding at two decimals.

Qwen is the prospectively frozen replicated core with two runs per arm. Kimi and gpt-oss are conditional checks with one new run per condition. The paper carries this table as Table 7, and these are the primary mixed-suite estimates. The suite-matched sensitivity above moves the Qwen E2, E3, and skilled-E4 values shown here; E1 and the reflective E4 are unchanged by construction, because no new data entered them.

E1 · Public-context skillQwen3-235BKimi K2.5gpt-oss-120bE2 · Full-brief skillQwen3-235BKimi K2.5gpt-oss-120bE3 · Interaction (rule: lower bound > -4)Qwen3-235BKimi K2.5gpt-oss-120bE4 · Public minus full, reflectiveQwen3-235BKimi K2.5gpt-oss-120bE4 · Public minus full, skilledQwen3-235BKimi K2.5gpt-oss-120b-20-10-401020
Points are composite-point differences; whiskers are 90% scenario-reweighting intervals. Filled dots: the replicated Qwen core (two runs per arm). Open dots: one run per condition. End caps mark fully contemporaneous contrasts; capless rows reuse a historical, protocol-matched full-brief skilled run. Dashed rule: the frozen -4-point E3 margin. Exact values are in the table below.
BackboneEvidencePublic-context skill (E1)Full-brief skill (E2)Interaction (E3)Public minus full (E4)
Qwen3-235BTwo runs per condition — the replicated frozen core Reps skilled-full / skilled-public / reflective-full / reflective-public: 2 / 2 / 2 / 2+13.17 [9.67, 17.55]+7.27 [1.31, 10.52]+5.90 [1.34, 13.81] Met the prospectively frozen four-point non-inferiority ruleReflective: -5.89 [-12.89, -2.52] Skilled: +0.01 [-4.15, 4.28]
Kimi K2.5One new run per condition; conditional Reps skilled-full / skilled-public / reflective-full / reflective-public: 2 / 1 / 1 / 1+9.66 [2.32, 16.21]+1.68 [-2.06, 5.85]+7.98 [-0.32, 15.47] Met the inherited frozen four-point rule; exploratoryReflective: -6.45 [-11.99, 0.60] Skilled: +1.53 [-1.64, 5.75]
gpt-oss-120bOne run per condition; conditional Reps skilled-full / skilled-public / reflective-full / reflective-public: 1 / 1 / 1 / 1-1.02 [-4.51, 5.01]+4.74 [2.19, 15.55]-5.75 [-16.75, 0.38] Did not meet the frozen non-inferiority ruleReflective: +4.12 [0.23, 15.22] Skilled: -1.64 [-4.59, 3.38]

Participant-axis transport check · snapshot 2026-07-26

Participant transport: the same 2×2 on Gemma participants

The frozen 2×2 was replayed with gemma-4-31b participants (cerebras) against the Qwen/Qwen3-235B-A22B-Instruct-2507 mediator backbone. Axis varied: participant configuration — mediator backbone, both policy prompts, information conditions, judge, scenario set, and suite held fixed. Policy x information 2x2 plus a scenario-matched no-mediator baseline; one run per arm (1/1/1/1); exploratory post-launch, pre-outcome amendment frozen 2026-07-25T04:19Z before any arm had evaluated scenarios.

Composite scores: no-mediator baseline 15; reflective 34 / 35; skilled 41 / 38 (full brief / public context). All four mediated arms scored above the baseline with 90% intervals excluding zero; per-run identifiers and baseline deltas are in the aggregate download.

Gemma participant-transport E1 to E4 contrasts with 90 percent intervals
Contrast (this participant column)Estimate · 90% interval
E1 · Public-context skill+5.70 [-3.79, 12.90] Failed the prospectively frozen positivity rule
E2 · Full-brief skill+11.39 [5.77, 16.92]
E3 · Interaction (rule: lower bound > -4)-5.68 [-15.95, 2.30] Failed the prospectively frozen four-point non-inferiority rule
E4 · Public minus full, reflective+0.73 [-3.66, 6.34]
E4 · Public minus full, skilled-4.95 [-13.69, 1.94]

Single run per arm, exploratory. Both prospectively frozen reads failed on this column: the public-context skill increment did not separate, and the interaction failed the -4-point non-inferiority margin. This is a participant-axis transport check, never a fourth row of the E1-E4 table above; the paper carries it as Table 8. First failed frozen-rule outcome on the participant axis (gpt-oss failed both rules on the mediator-backbone axis). Mediation still cleared the no-mediator baseline in all four arms. One run per arm; this does not establish that the private brief is necessary. This is a participant column, not a mediator-backbone row.

These are synthetic-dispute results, not evidence of effectiveness with human parties. The whiskers reweight the 62 observed scenarios; they do not estimate variability from fresh generations, different judges, provider or model updates, or human disputes. Only Qwen repeats every condition. Fine model rankings remain judge- and run-sensitive; the conditionality and historical-reuse boundaries are stated once in What is not established.

Analysis provenance
Design registry
2.5.0
Suite provenance
The v2.5 information-study registry combines contemporaneous v2.5 cells with protocol-matched historical full-brief skilled arms that retain their original suite stamps.
Snapshot
2026-07-24
Joint scenarios
62
Scale
62 scenarios × 17 study runs ≈ 1,050 scored conversations
Estimand and intervals
Each point estimate is a difference between arm medians of per-scenario replicate means across 62 scenarios jointly completed in all four arms. Each 90% percentile interval comes from the four-arm scenario-co-resampled bootstrap (B=10000; seed=20260706): scenarios are resampled once per draw across all arms, each arm's observed per-scenario replicate mean is held fixed, and every draw recomputes the point-estimate statistic. An earlier implementation independently resampled replicate identity within scenarios, creating hybrid pseudo-runs; this correction changes intervals only. All intervals are conditional on the observed generation runs and operational judge; they are not run-sampling or human-efficacy intervals.
Participant model
Sao10K/L3-8B-Stheno-v3.2
Operational judge
Qwen/Qwen3-235B-A22B-Instruct-2507
Methodology
seeded_25_v1
Sampling
unseeded
Temperatures
judge 0.3 / mediator 0.7 / participant 0.8
Decision rules
Two decision rules were frozen 2026-07-15, before any public-context run existed: E1 positivity and E3 non-inferiority. E1 passes when its 90% interval lower bound is greater than 0 composite points. E3 passes when its 90% interval lower bound is greater than -4.0 composite points. This is a non-inferiority rule, not an equivalence test; an interval spanning zero does not establish equivalence.
Source commit
abaff7086941f2d4a0da129e38b7019d88933db3 (internal HAI repository, not publicly resolvable)
Analysis artifact
docs/research/analysis/compute_claims_2026-07-28.json (internal HAI repository, not publicly resolvable; all 17 study run IDs, the five participant-transport run IDs, and the GPT-5.5 extension-run ID resolve in the public export, alongside the cross-family rejudge series)
Artifact SHA-256
c1c53164a1441186e7369b8a1a3592a3dcf55021510322a68f0ef1d0c7e266b9
Download SHA-256
ba6f96938f767d0609469d551a65cbd9d8c73d73df1e6022c661de4f9d83d995
Generation completion
All six newly frozen generation cells completed 62 of 62 scenarios after repair, with zero remaining failures. Every regeneration was triggered by completion status alone - a failed scenario or a zero-token generation - and never by the score a run produced (attested by the study author, 2026-07-26).

How the study is organized

The frozen table labels these contrasts E1 through E4 and its caption defines each one; a negative E4 means the mediator scored higher with the private briefs supplied. An arm is one policy-by-information cell of the 2 × 2 design; a replicate is one complete run of an arm across the scenario set; the operational judge is the model that scores every conversation.

Extensions and stress tests

  • On gpt-oss, the public-context skilled-minus-reflective interval included zero, and the interaction did not clear the -4-point non-inferiority margin; gpt-oss met neither frozen rule.
  • On the Gemma-4-31B participant configuration, the public-context result did not transport under the frozen rules: the interaction was -5.68. The participant-transport table above carries the five runs, their scores, and every contrast.
  • Two offline judges from other model families re-scored those fixed transcripts and also scored that interaction negative; the limits give the intervals and dependence details.
  • A registered post-freeze GPT-5.5 skilled/full-brief extension scored 50 on the Gemma participant configuration — one run, no matched reflective control, outside the frozen 2 × 2 analysis; a winner’s-curse caveat applies to selected extension rows.

The information experiment

In the frozen study the conditions are named full bilateral brief and public-context only. What each condition supplies to the mediator:

  • Full bilateral brief: the public scenario, shared facts, the dialogue so far, and both parties’ researcher-authored private scenario facts verbatim.
  • Public-context only: the public scenario, shared facts, and the dialogue so far. Any relevant private constraint must be elicited during the exchange.

The mediator’s brief includes the scenario’s authored focus areas in every arm, public-context-only included: of the 100 authored fixtures, 3 contain a focus area that overlaps substantially with an expected revelation, so a few scenarios give the public-context mediator a directional hint rather than an answer key.

Across Qwen, Kimi, and gpt-oss, the skilled policy’s public-minus-full point estimates were near zero (+0.01, +1.53, and -1.64), and every 90% interval included zero. That pattern is descriptive: it does not establish equivalence, no equivalence margin was specified, and the intervals permit nontrivial shifts. It is also suite-dependent for the one replicated backbone — under the suite-matched sensitivity , Qwen’s estimate moves from +0.01 to -2.66.

Methodology

In the E1-E4 table, the listed backbone is the mediator model used by both the skilled and reflective policies; the simulated disputants and operational judge stay fixed, so the contrast is model-matched. The tested backbones are third-party models. Exact model ids, harness version, role temperatures, and the bootstrap specification are in the readout’s provenance panel; sampling is unseeded.

Runs are regenerated only for completion reasons — a scenario that failed, or a generation that returned zero tokens — and never because of the score a run produced, as attested by the study author on 2026-07-26.

How synthetic disputes are generated

Each authored scenario has exactly two parties: public material establishes the dispute and opening exchange, private material gives each party a backstory, position, hidden interests, and a walk-away alternative, and the same fixture goes to every matched arm. The evaluated partition is 62 dyadic English text scenarios drawn from a 100-scenario corpus that also informed benchmark development — not a contamination-free holdout — with a sequestered, unrun reserve of 38. Synthetic control gives the harness known authored private-information assignments and matched comparisons; it does not prove a participant model acted from every assigned fact, or make the generated people representative of humans.

Construct and scoring

MediationBench measures conversation process, not only whether a settlement appears at the end: whether the dispute moves from zero-sum toward cooperative, the movement the philosophy sets out as the construct. Scores are a 0–100 composite, the explorer’s HAI Score; every E1-E4 estimate and the -4-point margin is a difference in points on it. Its expert-asserted composite weights are:

  • Cooperative Dimensions (cooperative movement between the parties), 25%;
  • Resolution Depth (how fully the dispute is resolved), 25%;
  • Hidden Revelations (discovery of private interests), 20%;
  • Commitment Symmetry (reciprocal commitment), 15%;
  • Moderator Quality (mediator process quality), 15%.

A high composite is not necessarily good mediation. This score does not assess voluntary consent, procedural fairness, power imbalance, safe non-agreement, legality, feasibility, or durability; a Hidden Revelations score does not establish that disclosure was safe or authorized.

The protocol has known design effects: mediated conversations may stop earlier than no-mediator ones, a mediator’s messages do not use up participant turns, and Moderator Quality is zero without a mediator and high for almost every mediated run, so a mediated-versus-baseline difference partly reflects that built-in gap. The frozen analysis treats stopping policy and composite construction as part of the measured package.

Two decision rules were frozen 2026-07-15, before any public-context run existed: E1 positivity, whose interval lower bound must exceed zero, and E3 non-inferiority, whose lower bound must exceed -4 points — the skill increment under public context only must not fall more than 4 composite points below the increment measured with the full bilateral brief. An interval crossing zero does not establish equivalence.

What is not established

  • All disputants are language models, not people. Every interval is conditional on the observed generation runs and the operational judge; none samples fresh generations, serving implementations, or human disputes.
  • The Qwen backbone is the checkpoint also used as the operational judge, so the replicated core is self-judged: a constant self-preference offset cancels in the within-backbone contrast, a style-varying one would not, and cross-backbone readings including the gpt-oss non-replication are unprotected. The export flags degraded judge reliability on two of five categories (κ 0.42 and 0.48, n=40), 45% of the composite weight.
  • Two offline non-Qwen scorers of the fixed transcripts also score the Gemma-participant interaction negative (-8.50 and -4.50 versus -5.68 operationally), corroborating its direction, not its magnitude; only the gpt-oss interval excludes zero (crossed-judge readout ). Scale and interval width stay evaluator-dependent, and offline scoring cannot alter trajectories or stop decisions the operational judge already influenced.
  • Qwen is the only replicated information core; Kimi and gpt-oss hold one run per condition and estimate scenario, not run-to-run generation, variation. For every tested backbone, E2, E3, and the skilled-policy information contrast reuse an earlier protocol-matched full-brief skilled run, so the factorial is not wholly contemporaneous; the cleanest contemporaneous readings are the public-context skill increments and the reflective-policy information effects.
  • The estimands and rules were frozen in an internal registry, not a conventional public preregistration, and the -4-point E3 margin rationale used already-observed full-brief results; no primary estimand, margin, or rule changed after the public-context outcomes.
  • No human criterion study establishes that the composite corresponds to professional mediator or participant judgments, so the study supports configuration-specific policy comparisons, not a universal model-family ranking or a claim of safe deployment.
  • HAI.AI designed and operates the benchmark, develops mediation policies and a mediation product, and publishes this site. These results are not an independent third-party evaluation.

For model labs

MediationBench is a versioned benchmark and research artifact. Aggregate results and the versioned E1-E4 JSON are public; the run harness, judge prompts, scenario-level records, transcripts, and held-out scenarios are not, so third parties cannot independently recompute E1-E4 or their intervals. Access for a new evaluation is arranged by email, as a development loop rather than a trophy table:

  • estimate your mediation prompt’s score contrast against our reflective control;
  • measure what a bilateral private-fact brief changes;
  • inspect failures by scenario type, difficulty, and participant resistance;
  • re-score fixed transcripts without regenerating the conversation.

Data-pipeline invariants can be deterministic product release gates. Fixed-transcript LLM scoring is instead a controlled statistical regression that needs frozen tolerances and calibrated false-alarm rates. Fresh end-to-end generated-score movement remains advisory until paired-seed and repeated-run false-alarm behavior is calibrated; it should not autonomously ship or block a product.

Useful feedback includes missing conflict domains, more credible participant manipulations, human-review design, and the metadata needed for independent reproduction. To discuss an evaluation or methodology contribution, contact hello@hai.io .

Publication boundary

The public site publishes aggregate, versioned evidence. It does not publish the held-out set, raw private briefs, unreleased prompts, private transcripts, canary material, or personally identifying information.

Full-file SHA-256 checksums for every circulated dated JSON and PDF are in the public checksum manifest .

Every run published here is HAI’s own: HAI commissioned it, on HAI’s scenarios, under the conditions described above. No outside party has submitted a system to MediationBench. Evaluation of an outside system is arranged directly with HAI by email, and what may be published from such an evaluation is agreed at the time it is arranged rather than by any standing rule on this page.

Terms and data use

Participation is optional. HAI does not sell personal data and does not publish personally identifying information. Use of hai.ai is governed by the HAI.AI Terms of Service .

Corrections and responses

Operators of evaluated systems may request a factual correction or submit a response for publication by contacting hello@hai.io .

For a safe public fixture excerpt, read the sample conversation . To propose a synthetic dispute, see Contribute a Scenario .

Paper

The paper carries the complete claim boundaries, the prospectively frozen estimands, the uncertainty protocol, and the judge-sensitivity analysis. The current revision (10 August 2026) incorporates the July 31 Sonnet evaluator-context readout and the August 3 Qwen suite-matched sensitivity reported above, and adds a related-work discussion of participant alignment tuning as a validity variable (Kim et al., arXiv:2607.28607). It is hosted here as a working draft with dated revisions; numbers may change between revisions until the preprint release is complete, and this page and the versioned aggregate JSON files remain the citable web record.

Paper revision 2026-08-10 (10 August 2026)

Download the current working draft (PDF)

Draft revisions

License

Two licenses apply. The aggregate results — the E1-E4 table, the current versioned aggregate JSON , the Qwen suite-matched sensitivity , the Sonnet evaluator-context readout , and the crossed-judge readout — are licensed under CC BY 4.0 : reuse the numbers anywhere, including commercially, with attribution. The narrative text and the figures on this page remain licensed under CC BY-NC-SA 4.0 .

Neither license covers the hidden evaluation set, private transcripts, raw prompts, or scoring internals.

How to cite: Jonathan Hendler. MediationBench. Human Assisted Intelligence, PBC. Frozen study snapshot 2026-07-24. https://mediationbench.com/leaderboard/