MediationBench is a synthetic-conflict testbed for building and evaluating AI mediators. It generates controlled two-party disputes with public positions, private interests, emotional state, and concrete constraints; then it records how policy, information, participant configuration, and evaluator context alter the conversation and its measured score.

The aim is not another model leaderboard. The aim is to separate four questions that ordinary end-state scoring mixes together:

  1. Did adding any mediator change the conversation?
  2. Did a skilled policy add value beyond minimal reflective listening?
  3. Did supplying both parties’ private-fact brief change what the mediator could accomplish?
  4. Would the conclusion survive a different run or evaluator?

Prospectively frozen study ยท snapshot 2026-07-24

The observed pattern varied across tested backbones.

MediationBench crosses a skilled policy with a minimal reflective-listening control under two information conditions. Its full bilateral brief is a controlled analogue of preparation, not an observed human interview. The replicated core asks whether skill still adds value when relevant private constraints must be elicited from the conversation.

Policy Information
Full bilateral brief Researcher-authored private facts supplied
Public-context only Public scenario, shared facts, and dialogue
Reflective control Acknowledge and reflect
Basic facilitation with the bilateral brief
Basic facilitation using public context
Skilled policy Probe interests and tradeoffs
Skilled mediation with the bilateral brief
Skilled mediation using public context

Factorial readout

Skill above reflection, with and without the bilateral brief

Each point estimate is a difference between arm medians of per-scenario replicate means across 62 scenarios jointly completed in all four arms. Each 90% percentile interval comes from the frozen four-arm scenario-co-resampled bootstrap (B=10000; seed=20260706), drawing one observed replicate per (scenario, arm) within each draw. All intervals are conditional on the observed generation runs and operational judge; they are not run-sampling or human-efficacy intervals. Qwen is the prospectively frozen replicated core. Kimi and gpt-oss are conditional breadth checks. Replicate order below is skilled full brief / skilled public context / reflective full brief / reflective public context.

BackboneEvidenceFull-brief skillPublic-context skillInteraction / non-inferiorityPublic minus full
Qwen3-235BReplicated prospectively frozen core Reps skilled-full / skilled-public / reflective-full / reflective-public: 2 / 2 / 2 / 2+7.27 [1.68, 15.07]+13.17 [9.76, 20.18]+5.90 [-0.42, 16.11] Met the prospectively frozen four-point non-inferiority ruleReflective: -5.89 [-17.56, -2.32] Skilled: +0.01 [-6.24, 0.85]
Kimi K2.5Conditional single-run breadth Reps skilled-full / skilled-public / reflective-full / reflective-public: 2 / 1 / 1 / 1+1.68 [-0.45, 7.56]+9.66 [2.20, 16.39]+7.98 [-1.74, 13.65] Met the inherited frozen four-point rule; exploratoryReflective: -6.45 [-12.08, 0.74] Skilled: +1.53 [-3.00, 3.28]
gpt-oss-120bConditional single-run breadth Reps skilled-full / skilled-public / reflective-full / reflective-public: 1 / 1 / 1 / 1+4.74 [2.15, 15.53]-1.02 [-4.51, 5.01]-5.75 [-16.63, 0.32] Did not meet the frozen non-inferiority ruleReflective: +4.12 [0.33, 15.10] Skilled: -1.64 [-4.57, 3.38]

Replicated core

Skill survived with public context only.

On Qwen, the skilled policy remained above the model-matched reflective control without the bilateral brief. The prospectively frozen interaction non-inferiority rule also passed: its lower 90% bound exceeded -4 points. This was not an equivalence test.

Failed generalization

The observed pattern was not backbone-universal.

Kimi supported the same direction. gpt-oss did not: its public-context reflective control improved enough to erase the measured skill increment.

Why preparation is tested

The brief models one preparation advantage.

Real mediators may conduct intake or caucus before a joint session. Here, researcher-authored private facts provide a controlled analogue; the public-context condition tests whether the policy can elicit what the brief otherwise supplies.

These are synthetic-dispute results, not evidence of effectiveness with human parties. Every interval is conditional on the finite observed generation runs and the operational judge. For every backbone, E2, E3, and the structured-policy information contrast reuse historical, protocol-matched full-brief skilled runs; E1 and the reflective-policy information contrast use contemporaneous v2.5 cells. Fine model rankings remain judge- and run-sensitive.

Analysis provenance
Design registry
2.5.0
Suite provenance
The v2.5 information-study registry combines contemporaneous v2.5 cells with protocol-matched historical full-brief skilled arms that retain their original suite stamps.
Snapshot
2026-07-24
Joint scenarios
62
Participant model
Sao10K/L3-8B-Stheno-v3.2
Operational judge
Qwen/Qwen3-235B-A22B-Instruct-2507
Methodology
seeded_25_v1
Sampling
unseeded
Temperatures
judge 0.3 / mediator 0.7 / participant 0.8
Decision rule
E3 passes when its 90% interval lower bound is greater than -4.0 composite points. This is a non-inferiority rule, not an equivalence test; an interval spanning zero does not establish equivalence.
Source commit
27c14ebcac7b49dcd3da569bce5cc52916f7fdf3
Analysis artifact
docs/research/analysis/compute_claims_2026-07-24.json
Artifact SHA-256
d9a6c84f6c6bb448a348bc1fd02c8ed9e68490bf3320b9cdcbea9001d0ddbeff
Breadth completion
All six newly frozen breadth cells completed 62 of 62 scenarios after repair, with zero remaining failures.

What the study says now

The replicated Qwen core supports a bounded within-harness contrast: under the high-resistance participant configuration, a skilled mediation policy improved the synthetic-study composite score relative to a model-matched reflective-listening control when both received only public context.

That result did not become a universal model claim. Conditional breadth supported the same direction for Kimi, while gpt-oss did not separate skilled mediation from reflection in the public-context-only condition. Across the tested configurations, the observed pattern was not backbone-universal: the active control was not inert, and policy-by-information estimates varied by backbone.

The participant result is similarly bounded. The benchmark discriminated among mediators under one high-resistance configuration but compressed them into a narrow band under one cooperative assistant-aligned configuration. Capability, disposition, and alignment differ between those participant models, so MediationBench treats the simulated disputant as an experimental factor, not as a stand-in for a population of people.

The information experiment

The run metadata calls the conditions full and blind. In the study, the literal names are full bilateral brief and public-context only:

  • Full bilateral brief: the mediator receives both parties' researcher-authored private scenario facts verbatim alongside public context.
  • Public-context only: the mediator receives the public scenario, shared facts, and dialogue. It must elicit any relevant private constraint during the exchange.

Neither condition is the universally correct way to mediate. Their contrast asks what supplying the bilateral brief changes, whether skilled elicitation can recover information absent from that brief, and whether a result depends on information an ordinary observer would not possess. The full brief is a controlled analogue of one advantage that intake, interview, or caucus may provide; the experiment did not observe any human preparation process.

In the replicated core, the skilled policy’s point score changed little between the two conditions, while the reflective control was lower without its brief. The breadth backbones did not reproduce that exact pattern, so we report policy-by-backbone heterogeneity rather than a law that bilateral briefs always help or that skill always substitutes for them.

Methodology

The frozen information study uses Sao10K/L3-8B-Stheno-v3.2 for both simulated disputants and Qwen/Qwen3-235B-A22B-Instruct-2507 as its operational judge. The harness methodology is seeded_25_v1: fixture dialogue is fixed, but language-model sampling is unseeded. Temperatures are 0.3 for judging, 0.7 for mediation, and 0.8 for participant turns. These values describe the tested configuration; they do not make stochastic generation deterministic.

How synthetic disputes are generated

Each authored fixture contains exactly two parties. Public material establishes the dispute and opening exchange. Private material gives each simulated party a backstory, stated position, hidden interests, emotions, and a walk-away alternative. Scenario metadata identifies the conflict domain, difficulty, topics expected to surface, and observable evidence of progress.

The same fixture is supplied to every matched arm. Participant models then continue the dispute in character. This produces comparable simulated trajectories from a shared starting point: no mediator, reflective control, and skilled mediation; full-brief and public-context mediation; different participant and evaluator configurations.

Synthetic control is the advantage and the limitation. It makes hidden state, matched comparisons, and repeatable stress tests possible. It does not make the generated people representative of humans. MediationBench is therefore a development and evaluation surface for mediation models, not a clinical, legal, or deployment trial.

Construct and scoring

MediationBench measures conversation process, not only whether a settlement appears at the end. Its public composite includes:

  • cooperative movement;
  • resolution depth;
  • discovery of private interests;
  • reciprocal commitment;
  • mediator process quality.

The public categories and aggregate scoring structure remain inspectable. Raw judge prompts, held-out fixtures, private transcripts, and anti-gaming mechanics remain private so the test set cannot become a training answer key.

The current protocol also has known design effects. Mediated runs may stop earlier than no-mediator runs, mediator messages do not consume participant turns, and the no-mediator row has no mediator-quality component. The paper treats stopping policy and the composite construction as part of the measured package and calls for common-horizon and human-calibrated follow-up analyses.

The E1-E4 table reports differences between arm medians of per-scenario replicate means over the 62 scenarios completed in all four arms. Its 90% intervals use one frozen four-arm, scenario-co-resampled bootstrap (B=10000, seed 20260706), with one observed replicate selected per (scenario, arm) in each draw. The intervals condition on the observed generation runs and operational judge; they do not sample fresh generations, serving implementations, or human disputes. The E3 rule is non-inferiority: its lower interval bound must exceed -4 composite points. An interval crossing zero does not establish equivalence.

MediationBench is complementary to three close studies:

  • Robots in the Middle evaluates the choice and wording of a single intervention on manually authored short disputes.
  • ProMediate evaluates when and how proactive mediators intervene in structured multi-party negotiation under a shared conversational budget.
  • SoCRATES expands proactive mediation evaluation across domains and socio-cognitive conditions with matched unmediated runs and a human-validated topic evaluator.

MediationBench focuses on a different measurement question: how participant configuration, a model-matched reflective control, intake information, run variation, and evaluator context change the estimated benefit of a mediation policy.

What is not established

  • All disputants are language models, not people.
  • The cooperative and high-resistance participant configurations differ in more than disposition; a same-backbone persona experiment is still needed.
  • No human criterion study yet establishes that the composite corresponds to professional mediator or participant judgments.
  • The replicated information core is Qwen. Kimi and gpt-oss breadth cells contain one new run in each new breadth arm and estimate scenario variation conditional on the finite observed runs and the operational judge, not run-to-run generation variation.
  • For every tested backbone, E2, E3, and the structured-policy information contrast reuse at least one earlier, protocol-matched full-brief skilled run. The factorial is therefore not wholly contemporaneous. The cleanest contemporaneous readings are the public-context skill increments and the reflective-policy information effects.
  • Passing E3 means that its lower 90% interval bound exceeded the frozen -4-point non-inferiority margin. It does not establish equivalence; neither does an interval that crosses zero.
  • The study supports configuration-specific policy comparisons, not a universal ranking of model families or a claim of safe deployment.
  • HAI.AI designed and operates the benchmark, develops mediation policies and a mediation product, and publishes this site. These results are not an independent third-party evaluation.

For model labs

MediationBench can be used as a development loop rather than a trophy table:

  • test whether a mediation prompt adds value beyond reflection;
  • measure what a bilateral private-fact brief changes;
  • inspect failures by scenario type and difficulty;
  • compare intervention policies under cooperative and resistant simulations;
  • generate controlled trajectories for evaluator and mediator development;
  • re-score fixed transcripts without regenerating the conversation.

Useful feedback includes missing conflict domains, more credible participant manipulations, whole-run replication requirements, human-review design, and the configuration metadata needed for independent reproduction. To discuss an evaluation or methodology contribution, contact hello@hai.io .

Integrity and data use

The public site publishes aggregate, versioned evidence. It does not publish the held-out set, raw private briefs, unreleased prompts, private transcripts, canary material, or personally identifying information. HAI’s separate Benchmark Terms permit uses of benchmark submissions including AI-model training, benchmark development, research/publication, and product, feature, and service development. Individual benchmark details are not made public without consent, and HAI does not sell personal data.

Participation is optional. The HAI.AI Terms of Service apply generally; use of the benchmark platform is additionally governed by the Benchmark Terms . Operators of evaluated systems may request a factual correction or submit a response for publication by contacting hello@hai.io .

For a safe public fixture excerpt, read the sample conversation . To propose a synthetic dispute, see Contribute a Scenario .

Public run explorer

The explorer is the currently published aggregate run snapshot, separate from the frozen E1-E4 study readout above. It can show versioned no-mediator, reflective-control, and skilled-mediator generation rows, plus fixed-transcript re-scores when the public snapshot contains them. A snapshot may lag a newly completed study cell until the publication step creates the next version.

The publication rule is outcome-independent: every registered clean generation arm included in the frozen study is public, including public-context-only (blind) mediator arms and null or negative outcomes. The operational Qwen judge is the primary study instrument. Preplanned robustness measurements are presented separately from clearly labeled exploratory diagnostics. A rejudge re-scores an existing transcript and is never counted as a new generation run or generation replicate.

Public snapshot

Loading results…

The latest aggregate HAI Score snapshot is loading.

How to read the run browser

  • Version first. Compare runs inside the same suite, fixture version, and information condition before comparing across time.
  • The active control is primary. Skilled minus reflective control asks what mediation strategy adds beyond acknowledgment and reflection. Reflective control minus no mediator asks what basic facilitation contributes.
  • Baseline delta is practical, not statistical. It is a descriptive difference from a matching no-mediator row. Baseline has no mediator-quality score, so it is not the cleanest endpoint for comparing policy skill.
  • Treat small model gaps as ties. End-to-end re-runs and alternative evaluators changed some individual scores and orderings; the browser is not a stable model ranking.
  • Judge context is part of the instrument. The blind-evaluator pass changed individual scores and reduced rank agreement relative to sighted scoring. It is a fixed-transcript sensitivity analysis, not a preregistered E1-E4 estimand; no claim is made here about its effect on arm contrasts.

Every score is date-stamped and belongs to a named configuration. It is not a certification, safety guarantee, or permanent judgment of a named model. The versioned E1-E4 aggregate JSON is downloadable separately from the rolling explorer snapshot.

Paper

The paper gives the complete claim boundaries, prospectively frozen estimands, uncertainty protocol, and judge-sensitivity analysis.

The whitepaper is in preparation. It will be published here as a downloadable PDF.

License

The aggregate public data report is licensed under CC BY-NC-SA 4.0 . The license does not cover the hidden evaluation set, private transcripts, raw prompts, or scoring internals.


Run a system through MediationBench via hai.ai . This site publishes results and does not accept evaluation-run submissions directly.

Frequently asked questions

What is MediationBench?
MediationBench is a controlled synthetic-conflict testbed for evaluating AI mediation policies. It compares the same authored fixture and starting transcript across no-mediator, reflective-control, and skilled-mediator conditions while recording the participant, mediator, information, and judge configuration.
Why compare two information conditions?
The full-bilateral-brief condition supplies both parties' researcher-authored private scenario facts verbatim. The public-context-only condition supplies the public scenario, shared facts, and dialogue. This is a controlled analogue of one advantage that intake or caucus may provide; it is not an observed human interview.
Are these statistically significant results?
The study reports prospectively frozen paired contrasts with 90 percent scenario-reweighting intervals. The Qwen information study is the replicated frozen core; Kimi and gpt-oss are conditional breadth checks. The public run explorer remains descriptive, and no human-outcome validation is claimed.
Why test more than one participant configuration?
The simulated disputant is part of the measurement system. Assistant-aligned participants may resolve a conflict on their own, while a high-resistance configuration leaves more room for mediation. The current configurations also differ in capability and training, so the study treats the result as configuration sensitivity rather than proof that one participant model is more realistic.
Does MediationBench prove that AI mediation works with people?
No. Every dispute in this study is synthetic and every party is played by a language model. The results test controlled model behavior, not safety or effectiveness with human parties.