Human Assisted Intelligence, PBC · MediationBench study 01 · frozen study snapshot 2026-07-24

MediationBench

A synthetic-conflict testbed for building and evaluating AI mediators. Every party is played by a language model, and every run replays the same authored two-party disputes and opening exchanges. Only two things change: the mediator's instructions, and what it knows going in.

Read the working-draft paper (PDF)

Why this study exists

How do mediation policies compare in a controlled simulation?

The answer depends on choices that usually stay hidden. Six of them can change it outright:

  • who plays the two disputants,
  • what the mediator knows before it starts,
  • what the mediator is compared against,
  • when the exchange stops,
  • which AI model scores the conversation,
  • and what that judge is shown.

MediationBench makes those choices explicit and holds them fixed. It runs 62 authored two-party disputes with both sides played by language models. Two mediation policies — a reflective-listening control and a skilled policy — run as different instructions on the same underlying model, and a fixed judge model scores each conversation from 0 to 100.

For mediator and foundation-model builders

Use synthetic conflict as a development loop.

These runs are lab equipment, not evidence about people. Synthetic conversations and synthetic mediation are experimental instruments and product test harnesses: they enable controlled comparisons and repeatable failure analysis, but do not substitute for validation with people in real conflicts.

Use it to compare policies and information conditions on matched runs, to regression-test a judge against fixed transcripts, to check data pipelines, and to discover failures worth turning into hypotheses.

What the experiment found

Three findings, each with its limit.

01

The public-context result did not transport to a second participant configuration.

Change who plays the disputants, and the pattern changes. Under the frozen rules, the public-context Skilled advantage was 5.68 points smaller than the full-brief advantage with Gemma participants. Two offline judge families re-scoring the same transcripts also put that interaction below zero (-8.50 and -4.50), though magnitudes differed.

02

Reflection is the active control, so policy is tested above it.

The bar is a competent baseline, not silence. The benchmark compares skilled-policy scores with a reflective mediator running on the same model. That within-harness contrast isolates instructions more cleanly than a no-mediator comparison; it does not establish benefit in human disputes.

03

A suite-matched recheck shifted the interaction, within the frozen rules.

The follow-up was chosen after the original result; its two new runs and estimator were registered before those new outcomes. The additions gave every Qwen arm two suite-2.5 runs but changed only the Skilled full-brief arm, so the public-context difference of +13.17 points (90% interval [9.67, 17.55]) carries over unchanged by construction — those arms kept their original runs, so that number could not move. The full-brief difference rose to +9.93, and the interaction fell to +3.23 — its interval now includes zero while clearing the frozen -4-point margin. It is a sensitivity, not a replication, and the effect is still not backbone-universal: gpt-oss met neither frozen rule.

Full bilateral brief

Private facts handed over up front.

Both parties' researcher-authored private facts are provided verbatim alongside public context.

Public-context only

Private facts must be drawn out.

The mediator receives the public scenario, shared facts, and dialogue, but no private-fact brief.

The preparation experiment

How do scores change when both sides' private facts are supplied?

One condition supplies both parties' authored private facts; the other supplies only public context and the opening dialogue. This is an information-provisioning comparison inside the simulation. It does not model intake, confidentiality, consent, or permission to disclose.

See the prospectively frozen 2 × 2 study

Safe experimentation

Models, not people.

Every party in this study is simulated. MediationBench measures controlled model behavior; it does not establish safety, legal validity, or effectiveness with human participants.

MediationBench is a benchmark and research program owned and operated by Human Assisted Intelligence, PBC (HAI.AI). HAI.AI designed and operates MediationBench, develops mediation policies and a mediation product, and publishes this site. These results are not an independent third-party evaluation. Not affiliated with Stanford HAI.