Published by Human Assisted Intelligence, PBC

MediationBench

MediationBench compares AI mediators on the same 62 simulated two-party conflicts. Within each test, every system works with the same model-played parties under the same rules.

Early results are labeled provisional: one complete run, scored by one or two judge models.

How it works

Each run covers 62 authored disputes. Language models play both parties, and the system being tested acts as mediator. The mediator decides when to speak, wait, or end within a fixed 3-to-25-round window.

Sol scores every finished conversation after the run. A row scored by Sol alone is Sol's average across all 62 conflicts. Other judges can be added later and are shown separately under the row; each judge that scores all 62 conflicts is averaged into the row score, so early results stay inexpensive without pretending they are complete.

How to read the results

Each provisional row shows a score from 0 to 100 and how often the final judge found that the conversation ended appropriately. Rows are ordered by score. Do not assume a small score gap is meaningful.

All parties are AI. The benchmark does not measure human safety, fairness, legality, or whether an agreement will last.

Research and data

MediationBench: Auditing AI Mediation in Synthetic Conflict

Jonathan Hendler · Working draft · Paper revision 2026-09-05 (5 September 2026)

The research page contains the exact method, earlier studies, limitations, corrections, and downloadable data. The working paper gives the complete design and evidence.

Numbers may change while the paper remains a working draft.