Leaderboard
AI mediator performance across 62 simulated conflicts
Current leaderboard · updated 2026-09-06
AI mediator scores · 0–100
Provisional · ordered by score
Current test · 62 simulated conflicts
Each result is one complete run, scored across all 62 conflicts by Sol and by any other judge that covers all 62. One or two judges make a result provisional. Hover over a row for its judges and endings, or open it to keep them in view.
1 GLM 5.3 Flash OpenRouter · simple prompt Score 66.0 Judges 2 Ended appropriately 42/62
Provisional · 1 run · scored by Sol and Gemini
- Sol · 62/62Score 61.6
- Gemini · 62/62Score 70.4
Ended appropriately: 42 of 62 · 0 unclear
Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.
More evidence about this run
Conversations lasted 5.5 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.
What Sol found at the end
- 31 supported agreements
- 5 useful resolutions
- 6 reasonable endings without resolution
- 20 endings that came too soon
- 0 unclear endings
The moderator said it had reached agreement in 52 of 62 conversations. Sol agreed in 30 of those cases.
The moderator chose when to stop in 62 conversations; 0 reached the limit instead.
Wording overlap check: 273 of 46250 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.
Full counts and judge recipes
Conversation length
Run 1 averaged 5.5 rounds. Round 3: 10 · Round 4: 13 · Round 5: 16 · Round 6: 12 · Round 7: 3 · Round 9: 3 · Round 10: 3 · Round 12: 1 · Round 14: 1
Moderator ending claims compared with Sol
- Claimed agreement: 30 supported agreements, 4 useful resolutions, 0 reasonable endings without resolution, 18 too soon, 0 unclear.
- Claimed useful resolution: 1 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
- Claimed no resolution: 0 supported agreements, 0 useful resolutions, 6 reasonable endings without resolution, 1 too soon, 0 unclear.
- No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
Judge evidence
- Sol · score 61.6
Evaluator recipe:
320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddbEndings: 31 supported agreements, 5 useful resolutions, 6 reasonable endings without resolution, 20 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 195 extra; recall 4.6%. - Gemini · score 70.4
Evaluator recipe:
7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559Endings: 49 supported agreements, 5 useful resolutions, 6 reasonable endings without resolution, 2 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 6 found, 232 missed, 115 extra; recall 2.5%.
2 Qwen3.8 Flash OpenRouter · simple prompt Score 60.0 Judges 2 Ended appropriately 45/62
Provisional · 1 run · scored by Sol and Gemini
- Sol · 62/62Score 58.2
- Gemini · 62/62Score 61.7
Ended appropriately: 45 of 62 · 0 unclear
Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.
More evidence about this run
Conversations lasted 4.8 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.
What Sol found at the end
- 40 supported agreements
- 1 useful resolutions
- 4 reasonable endings without resolution
- 17 endings that came too soon
- 0 unclear endings
The moderator said it had reached agreement in 42 of 62 conversations. Sol agreed in 37 of those cases.
The moderator chose when to stop in 62 conversations; 0 reached the limit instead.
Wording overlap check: 130 of 32624 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.
Full counts and judge recipes
Conversation length
Run 1 averaged 4.8 rounds. Round 3: 17 · Round 4: 13 · Round 5: 16 · Round 6: 7 · Round 7: 3 · Round 8: 4 · Round 9: 2
Moderator ending claims compared with Sol
- Claimed agreement: 37 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 5 too soon, 0 unclear.
- Claimed useful resolution: 3 supported agreements, 0 useful resolutions, 1 reasonable endings without resolution, 0 too soon, 0 unclear.
- Claimed no resolution: 0 supported agreements, 1 useful resolutions, 3 reasonable endings without resolution, 12 too soon, 0 unclear.
- No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
Judge evidence
- Sol · score 58.2
Evaluator recipe:
320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddbEndings: 40 supported agreements, 1 useful resolutions, 4 reasonable endings without resolution, 17 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 12 found, 226 missed, 192 extra; recall 5.0%. - Gemini · score 61.7
Evaluator recipe:
7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559Endings: 43 supported agreements, 3 useful resolutions, 11 reasonable endings without resolution, 5 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 104 extra; recall 4.6%.
3 Grok 4.6 OpenRouter · simple prompt Score 54.6 Judges 2 Ended appropriately 40/62
Provisional · 1 run · scored by Sol and Gemini
- Sol · 62/62Score 56.1
- Gemini · 62/62Score 53.2
Ended appropriately: 40 of 62 · 0 unclear
Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.
More evidence about this run
Conversations lasted 4.4 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.
What Sol found at the end
- 32 supported agreements
- 0 useful resolutions
- 8 reasonable endings without resolution
- 22 endings that came too soon
- 0 unclear endings
The moderator said it had reached agreement in 35 of 62 conversations. Sol agreed in 31 of those cases.
The moderator chose when to stop in 62 conversations; 0 reached the limit instead.
Wording overlap check: 223 of 18725 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.
Full counts and judge recipes
Conversation length
Run 1 averaged 4.4 rounds. Round 3: 22 · Round 4: 15 · Round 5: 12 · Round 6: 7 · Round 7: 2 · Round 8: 4
Moderator ending claims compared with Sol
- Claimed agreement: 31 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 4 too soon, 0 unclear.
- Claimed useful resolution: 1 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
- Claimed no resolution: 0 supported agreements, 0 useful resolutions, 8 reasonable endings without resolution, 18 too soon, 0 unclear.
- No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
Judge evidence
- Sol · score 56.1
Evaluator recipe:
320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddbEndings: 32 supported agreements, 0 useful resolutions, 8 reasonable endings without resolution, 22 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 195 extra; recall 3.4%. - Gemini · score 53.2
Evaluator recipe:
7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559Endings: 36 supported agreements, 0 useful resolutions, 12 reasonable endings without resolution, 14 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 106 extra; recall 3.4%.
4 DeepSeek V4 Pro DeepInfra · simple prompt Score 46.5 Judges 2 Ended appropriately 9/62
Provisional · 1 run · scored by Sol and Gemini
- Sol · 62/62Score 49.1
- Gemini · 62/62Score 44.0
Ended appropriately: 9 of 62 · 0 unclear
Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.
More evidence about this run
Conversations lasted 3.7 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.
What Sol found at the end
- 6 supported agreements
- 0 useful resolutions
- 3 reasonable endings without resolution
- 53 endings that came too soon
- 0 unclear endings
The moderator said it had reached agreement in 25 of 62 conversations. Sol agreed in 5 of those cases.
The moderator chose when to stop in 62 conversations; 0 reached the limit instead.
Wording overlap check: 259 of 24204 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.
Full counts and judge recipes
Conversation length
Run 1 averaged 3.7 rounds. Round 3: 33 · Round 4: 18 · Round 5: 8 · Round 6: 2 · Round 7: 1
Moderator ending claims compared with Sol
- Claimed agreement: 5 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 20 too soon, 0 unclear.
- Claimed useful resolution: 1 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 4 too soon, 0 unclear.
- Claimed no resolution: 0 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 29 too soon, 0 unclear.
- No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
Judge evidence
- Sol · score 49.1
Evaluator recipe:
320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddbEndings: 6 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 53 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 175 extra; recall 3.4%. - Gemini · score 44.0
Evaluator recipe:
7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559Endings: 9 supported agreements, 2 useful resolutions, 5 reasonable endings without resolution, 46 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 108 extra; recall 3.4%.
5 Gemma 4 31B Cerebras · simple prompt Score 42.1 Judges 2 Ended appropriately 28/62
Provisional · 1 run · scored by Sol and Gemini
- Sol · 62/62Score 43.9
- Gemini · 62/62Score 40.2 Same provider lab as mediator
Ended appropriately: 28 of 62 · 0 unclear
Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.
More evidence about this run
Conversations lasted 17.4 rounds on average. 19 of 62 reached the 25-round limit; 22 ended at round 24 or later.
What Sol found at the end
- 8 supported agreements
- 7 useful resolutions
- 13 reasonable endings without resolution
- 34 endings that came too soon
- 0 unclear endings
The moderator said it had reached agreement in 20 of 62 conversations. Sol agreed in 6 of those cases.
The moderator chose when to stop in 44 conversations; 18 reached the limit instead.
Wording overlap check: 144 of 14521 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.
Full counts and judge recipes
Conversation length
Run 1 averaged 17.4 rounds. Round 3: 1 · Round 4: 1 · Round 5: 1 · Round 8: 1 · Round 9: 4 · Round 10: 4 · Round 12: 4 · Round 13: 6 · Round 14: 4 · Round 15: 1 · Round 16: 4 · Round 17: 2 · Round 19: 2 · Round 20: 2 · Round 21: 3 · Round 24: 3 · Round 25: 19
Moderator ending claims compared with Sol
- Claimed agreement: 6 supported agreements, 5 useful resolutions, 0 reasonable endings without resolution, 9 too soon, 0 unclear.
- Claimed useful resolution: 1 supported agreements, 2 useful resolutions, 0 reasonable endings without resolution, 3 too soon, 0 unclear.
- Claimed no resolution: 0 supported agreements, 0 useful resolutions, 8 reasonable endings without resolution, 10 too soon, 0 unclear.
- No ending claim: 1 supported agreements, 0 useful resolutions, 5 reasonable endings without resolution, 12 too soon, 0 unclear.
Judge evidence
- Sol · score 43.9
Evaluator recipe:
320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddbEndings: 8 supported agreements, 7 useful resolutions, 13 reasonable endings without resolution, 34 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 20 found, 218 missed, 196 extra; recall 8.4%. - Gemini · score 40.2 · same provider lab as mediator
Evaluator recipe:
7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559Endings: 12 supported agreements, 3 useful resolutions, 13 reasonable endings without resolution, 34 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 17 found, 221 missed, 107 extra; recall 7.1%.
Nemotron 3 Ultra Together · simple prompt Incomplete · Sol scored 0/62
- Sol · 0/62
Reference scores · not ranked 2
What the same kind of conversation reaches with no mediator at all. Read it as context for the results above: it is never ranked, and no difference is computed.
- No mediator · Gemma parties July 2026 historical reference 15
Earlier result · kept from the original publication
- No mediator · Stheno parties July 2026 reference 17
Earlier result
Earlier results · not compared 8
These results used earlier test or scoring rules. They stay public as separate records.
- GLM 5.3 Flash August 2026 method 43.6
Earlier result
- GLM-5.2 July 2026 method 52
Earlier result
- gpt-oss-120b July 2026 method 49
Earlier result
- Kimi K2.5 July 2026 method 45
Earlier result
- Qwen3-235B July 2026 method · Stheno parties 46
Earlier result
- Qwen3-235B · Gemma parties July 2026 participant check 39.5
Earlier result
- Reflective control July 2026 reference 40.5
Earlier result
- Sonnet 5 July 2026 method 51
Earlier result
Each score is the equal-weight average of every judge that scored all 62 simulated conflicts in one complete run; the judges counted are drawn as colored rows in the bar and named in the row's evidence. One or two judges make a result provisional. Rows are ordered by score; a small gap between rows is not evidence that one system is better. Reference scores are never ranked; the no-mediator run was scored under earlier rules, so read the gap as context rather than a measured difference. Earlier results used different rules and are shown only as separate records. Read the full scoring method and limitations.
These results describe model simulations, not performance in real disputes between people. The research page gives the complete method and limitations.
How it works
- Same test: every system faces the same 62 authored conflicts, and language models play both parties.
- Unguided ending: the mediator decides when to speak, wait, or end inside a fixed 3-to-25 round window it is not told about.
- Score: afterwards, Sol scores each finished conversation from 0 to 100. Other judges can score the same conversations; the row score is the equal-weight average of every judge that scored all 62.
- Ended appropriately: Sol also says whether the ending made sense. It does not mean every conflict was solved.
- Provisional: one run scored by one or two judges. Rows are ordered by score; a small gap between rows is not evidence that one system is better.
- Other judges: opened from the row, each with its own score. Three or more judges make a result no longer provisional.
A sample of the benchmark data is published on Hugging Face: mediationbench-sample .
Earlier results, made under different rules, are on the research page and are not compared with the current leaderboard.
Run a system through MediationBench via hai.ai ; to arrange an evaluation, email hello@hai.io . This site publishes results and does not accept evaluation-run submissions directly.
Frequently asked questions
- What is MediationBench?
- MediationBench compares AI mediators on the same 62 simulated conflicts. Each system works with the same model-played parties under the same test rules. Sol scores the completed conversations; extra judge models can be added.
- How certain are these results?
- A provisional result is one complete run scored by one or two judge models. It is useful early evidence. Rows are ordered by score, but a small gap is not a claim that one system is better than another.
- Does MediationBench prove that AI mediation works with people?
- No. Every dispute in this study is synthetic and every party is played by a language model. The results test controlled model behavior, not safety or effectiveness with human parties.