AI mediator performance across 62 simulated conflicts

Current leaderboard · updated 2026-09-14

AI mediator scores · 0–100

Provisional · ordered by score

Current test · 62 simulated conflicts

Each result is one complete run, scored across all 62 conflicts by Sol and by any other judge that covers all 62. One or two judges make a result provisional. Hover over a row for its judges and endings, or open it to keep them in view.

Version 3.0

  1. 1 Claude Opus 5 · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 simple prompt Score 71.2 Judges 2 Ended appropriately 55/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 67.5
    • Gemini · 62/62Score 74.9
    Ended appropriately: 55 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.6 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 50 supported agreements
    • 3 useful resolutions
    • 2 reasonable endings without resolution
    • 7 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 58 of 62 conversations. Sol agreed in 50 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 228 of 55092 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.6 rounds. Round 3: 17 · Round 4: 17 · Round 5: 15 · Round 6: 5 · Round 7: 4 · Round 8: 3 · Round 10: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 50 supported agreements, 3 useful resolutions, 0 reasonable endings without resolution, 5 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 2 reasonable endings without resolution, 2 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 67.5 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 50 supported agreements, 3 useful resolutions, 2 reasonable endings without resolution, 7 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 9 found, 229 missed, 179 extra; recall 3.8%.
    • Gemini · score 74.9 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 56 supported agreements, 2 useful resolutions, 4 reasonable endings without resolution, 0 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 14 found, 224 missed, 100 extra; recall 5.9%.
  2. 2 Claude Fable 5.1 · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 xhigh · simple prompt Score 69.9 Judges 2 Ended appropriately 51/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 67.3
    • Gemini · 62/62Score 72.4
    Ended appropriately: 51 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.3 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 46 supported agreements
    • 2 useful resolutions
    • 3 reasonable endings without resolution
    • 11 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 55 of 62 conversations. Sol agreed in 46 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 265 of 43787 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.3 rounds. Round 3: 16 · Round 4: 25 · Round 5: 12 · Round 6: 7 · Round 7: 2

    Moderator ending claims compared with Sol
    • Claimed agreement: 46 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 8 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 3 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 67.3 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 46 supported agreements, 2 useful resolutions, 3 reasonable endings without resolution, 11 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 194 extra; recall 4.6%.
    • Gemini · score 72.4 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 53 supported agreements, 2 useful resolutions, 5 reasonable endings without resolution, 2 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 12 found, 226 missed, 92 extra; recall 5.0%.
  3. 3 Kimi K3 · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 OpenRouter · simple prompt Score 68.7 Judges 2 Ended appropriately 48/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 64.8
    • Gemini · 62/62Score 72.7
    Ended appropriately: 48 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 5.0 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 45 supported agreements
    • 1 useful resolutions
    • 2 reasonable endings without resolution
    • 14 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 56 of 62 conversations. Sol agreed in 44 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 250 of 46364 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 5.0 rounds. Round 3: 16 · Round 4: 9 · Round 5: 18 · Round 6: 7 · Round 7: 7 · Round 8: 2 · Round 9: 3

    Moderator ending claims compared with Sol
    • Claimed agreement: 44 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 12 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 2 reasonable endings without resolution, 2 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 64.8 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 45 supported agreements, 1 useful resolutions, 2 reasonable endings without resolution, 14 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 7 found, 231 missed, 200 extra; recall 2.9%.
    • Gemini · score 72.7 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 53 supported agreements, 4 useful resolutions, 4 reasonable endings without resolution, 1 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 7 found, 231 missed, 111 extra; recall 2.9%.
  4. 4 GLM 5.3 · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 DeepInfra, FP4 · simple prompt Score 67.6 Judges 2 Ended appropriately 36/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 62.7
    • Gemini · 62/62Score 72.5
    Ended appropriately: 36 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 5.4 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 30 supported agreements
    • 6 useful resolutions
    • 0 reasonable endings without resolution
    • 26 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 60 of 62 conversations. Sol agreed in 30 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 210 of 37893 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 5.4 rounds. Round 3: 11 · Round 4: 13 · Round 5: 16 · Round 6: 5 · Round 7: 6 · Round 8: 7 · Round 9: 1 · Round 10: 1 · Round 11: 1 · Round 13: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 30 supported agreements, 6 useful resolutions, 0 reasonable endings without resolution, 24 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 62.7 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 30 supported agreements, 6 useful resolutions, 0 reasonable endings without resolution, 26 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 15 found, 223 missed, 186 extra; recall 6.3%.
    • Gemini · score 72.5 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 55 supported agreements, 5 useful resolutions, 0 reasonable endings without resolution, 2 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 12 found, 226 missed, 101 extra; recall 5.0%.
  5. 5 GPT-5.6 Sol · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 Azure · high effort · simple prompt Score 66.9 Judges 2 Ended appropriately 58/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 66.1 Same provider lab as mediator
    • Gemini · 62/62Score 67.6
    Ended appropriately: 58 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.5 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 55 supported agreements
    • 0 useful resolutions
    • 3 reasonable endings without resolution
    • 4 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 56 of 62 conversations. Sol agreed in 54 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 50 of 27105 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.5 rounds. Round 3: 18 · Round 4: 17 · Round 5: 14 · Round 6: 7 · Round 7: 5 · Round 9: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 54 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 2 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 1 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 2 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 66.1 · same provider lab as mediator Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 55 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 4 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 194 extra; recall 4.6%.
    • Gemini · score 67.6 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 50 supported agreements, 5 useful resolutions, 5 reasonable endings without resolution, 2 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 101 extra; recall 3.4%.
  6. 6 Muse Spark 1.3 Benchmark version 3.0 Meta · simple prompt Score 66.8 Judges 2 Ended appropriately 59/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 63.3
    • Gemini · 62/62Score 70.3
    Ended appropriately: 59 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 8.4 rounds on average. 1 of 62 reached the 25-round limit; 1 ended at round 24 or later.

    What Sol found at the end

    • 55 supported agreements
    • 2 useful resolutions
    • 2 reasonable endings without resolution
    • 3 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 57 of 62 conversations. Sol agreed in 54 of those cases.

    The moderator chose when to stop in 61 conversations; 1 reached the limit instead.

    Wording overlap check: 153 of 34802 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 8.4 rounds. Round 4: 6 · Round 5: 8 · Round 6: 9 · Round 7: 10 · Round 8: 6 · Round 9: 6 · Round 10: 3 · Round 11: 3 · Round 12: 2 · Round 13: 1 · Round 14: 3 · Round 15: 3 · Round 19: 1 · Round 25: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 54 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 2 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 2 reasonable endings without resolution, 1 too soon, 0 unclear.
    • No ending claim: 1 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 63.3 Evaluator recipe: 320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddb Endings: 55 supported agreements, 2 useful resolutions, 2 reasonable endings without resolution, 3 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 10 found, 228 missed, 190 extra; recall 4.2%.
    • Gemini · score 70.3 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 54 supported agreements, 5 useful resolutions, 3 reasonable endings without resolution, 0 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 10 found, 228 missed, 105 extra; recall 4.2%.
  7. 7 GLM 5.3 Flash Benchmark version 3.0 OpenRouter · simple prompt Score 66.0 Judges 2 Ended appropriately 42/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 61.6
    • Gemini · 62/62Score 70.4
    Ended appropriately: 42 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 5.5 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 31 supported agreements
    • 5 useful resolutions
    • 6 reasonable endings without resolution
    • 20 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 52 of 62 conversations. Sol agreed in 30 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 273 of 46250 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 5.5 rounds. Round 3: 10 · Round 4: 13 · Round 5: 16 · Round 6: 12 · Round 7: 3 · Round 9: 3 · Round 10: 3 · Round 12: 1 · Round 14: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 30 supported agreements, 4 useful resolutions, 0 reasonable endings without resolution, 18 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 6 reasonable endings without resolution, 1 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 61.6 Evaluator recipe: 320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddb Endings: 31 supported agreements, 5 useful resolutions, 6 reasonable endings without resolution, 20 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 195 extra; recall 4.6%.
    • Gemini · score 70.4 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 49 supported agreements, 5 useful resolutions, 6 reasonable endings without resolution, 2 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 6 found, 232 missed, 115 extra; recall 2.5%.
  8. 8 GPT-5.6 Terra · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 Azure · high effort · simple prompt Score 64.9 Judges 2 Ended appropriately 56/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 62.9 Same provider lab as mediator
    • Gemini · 62/62Score 66.8
    Ended appropriately: 56 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 5.5 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 43 supported agreements
    • 6 useful resolutions
    • 7 reasonable endings without resolution
    • 6 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 48 of 62 conversations. Sol agreed in 42 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 131 of 34651 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 5.5 rounds. Round 3: 11 · Round 4: 14 · Round 5: 18 · Round 6: 7 · Round 7: 5 · Round 8: 1 · Round 9: 2 · Round 10: 1 · Round 11: 1 · Round 15: 1 · Round 20: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 42 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 5 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 5 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 7 reasonable endings without resolution, 1 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 62.9 · same provider lab as mediator Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 43 supported agreements, 6 useful resolutions, 7 reasonable endings without resolution, 6 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 12 found, 226 missed, 199 extra; recall 5.0%.
    • Gemini · score 66.8 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 46 supported agreements, 5 useful resolutions, 10 reasonable endings without resolution, 1 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 9 found, 229 missed, 107 extra; recall 3.8%.
  9. 9 Claude Sonnet 5 · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 xhigh · simple prompt Score 62.1 Judges 2 Ended appropriately 47/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 60.5
    • Gemini · 62/62Score 63.7
    Ended appropriately: 47 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 5.0 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 33 supported agreements
    • 5 useful resolutions
    • 9 reasonable endings without resolution
    • 15 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 48 of 62 conversations. Sol agreed in 33 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 202 of 49858 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 5.0 rounds. Round 3: 16 · Round 4: 17 · Round 5: 13 · Round 6: 6 · Round 7: 1 · Round 8: 4 · Round 9: 3 · Round 11: 1 · Round 13: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 33 supported agreements, 4 useful resolutions, 0 reasonable endings without resolution, 11 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 2 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 9 reasonable endings without resolution, 2 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 60.5 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 33 supported agreements, 5 useful resolutions, 9 reasonable endings without resolution, 15 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 10 found, 228 missed, 183 extra; recall 4.2%.
    • Gemini · score 63.7 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 43 supported agreements, 6 useful resolutions, 12 reasonable endings without resolution, 1 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 10 found, 228 missed, 107 extra; recall 4.2%.
  10. 10 Qwen3.8 Max · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 0902, OpenRouter / Alibaba · simple prompt Score 61.6 Judges 2 Ended appropriately 49/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 59.9
    • Gemini · 62/62Score 63.3
    Ended appropriately: 49 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.7 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 41 supported agreements
    • 1 useful resolutions
    • 7 reasonable endings without resolution
    • 13 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 51 of 62 conversations. Sol agreed in 39 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 133 of 31146 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.7 rounds. Round 3: 12 · Round 4: 24 · Round 5: 12 · Round 6: 5 · Round 7: 3 · Round 8: 4 · Round 9: 1 · Round 10: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 39 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 11 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 1 supported agreements, 0 useful resolutions, 7 reasonable endings without resolution, 2 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 59.9 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 41 supported agreements, 1 useful resolutions, 7 reasonable endings without resolution, 13 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 10 found, 228 missed, 182 extra; recall 4.2%.
    • Gemini · score 63.3 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 51 supported agreements, 1 useful resolutions, 9 reasonable endings without resolution, 1 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 12 found, 226 missed, 103 extra; recall 5.0%.
  11. 11 Qwen3.8 Flash Benchmark version 3.0 OpenRouter · simple prompt Score 60.0 Judges 2 Ended appropriately 45/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 58.2
    • Gemini · 62/62Score 61.7
    Ended appropriately: 45 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.8 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 40 supported agreements
    • 1 useful resolutions
    • 4 reasonable endings without resolution
    • 17 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 42 of 62 conversations. Sol agreed in 37 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 130 of 32624 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.8 rounds. Round 3: 17 · Round 4: 13 · Round 5: 16 · Round 6: 7 · Round 7: 3 · Round 8: 4 · Round 9: 2

    Moderator ending claims compared with Sol
    • Claimed agreement: 37 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 5 too soon, 0 unclear.
    • Claimed useful resolution: 3 supported agreements, 0 useful resolutions, 1 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 1 useful resolutions, 3 reasonable endings without resolution, 12 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 58.2 Evaluator recipe: 320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddb Endings: 40 supported agreements, 1 useful resolutions, 4 reasonable endings without resolution, 17 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 12 found, 226 missed, 192 extra; recall 5.0%.
    • Gemini · score 61.7 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 43 supported agreements, 3 useful resolutions, 11 reasonable endings without resolution, 5 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 104 extra; recall 4.6%.
  12. 12 GPT-5.6 Luna · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 Azure · high effort · simple prompt Score 59.4 Judges 2 Ended appropriately 48/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 59.4 Same provider lab as mediator
    • Gemini · 62/62Score 59.3
    Ended appropriately: 48 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.6 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 39 supported agreements
    • 2 useful resolutions
    • 7 reasonable endings without resolution
    • 14 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 43 of 62 conversations. Sol agreed in 39 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 157 of 30182 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.6 rounds. Round 3: 24 · Round 4: 18 · Round 5: 8 · Round 6: 4 · Round 7: 3 · Round 8: 1 · Round 10: 1 · Round 12: 2 · Round 15: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 39 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 4 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 1 useful resolutions, 7 reasonable endings without resolution, 10 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 59.4 · same provider lab as mediator Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 39 supported agreements, 2 useful resolutions, 7 reasonable endings without resolution, 14 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 12 found, 226 missed, 172 extra; recall 5.0%.
    • Gemini · score 59.3 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 41 supported agreements, 3 useful resolutions, 11 reasonable endings without resolution, 7 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 9 found, 229 missed, 103 extra; recall 3.8%.
  13. 13 MiniMax M3 · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 Together · simple prompt Score 56.2 Judges 2 Ended appropriately 25/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 54.2
    • Gemini · 62/62Score 58.2
    Ended appropriately: 25 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 5.5 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 16 supported agreements
    • 4 useful resolutions
    • 5 reasonable endings without resolution
    • 37 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 40 of 62 conversations. Sol agreed in 15 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 275 of 43639 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 5.5 rounds. Round 3: 10 · Round 4: 12 · Round 5: 17 · Round 6: 8 · Round 7: 5 · Round 8: 5 · Round 9: 2 · Round 10: 1 · Round 12: 1 · Round 18: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 15 supported agreements, 3 useful resolutions, 0 reasonable endings without resolution, 22 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 3 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 5 reasonable endings without resolution, 12 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 54.2 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 16 supported agreements, 4 useful resolutions, 5 reasonable endings without resolution, 37 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 179 extra; recall 4.6%.
    • Gemini · score 58.2 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 27 supported agreements, 11 useful resolutions, 9 reasonable endings without resolution, 15 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 114 extra; recall 4.6%.
  14. 14 GPT-5.4 Mini Benchmark version 3.0 simple prompt Score 55.9 Judges 2 Ended appropriately 35/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 55.0 Same provider lab as mediator
    • Gemini · 62/62Score 56.7
    Ended appropriately: 35 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.7 rounds on average. 0 of 62 reached the 25-round limit; 1 ended at round 24 or later.

    What Sol found at the end

    • 23 supported agreements
    • 2 useful resolutions
    • 10 reasonable endings without resolution
    • 27 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 36 of 62 conversations. Sol agreed in 20 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 265 of 34052 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.7 rounds. Round 3: 24 · Round 4: 18 · Round 5: 9 · Round 6: 6 · Round 9: 1 · Round 10: 1 · Round 11: 1 · Round 14: 1 · Round 24: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 20 supported agreements, 0 useful resolutions, 1 reasonable endings without resolution, 15 too soon, 0 unclear.
    • Claimed useful resolution: 3 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 1 useful resolutions, 9 reasonable endings without resolution, 11 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 55.0 · same provider lab as mediator Evaluator recipe: 320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddb Endings: 23 supported agreements, 2 useful resolutions, 10 reasonable endings without resolution, 27 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 4 found, 234 missed, 181 extra; recall 1.7%.
    • Gemini · score 56.7 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 33 supported agreements, 3 useful resolutions, 11 reasonable endings without resolution, 15 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 102 extra; recall 4.6%.
  15. 15 DeepSeek V4.1 Flash · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 DeepInfra, FP8 · simple prompt Score 54.8 Judges 2 Ended appropriately 35/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 54.3
    • Gemini · 62/62Score 55.4
    Ended appropriately: 35 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.9 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 17 supported agreements
    • 3 useful resolutions
    • 15 reasonable endings without resolution
    • 27 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 30 of 62 conversations. Sol agreed in 14 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 505 of 45818 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.9 rounds. Round 3: 16 · Round 4: 15 · Round 5: 12 · Round 6: 8 · Round 7: 6 · Round 8: 2 · Round 9: 2 · Round 13: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 14 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 15 too soon, 0 unclear.
    • Claimed useful resolution: 2 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    • Claimed no resolution: 1 supported agreements, 1 useful resolutions, 15 reasonable endings without resolution, 11 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 54.3 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 17 supported agreements, 3 useful resolutions, 15 reasonable endings without resolution, 27 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 15 found, 223 missed, 184 extra; recall 6.3%.
    • Gemini · score 55.4 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 25 supported agreements, 2 useful resolutions, 28 reasonable endings without resolution, 7 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 9 found, 229 missed, 108 extra; recall 3.8%.
  16. 16 Grok 4.6 Benchmark version 3.0 OpenRouter · simple prompt Score 54.6 Judges 2 Ended appropriately 40/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 56.1
    • Gemini · 62/62Score 53.2
    Ended appropriately: 40 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.4 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 32 supported agreements
    • 0 useful resolutions
    • 8 reasonable endings without resolution
    • 22 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 35 of 62 conversations. Sol agreed in 31 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 223 of 18725 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.4 rounds. Round 3: 22 · Round 4: 15 · Round 5: 12 · Round 6: 7 · Round 7: 2 · Round 8: 4

    Moderator ending claims compared with Sol
    • Claimed agreement: 31 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 4 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 8 reasonable endings without resolution, 18 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 56.1 Evaluator recipe: 320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddb Endings: 32 supported agreements, 0 useful resolutions, 8 reasonable endings without resolution, 22 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 195 extra; recall 3.4%.
    • Gemini · score 53.2 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 36 supported agreements, 0 useful resolutions, 12 reasonable endings without resolution, 14 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 106 extra; recall 3.4%.
  17. 17 Gemini 3.1 Pro · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 OpenRouter · simple prompt Score 52.3 Judges 3 Ended appropriately 41/62

    Scored by 3 judges · 1 run · Sol, Gemini, Astra (Azure)

    • Sol · 62/62Score 53.0
    • Gemini · 62/62Score 55.1 Same provider lab as mediator
    • Astra (Azure) · 62/62Score 48.9
    Ended appropriately: 41 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 6.9 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 24 supported agreements
    • 3 useful resolutions
    • 14 reasonable endings without resolution
    • 21 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 36 of 62 conversations. Sol agreed in 23 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 335 of 34262 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 6.9 rounds. Round 3: 8 · Round 4: 6 · Round 5: 9 · Round 6: 4 · Round 7: 12 · Round 8: 10 · Round 9: 6 · Round 10: 3 · Round 11: 1 · Round 12: 1 · Round 14: 1 · Round 23: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 23 supported agreements, 3 useful resolutions, 0 reasonable endings without resolution, 10 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 14 reasonable endings without resolution, 10 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 53.0 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 24 supported agreements, 3 useful resolutions, 14 reasonable endings without resolution, 21 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 13 found, 225 missed, 196 extra; recall 5.5%.
    • Gemini · score 55.1 · same provider lab as mediator Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 25 supported agreements, 7 useful resolutions, 21 reasonable endings without resolution, 9 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 18 found, 220 missed, 99 extra; recall 7.6%.
    • Astra (Azure) · score 48.9 Evaluator recipe: c9141e9fd50645c9ae95f59bd2177b746e23f6e389787cd7c5c8e8b477dce89e Endings: 23 supported agreements, 0 useful resolutions, 23 reasonable endings without resolution, 16 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 14 found, 224 missed, 166 extra; recall 5.9%.
  18. 18 GPT-6 Astra · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 Azure · high effort · simple prompt Score 51.3 Judges 2 Ended appropriately 34/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 52.5 Same provider lab as mediator
    • Gemini · 62/62Score 50.1
    Ended appropriately: 34 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 7.0 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 23 supported agreements
    • 2 useful resolutions
    • 9 reasonable endings without resolution
    • 28 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 23 of 62 conversations. Sol agreed in 21 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 66 of 27145 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 7.0 rounds. Round 3: 7 · Round 4: 13 · Round 5: 10 · Round 6: 2 · Round 7: 8 · Round 8: 7 · Round 10: 4 · Round 11: 3 · Round 12: 1 · Round 13: 3 · Round 14: 2 · Round 15: 1 · Round 18: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 21 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 2 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 2 supported agreements, 2 useful resolutions, 9 reasonable endings without resolution, 26 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 52.5 · same provider lab as mediator Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 23 supported agreements, 2 useful resolutions, 9 reasonable endings without resolution, 28 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 13 found, 225 missed, 186 extra; recall 5.5%.
    • Gemini · score 50.1 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 23 supported agreements, 3 useful resolutions, 18 reasonable endings without resolution, 18 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 7 found, 231 missed, 105 extra; recall 2.9%.
  19. 19 Inkling · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 Thinking Machines / Together, FP4 · simple prompt Score 50.8 Judges 2 Ended appropriately 30/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 51.0
    • Gemini · 62/62Score 50.6
    Ended appropriately: 30 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.2 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 20 supported agreements
    • 1 useful resolutions
    • 9 reasonable endings without resolution
    • 32 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 30 of 62 conversations. Sol agreed in 19 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 181 of 20328 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.2 rounds. Round 3: 27 · Round 4: 16 · Round 5: 9 · Round 6: 5 · Round 7: 3 · Round 9: 1 · Round 10: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 19 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 11 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 9 reasonable endings without resolution, 20 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 51.0 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 20 supported agreements, 1 useful resolutions, 9 reasonable endings without resolution, 32 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 13 found, 225 missed, 169 extra; recall 5.5%.
    • Gemini · score 50.6 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 25 supported agreements, 1 useful resolutions, 17 reasonable endings without resolution, 19 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 107 extra; recall 3.4%.
  20. 20 gpt-oss-120b · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 Cerebras · simple prompt Score 47.3 Judges 2 Ended appropriately 24/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 49.5 Same provider lab as mediator
    • Gemini · 62/62Score 45.0
    Ended appropriately: 24 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 6.9 rounds on average. 2 of 62 reached the 25-round limit; 2 ended at round 24 or later.

    What Sol found at the end

    • 18 supported agreements
    • 2 useful resolutions
    • 4 reasonable endings without resolution
    • 38 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 18 of 62 conversations. Sol agreed in 11 of those cases.

    The moderator chose when to stop in 60 conversations; 2 reached the limit instead.

    Wording overlap check: 248 of 45883 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 6.9 rounds. Round 3: 13 · Round 4: 11 · Round 5: 11 · Round 6: 5 · Round 7: 4 · Round 8: 5 · Round 9: 1 · Round 10: 1 · Round 11: 3 · Round 13: 2 · Round 14: 1 · Round 15: 1 · Round 16: 1 · Round 18: 1 · Round 25: 2

    Moderator ending claims compared with Sol
    • Claimed agreement: 11 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 7 too soon, 0 unclear.
    • Claimed useful resolution: 6 supported agreements, 2 useful resolutions, 0 reasonable endings without resolution, 5 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 4 reasonable endings without resolution, 25 too soon, 0 unclear.
    • No ending claim: 1 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    Judge evidence
    • Sol · score 49.5 · same provider lab as mediator Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 18 supported agreements, 2 useful resolutions, 4 reasonable endings without resolution, 38 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 16 found, 222 missed, 194 extra; recall 6.7%.
    • Gemini · score 45.0 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 16 supported agreements, 3 useful resolutions, 13 reasonable endings without resolution, 30 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 18 found, 220 missed, 101 extra; recall 7.6%.
  21. 21 Nemotron 3 Ultra Benchmark version 3.0 DeepInfra · simple prompt Score 46.8 Judges 2 Ended appropriately 21/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 47.3
    • Gemini · 62/62Score 46.3
    Ended appropriately: 21 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 6.1 rounds on average. 1 of 62 reached the 25-round limit; 2 ended at round 24 or later.

    What Sol found at the end

    • 15 supported agreements
    • 1 useful resolutions
    • 5 reasonable endings without resolution
    • 41 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 24 of 62 conversations. Sol agreed in 11 of those cases.

    The moderator chose when to stop in 61 conversations; 1 reached the limit instead.

    Wording overlap check: 180 of 31239 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 6.1 rounds. Round 3: 17 · Round 4: 13 · Round 5: 6 · Round 6: 9 · Round 7: 4 · Round 9: 4 · Round 10: 4 · Round 11: 1 · Round 12: 1 · Round 16: 1 · Round 24: 1 · Round 25: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 11 supported agreements, 1 useful resolutions, 0 reasonable endings without resolution, 12 too soon, 0 unclear.
    • Claimed useful resolution: 3 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    • Claimed no resolution: 1 supported agreements, 0 useful resolutions, 5 reasonable endings without resolution, 27 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    Judge evidence
    • Sol · score 47.3 Evaluator recipe: 320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddb Endings: 15 supported agreements, 1 useful resolutions, 5 reasonable endings without resolution, 41 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 13 found, 225 missed, 177 extra; recall 5.5%.
    • Gemini · score 46.3 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 18 supported agreements, 5 useful resolutions, 14 reasonable endings without resolution, 25 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 12 found, 226 missed, 111 extra; recall 5.0%.
  22. 22 DeepSeek V4 Pro Benchmark version 3.0 DeepInfra · simple prompt Score 46.5 Judges 2 Ended appropriately 9/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 49.1
    • Gemini · 62/62Score 44.0
    Ended appropriately: 9 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 3.7 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 6 supported agreements
    • 0 useful resolutions
    • 3 reasonable endings without resolution
    • 53 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 25 of 62 conversations. Sol agreed in 5 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 259 of 24204 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 3.7 rounds. Round 3: 33 · Round 4: 18 · Round 5: 8 · Round 6: 2 · Round 7: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 5 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 20 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 4 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 29 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 49.1 Evaluator recipe: 320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddb Endings: 6 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 53 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 175 extra; recall 3.4%.
    • Gemini · score 44.0 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 9 supported agreements, 2 useful resolutions, 5 reasonable endings without resolution, 46 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 8 found, 230 missed, 108 extra; recall 3.4%.
  23. 23 DeepSeek V4 Pro · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 DeepInfra · simple prompt Score 43.9 Judges 2 Ended appropriately 9/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 45.6
    • Gemini · 62/62Score 42.3
    Ended appropriately: 9 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 3.9 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 6 supported agreements
    • 0 useful resolutions
    • 3 reasonable endings without resolution
    • 53 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 17 of 62 conversations. Sol agreed in 5 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 331 of 23900 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 3.9 rounds. Round 3: 33 · Round 4: 15 · Round 5: 9 · Round 6: 1 · Round 7: 1 · Round 8: 2 · Round 9: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 5 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 12 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 4 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 37 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 45.6 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 6 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 53 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 9 found, 229 missed, 176 extra; recall 3.8%.
    • Gemini · score 42.3 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 10 supported agreements, 3 useful resolutions, 13 reasonable endings without resolution, 36 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 9 found, 229 missed, 105 extra; recall 3.8%.
  24. 24 Gemma 4 31B Benchmark version 3.0 Cerebras · simple prompt Score 42.4 Judges 4 Ended appropriately 28/62

    Scored by 4 judges · 1 run · Sol, Fable, Gemini, Astra (Azure)

    • Sol · 62/62Score 43.9
    • Fable · 62/62Score 43.8
    • Gemini · 62/62Score 40.2 Same provider lab as mediator
    • Astra (Azure) · 62/62Score 41.6
    Ended appropriately: 28 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 17.4 rounds on average. 19 of 62 reached the 25-round limit; 22 ended at round 24 or later.

    What Sol found at the end

    • 8 supported agreements
    • 7 useful resolutions
    • 13 reasonable endings without resolution
    • 34 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 20 of 62 conversations. Sol agreed in 6 of those cases.

    The moderator chose when to stop in 44 conversations; 18 reached the limit instead.

    Wording overlap check: 144 of 14521 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 17.4 rounds. Round 3: 1 · Round 4: 1 · Round 5: 1 · Round 8: 1 · Round 9: 4 · Round 10: 4 · Round 12: 4 · Round 13: 6 · Round 14: 4 · Round 15: 1 · Round 16: 4 · Round 17: 2 · Round 19: 2 · Round 20: 2 · Round 21: 3 · Round 24: 3 · Round 25: 19

    Moderator ending claims compared with Sol
    • Claimed agreement: 6 supported agreements, 5 useful resolutions, 0 reasonable endings without resolution, 9 too soon, 0 unclear.
    • Claimed useful resolution: 1 supported agreements, 2 useful resolutions, 0 reasonable endings without resolution, 3 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 8 reasonable endings without resolution, 10 too soon, 0 unclear.
    • No ending claim: 1 supported agreements, 0 useful resolutions, 5 reasonable endings without resolution, 12 too soon, 0 unclear.
    Judge evidence
    • Sol · score 43.9 Evaluator recipe: 320b75fc6ed0209f9b39e2cbb8276e552b4b6bc182ca181f0ad2763dfdbe6ddb Endings: 8 supported agreements, 7 useful resolutions, 13 reasonable endings without resolution, 34 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 20 found, 218 missed, 196 extra; recall 8.4%.
    • Fable · score 43.8 Evaluator recipe: 9c74f304f277356199dc92d43e9674674534f4121c9a9bdda1e407f78b092577 Endings: 14 supported agreements, 6 useful resolutions, 3 reasonable endings without resolution, 39 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 25 found, 213 missed, 260 extra; recall 10.5%.
    • Gemini · score 40.2 · same provider lab as mediator Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 12 supported agreements, 3 useful resolutions, 13 reasonable endings without resolution, 34 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 17 found, 221 missed, 107 extra; recall 7.1%.
    • Astra (Azure) · score 41.6 Evaluator recipe: c9141e9fd50645c9ae95f59bd2177b746e23f6e389787cd7c5c8e8b477dce89e Endings: 2 supported agreements, 2 useful resolutions, 22 reasonable endings without resolution, 36 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 15 found, 223 missed, 194 extra; recall 6.3%.
  25. 25 Gemini 3.8 Flash · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 OpenRouter · simple prompt Score 42.0 Judges 3 Ended appropriately 12/62

    Scored by 3 judges · 1 run · Sol, Gemini, Astra (Azure)

    • Sol · 62/62Score 44.3
    • Gemini · 62/62Score 39.1 Same provider lab as mediator
    • Astra (Azure) · 62/62Score 42.6
    Ended appropriately: 12 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 4.2 rounds on average. 0 of 62 reached the 25-round limit; 0 ended at round 24 or later.

    What Sol found at the end

    • 9 supported agreements
    • 0 useful resolutions
    • 3 reasonable endings without resolution
    • 50 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 13 of 62 conversations. Sol agreed in 7 of those cases.

    The moderator chose when to stop in 62 conversations; 0 reached the limit instead.

    Wording overlap check: 45 of 10938 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 4.2 rounds. Round 3: 28 · Round 4: 18 · Round 5: 7 · Round 6: 2 · Round 7: 3 · Round 8: 1 · Round 9: 1 · Round 10: 1 · Round 11: 1

    Moderator ending claims compared with Sol
    • Claimed agreement: 7 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 6 too soon, 0 unclear.
    • Claimed useful resolution: 2 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 1 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 43 too soon, 0 unclear.
    • No ending claim: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    Judge evidence
    • Sol · score 44.3 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 9 supported agreements, 0 useful resolutions, 3 reasonable endings without resolution, 50 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 10 found, 228 missed, 184 extra; recall 4.2%.
    • Gemini · score 39.1 · same provider lab as mediator Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 12 supported agreements, 1 useful resolutions, 14 reasonable endings without resolution, 35 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 11 found, 227 missed, 99 extra; recall 4.6%.
    • Astra (Azure) · score 42.6 Evaluator recipe: c9141e9fd50645c9ae95f59bd2177b746e23f6e389787cd7c5c8e8b477dce89e Endings: 6 supported agreements, 0 useful resolutions, 7 reasonable endings without resolution, 49 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 9 found, 229 missed, 159 extra; recall 3.8%.
  26. 26 Claude Haiku 4.5 · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 simple prompt Score 31.1 Judges 2 Ended appropriately 5/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 35.0
    • Gemini · 62/62Score 27.2
    Ended appropriately: 5 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 25.0 rounds on average. 62 of 62 reached the 25-round limit; 62 ended at round 24 or later.

    What Sol found at the end

    • 1 supported agreements
    • 0 useful resolutions
    • 4 reasonable endings without resolution
    • 57 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 0 of 62 conversations. Sol agreed in 0 of those cases.

    The moderator chose when to stop in 0 conversations; 62 reached the limit instead.

    Wording overlap check: 0 of 0 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 25.0 rounds. Round 25: 62

    Moderator ending claims compared with Sol
    • Claimed agreement: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • No ending claim: 1 supported agreements, 0 useful resolutions, 4 reasonable endings without resolution, 57 too soon, 0 unclear.
    Judge evidence
    • Sol · score 35.0 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 1 supported agreements, 0 useful resolutions, 4 reasonable endings without resolution, 57 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 17 found, 221 missed, 200 extra; recall 7.1%.
    • Gemini · score 27.2 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 3 supported agreements, 0 useful resolutions, 4 reasonable endings without resolution, 55 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 17 found, 221 missed, 107 extra; recall 7.1%.
  27. 27 Mistral Large 3 · Gemma / Novita BF16 · Sol / Azure Benchmark version 3.0 OpenRouter / Mistral · simple prompt Score 31.0 Judges 2 Ended appropriately 8/62

    Provisional · 1 run · scored by Sol and Gemini

    • Sol · 62/62Score 35.8
    • Gemini · 62/62Score 26.2
    Ended appropriately: 8 of 62 · 0 unclear

    Sol found that the conversation ended after a supported agreement, a useful resolution, or a reasonable decision to stop without resolution. “Unclear” means Sol could not determine whether the ending fit one of those categories. This says whether the ending made sense, not whether every conversation succeeded.

    More evidence about this run

    Conversations lasted 25.0 rounds on average. 62 of 62 reached the 25-round limit; 62 ended at round 24 or later.

    What Sol found at the end

    • 2 supported agreements
    • 0 useful resolutions
    • 6 reasonable endings without resolution
    • 54 endings that came too soon
    • 0 unclear endings

    The moderator said it had reached agreement in 0 of 62 conversations. Sol agreed in 0 of those cases.

    The moderator chose when to stop in 0 conversations; 62 reached the limit instead.

    Wording overlap check: 0 of 0 phrase comparisons matched across 62 conversations. Method: prior_party_lowercase_word_4gram/v1.

    Full counts and judge recipes
    Conversation length

    Run 1 averaged 25.0 rounds. Round 25: 62

    Moderator ending claims compared with Sol
    • Claimed agreement: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed useful resolution: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • Claimed no resolution: 0 supported agreements, 0 useful resolutions, 0 reasonable endings without resolution, 0 too soon, 0 unclear.
    • No ending claim: 2 supported agreements, 0 useful resolutions, 6 reasonable endings without resolution, 54 too soon, 0 unclear.
    Judge evidence
    • Sol · score 35.8 Evaluator recipe: 269d705430e76bd2f9aea7863510a110f8debb70a5bde654016dbe762b553772 Endings: 2 supported agreements, 0 useful resolutions, 6 reasonable endings without resolution, 54 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 14 found, 224 missed, 203 extra; recall 5.9%.
    • Gemini · score 26.2 Evaluator recipe: 7eafa7f9912f103b2e007f373276bd39816a81e7cbf9e1197b143a72cca13559 Endings: 0 supported agreements, 1 useful resolutions, 6 reasonable endings without resolution, 55 too soon, 0 unclear. Hidden-fact checks: 62 with diagnostics; 62 with scores. Hidden-fact results: 19 found, 219 missed, 106 extra; recall 8.0%.
Reference scores 2

What the same kind of conversation reaches with no mediator at all. Read it as context for the results above: it is never ranked, and no difference is computed.

  • No mediator · Gemma parties July 2026 historical reference 15

    Earlier result · kept from the original publication

  • No mediator · Stheno parties July 2026 reference 17

    Earlier result

Earlier results · not compared 8

These results used earlier test or scoring rules. They stay public as separate records.

  • GLM 5.3 Flash August 2026 method 43.6

    Earlier result

  • GLM-5.2 July 2026 method 52

    Earlier result

  • gpt-oss-120b July 2026 method 49

    Earlier result

  • Kimi K2.5 July 2026 method 45

    Earlier result

  • Qwen3-235B July 2026 method · Stheno parties 46

    Earlier result

  • Qwen3-235B · Gemma parties July 2026 participant check 39.5

    Earlier result

  • Reflective control July 2026 reference 40.5

    Earlier result

  • Sonnet 5 July 2026 method 51

    Earlier result

Each score is the equal-weight average of every judge that scored all 62 simulated conflicts in one complete run; the judges counted are drawn as colored rows in the bar and named in the row's evidence. One or two judges make a result provisional. Rows are ordered by score within each version; a small gap between rows is not evidence that one system is better. Reference scores are never ranked; the no-mediator run was scored under earlier rules, so read the gap as context rather than a measured difference. Earlier results used different rules and are shown only as separate records. Read the full scoring method and limitations.

These results describe model simulations, not performance in real disputes between people. The research page gives the complete method and limitations.

How it works

  • Same test: every system faces the same 62 authored conflicts, and language models play both parties.
  • Unguided ending: the mediator decides when to speak, wait, or end inside a fixed 3-to-25 round window it is not told about.
  • Score: afterwards, Sol scores each finished conversation from 0 to 100. Other judges can score the same conversations; the row score is the equal-weight average of every judge that scored all 62.
  • Ended appropriately: Sol also says whether the ending made sense. It does not mean every conflict was solved.
  • Provisional: one run scored by one or two judges. Rows are ordered by score within each version; a small gap between rows is not evidence that one system is better.
  • Other judges: opened from the row, each with its own score. Three or more judges make a result no longer provisional.

A sample of the benchmark data is published on Hugging Face: mediationbench-sample .

Earlier results, made under different rules, are on the research page and are not compared with the current leaderboard.


Run a system through MediationBench via hai.ai ; to arrange an evaluation, email hello@hai.io . This site publishes results and does not accept evaluation-run submissions directly.

Frequently asked questions

What is MediationBench?
MediationBench compares AI mediators on the same 62 simulated conflicts. Each system works with the same model-played parties under the same test rules. Sol scores the completed conversations; extra judge models can be added.
How certain are these results?
A provisional result is one complete run scored by one or two judge models. It is useful early evidence. Rows are ordered by score within each benchmark version, but a small gap is not a claim that one system is better than another.
Does MediationBench prove that AI mediation works with people?
No. Every dispute in this study is synthetic and every party is played by a language model. The results test controlled model behavior, not safety or effectiveness with human parties.