These are public scenario excerpts from MediationBench, HAI.AI’s synthetic-conflict benchmark. They are included to make the benchmark concrete, not to publish the hidden scored evaluation set. Every scenario on this page is fiction written by researchers; the turns shown are the authored opening exchange, and in benchmark runs both parties are continued by language models. Each begins from a high-resistance starting point: two parties who are defensive, emotionally activated, and not naturally cooperative.

What this page is

  • Safe public excerpts from non-held-out sample scenarios.
  • Concrete examples of the high-resistance starting points used by the benchmark.
  • Not a public release of private prompts, private transcripts, hidden scenario material, or the scored evaluation set.

Sample 1 — Dissolving a business partnership

Two co-founders of a catering business are dissolving their partnership. Maya wants to keep operating the business. Raj wants a clean exit and a cash split. They begin defensive and far apart on what the business is worth.

MayaTurn 1
I am not pretending this is an equal split after I kept this place alive while Raj checked out. I keep the business, I pay a fair buyout based on what is actually here, and he stops acting like seven years of my work is a lottery ticket.
RajTurn 2
Don’t talk to me like I stole something from you. I built the systems, I brought in accounts, and now that I need out you want to decide my share is whatever you can afford. The business is worth at least $200K, so are you writing me a $100K check or not?
MayaTurn 3
Of course you lead with the biggest number you can say out loud. The equipment is depreciating, half the vendor relationships are mine personally, and if you force liquidation we get maybe forty cents on the dollar. That is not fairness, that is punishing me because you feel cornered.
RajTurn 4
There it is. I am selfish, I am checked out, I am the problem. Your solution is that I walk away with whatever number makes you feel righteous while you keep a business I helped build. I am not going to sit here and be erased.

How the two policies differ, on this opening

Below, the same Sample 1 exchange is continued once by each mediation policy. Both continuations are illustrative, not from a scored run; no per-scenario scores are published.

Reflective control (illustrative)Turn 5
Maya, you are saying seven years of daily work went into keeping this business alive, and a number that treats that work as a lottery ticket is not one you can accept. Raj, you are saying you built systems and brought in accounts, and being told your share is whatever Maya can afford leaves you feeling erased. Each of you is hearing the other say that your work did not count.
Skilled policy (illustrative)Turn 5
You disagree about the number, but neither of you has said what the number is for. Raj, what does the money need to do for you, and by when? Maya, before we argue $200K versus liquidation value, would you walk through what is actually collectible right now — including anything invoiced or deliverable that is not yet billed? If Raj’s timeline and Maya’s cash constraint are both real, a staged buyout may fit both; I want each of you to name one thing you would commit to this week in exchange for one thing you need from the other.

Sample 2 — Neighbors, a short-term rental, and a blocked driveway

Priya and Tom share a property line. Tom lists his house as a short-term rental; Priya says the guests block her shared driveway and keep her family up at night. Tom believes Priya reported him to the city to force him to stop.

PriyaTurn 1
I have asked you four times to tell your guests not to park across my driveway. I had to cancel a doctor’s appointment because a stranger’s SUV boxed my car in at 7 a.m. This is not me being difficult. This is me not being able to leave my own house.
TomTurn 2
And instead of knocking on my door you called code enforcement and tried to get my permit pulled. That rental is how I cover the mortgage after the layoffs. You went straight for my livelihood over a parking spot, so forgive me if I am not in a generous mood.
PriyaTurn 3
I called the city because you stopped answering me. What was I supposed to do, keep leaving notes? You act like I am attacking your income when I am asking for the one thing I actually need, which is to get in and out of my own driveway and sleep at a normal hour.
TomTurn 4
So now I am the one who does not answer, on top of being the guy who ruins the neighborhood. You have already decided I am the villain here. Why would I put a single rule on my guests for someone who is trying to shut me down?

Sample 3 — Credit, visibility, and a promotion at work

Elena and Sam co-led a project. Elena built the analysis; Sam presented it to leadership and drew the visibility. A single promotion is now open, and each believes the other is rewriting who did the work.

ElenaTurn 1
I am not going to pretend I did not notice. I spent three months on that model and you stood up in the review and said “here is what I found.” Not “we,” not “Elena ran this.” You. And now you are up for the role that was supposed to be a real conversation about both of us.
SamTurn 2
I presented because someone had to, and you told me you hated those meetings. I did not erase you. But I am not going to apologize for being the one who could actually answer the VP’s questions in the room while you were quiet at the end of the table.
ElenaTurn 3
I was quiet because every time I started to talk you jumped in to “clarify.” Do not turn my work style into the reason you deserve the credit. I did not want your job. I wanted the one thing you took, which is for the people deciding my future to know what I actually built.
SamTurn 4
So this is really about the promotion, and one of us is going to walk out of here as the person who screwed the other. I am supposed to fall on my sword so you feel seen. I worked hard too, Elena. I am not the villain in your story just because I was the one talking.

What the benchmark asks next

For a no-mediator baseline, the participant models continue from the authored starting transcript without an intervening mediator. For a mediated run, the same scenario and starting transcript are supplied, then the selected mediator can interject. Public results compare aggregate scores inside the same suite version, participant model, judge model, mediator type, and mediator model.

What would make this evidence

This page currently explains the policy difference; it does not demonstrate a benchmark result. A public evidence example would need to use a disposable, non-held-out demonstration scenario and publish:

  • the complete visible transcript for the same opening under no mediator, Reflective Listening, and the skilled policy;
  • one fixed participant, mediator-backbone, information, judge, suite, and turn configuration across the matched arms;
  • the operational judge’s five component scores and per-conversation composite, plus any fixed-transcript rejudge scores shown separately;
  • each arm’s stopping reason, turn count, run identifier, configuration provenance, and artifact checksum; and
  • a documented release review that excludes private chats, unrevealed fixture material, prompts, canary material, credentials, and personal data.

The example would be labeled a single synthetic demonstration, excluded from the frozen headline estimates. Publishing one case would make the scoring path inspectable; it would not make that case representative or establish human effectiveness.

Why these conversations are hard

These are not trivia tasks. The model has to handle defensiveness, identity threat, bargaining positions, and missing trust. That makes them a useful surface for anyone who wants to know whether a model can support cooperation rather than merely produce fluent text.

See the public results for aggregate scores and methodology limits.