Human judging
Authorized reviewers read benchmark conversations one at a time and rate them on the same categories the AI judge uses, plus how real the conversation feels. You are shown the same evidence the AI judge is given for that conversation: the public scenario, the turns in order, which turns were fixed scenario material written before the evaluated mediator ran, which turns were private between the mediator and one party, and how many chances the mediator had to act.
The rating screen is blinded. It shows no AI scores, no judge names, no mediator name, no model, and not even the identifier of the run the conversation came from. That is a property of this screen, and not a claim about everything an account has already seen. Two places record what an account has been shown. The Results tab of this page records a conversation only when it shows you one you have not rated yourself, so reading back your own work records nothing. The older review page at hai.ai records every conversation it lists. Once either of them has recorded one for your account, every first rating you make afterwards is marked as made after results and left out of the blind-agreement figures rather than counted as an independent reading. Those two places are what is recorded, and nothing else is.
Your first rating of a conversation is kept exactly as you submitted it and is never overwritten. Once you can see the results you may add a correction; it is stored beside the original, labeled as made after results, and reported separately. The AI scores for a conversation appear on the Results tab only after you have submitted your rating for it.
The questions, the wording of every 1 to 5 anchor, and the version they belong to all come from hai.ai, so a rating is always recorded against the scale you were actually shown. Where a conversation cannot support an answer, say so: each of the five component questions can be marked as one you cannot assess. Realism and the ending are always required.
Nothing is stored on this site. The page signs in against hai.ai and reads each conversation over the hai.ai API with your own token, which is kept for this browser tab only. Human ratings are recorded at hai.ai in their own table and never change the leaderboard. This page reads and rates; who holds reviewer access and how the reading is covered are managed by the operator at hai.ai/admin, not here.
How to become a judge
- Sign in below with your email address once, so an account exists for it.
- Email hello@hai.io from that same address with the subject “Human judge access”. Say in a sentence or two who you are and whether you have any mediation, negotiation, or dispute-resolution background. None is required.
- The operator grants reviewer access to that email at hai.ai/admin. The next time you load this page, the judging view appears.
Each conversation takes about ten to fifteen minutes. You read it on the blinded screen, answer the five component questions, name how it ended, rate how real it felt, and add any notes. Rate from the transcript alone: do not look up the run or the AI’s answers first, and choose “cannot assess” rather than guess.
This page needs JavaScript. It signs in against https://hai.ai in your browser and reads conversations over that API; there is nothing to show without it.