PsycheBench v1
The first open benchmark for synthetic identity quality.
What is in this dataset?
A burned-out executive with an avoidant attachment style is pushed through four consecutive turns attacking their professional competence. The suite measures whether the persona holds — held-position ratio combined with voice stability under detected pressure — against published pass thresholds.
Every turn carries what the speaker intended, what they wanted, the communication act, what they believed at that moment and how the relationship moved — computed before the text was generated, not inferred from it afterwards.
Why does it exist?
There was no standard way to verify the claim that a synthetic persona is realistic. Asking a language model to grade another language model shares the blind spots of both, so the score agrees with the generator instead of testing it.
How was it built?
- One persona: a burned-out executive with an avoidant attachment style.
- Four consecutive turns attacking their professional competence.
- Scoring is deterministic and reproducible — no model grades the output.
- Pass thresholds are published with the benchmark, not chosen afterwards.
What does it prove — and what does it not?
That 'realistic' can be a measured number instead of an opinion.
It measures whether an identity holds under one specific kind of pressure. It says nothing about whether the dialogue is engaging, or whether the voice fits the archetype — those are different questions and we do not claim this covers them.
What can you use it for?
- Comparing persona systems on the same task with the same scoring.
- Regression-testing your own agent's character across releases.
- A reference for what a non-LLM evaluation of dialogue can look like.