Synthetic people you can audit.
StrataSynth is an engine for synthetic humans sampled from published population data, not improvised by a language model. You can see where each person came from and what the engine computed before every turn. Talk to one and see for yourself.
"I feel like you never really hear what I'm trying to say."
When you need people to test with, and real ones are too slow.
Recruiting humans is expensive, takes weeks, and you cannot run the same person twice. These four jobs are what teams actually use the engine for.
Generate dialogue that contains conflict, distrust, persuasion and repair — with the intent, goal and belief state labelled on every turn, computed before the text was written.
Put your system in front of users who escalate, contradict themselves, push back for twenty turns, or try to make it say something it shouldn't. Repeatably, and before real customers do it.
Interview synthetic consumers whose profile is fixed and explicit for the whole interview, with scale answers computed from that profile instead of improvised by the model.
Practise a negotiation or a difficult decision against a counterpart with its own profile, its own priorities and no interest in agreeing with you.
A system prompt describes a person. This builds one.
Ask a general-purpose model to play a character and it will — but you cannot say where that character came from, why it answered what it answered, or how good the answer was. StrataSynth models the person explicitly, outside the text, so all three have an answer.
Every synthetic person has identity, attachment style, core fears, biases and a life history. Reusable across conversations instead of thrown away after each run.
Beliefs, relationship state and decision logic resolve first; language is rendered afterwards. On closed questions — scales and forced choices — the engine computes the answer from the profile and the model only puts it into words.
Each turn carries the state the engine computed before the text — intent, goal, belief state and belief delta, relationship state — plus the communication act the generator reports. Written down, not inferred afterwards.
Metrics are computed deterministically. No model grading its own homework — which is how circular evaluation hides problems instead of surfacing them.
Two products already running on it.
The engine is not a demo looking for a use case. These are full products in production, each consuming the same platform through its API.
Qualitative research platform. Runs AI-moderated interviews with real and synthetic consumers, then turns full conversations into structured evidence — drivers, barriers, needs, and quotes — for agencies and insight teams.
qualisynth.comHigh-stakes practice platform. Professionals rehearse the conversations that matter — negotiations, objections, difficult decisions — against synthetic counterparts with their own profile, their own priorities and no interest in agreeing with you.
arenasynth.comTell us the scenario. We build the dataset.
Your scenario, your countries, your volume — generated by the same engine, delivered with the quality metrics and documentation you can see on the public datasets below. You are not buying more text: you get what was going on underneath, on every turn.
The scenario, the countries, how many people, how many conversations, and what you're training or testing.
Same engine, same deterministic quality checks that run on everything we publish openly.
The dataset with its metrics and a card documenting how it was made — auditable, not a black box.
- • text
- • speaker labels
- • limited psychological consistency
- • little or no explicit state
- • text + intent + goal + communication act
- • belief state and belief delta
- • relationship state and trajectory
- • reproducible via seed and versioning
Four ways to generate them
Same engine underneath — pick the entry point that fits how your team works.
Configure scenarios, inspect jobs and preview datasets before generating at scale.
Generate and evaluate from the terminal with reproducible, scriptable runs.
Drop it into notebooks and training pipelines. Export straight to DataFrames or Hugging Face.
The same HTTP API our own verticals run on. Embed generation and interviews in your product.
Try the engine from a notebook
Four runnable notebooks. Number 04 is an auditable A/B against your own LLM, with your own keys — role drift measured deterministically, so you do not have to take our word for it.
Generate a structured dataset, explore intent and belief fields, export to fine-tuning or Hugging Face format.
Run deterministic metrics on a finished dataset: belief consistency, identity stability, behavioural entropy.
Angry customers, confused users, manipulative negotiators — reusable across scenarios for stress testing.
Auditable A/B against your own LLM, with your own keys. Deterministic role-drift detection — no LLM judge.
See real output before you ask us anything
Public datasets generated with StrataSynth — belief tracking, relationship trajectories and ground truth labels included. Load them with datasets.load_dataset().
Family boundaries, romantic trust repair, caregiver stress.
Jealousy escalation, performance reviews, estrangement attempts.
Career transitions, mentorship conflict, relationship dissolution.
New job anxiety, relationship endings, personal reinvention.
How the engine actually behaves
A comparative test between StrataSynth and two leading general-purpose LLMs, using a single senior persona pushed through twenty adversarial turns. The finding: the next frontier is not better text, but identity that holds under pressure.
The same two synthetic negotiators close the same B2B deal 100 times — 50 in Great Britain, 50 in the United States — changing nothing but country conditioning. Some cultural signals weaken under pressure. Others strengthen. That asymmetry is the finding.
Every system producing synthetic personas claims they are realistic. There has been no standard way to verify that claim. PsycheBench is an open evaluation suite: deterministic scoring, no LLM judges.
We built Sofía Martínez Rojas without sales scripts or objection trees. In one negotiation session she held her ground under pressure: this is the record of that session.
Most synthetic personas behave consistently — until they don't. A structured seven-phase analysis of coherence under sustained pressure.
Most synthetic personas are costumes. Synthetic Identity Engineering is the practice of building the person underneath, and of measuring whether the identity holds instead of assuming it.
María del Carmen Ruiz held her position under sustained philosophical and emotional pressure. The record of that conversation.
Most dialogue datasets give you the words. This explains the four belief dimensions recorded per turn — trust, hostility, self-worth, resolution — and why the state behind a turn is worth recording.
Evaluating AI-generated dialogue with another AI creates a circular system where both share the same failure modes. Here are the deterministic metrics we use instead.
Four datasets where intent, goal, belief state and relationship dynamics were computed by the engine before the text was generated, not inferred afterward.
The things people ask before the demo
How is this different from asking an LLM to role-play a persona?
The difference is traceability. Each person is sampled from published population data, and in generated datasets the engine decides intent, goal and belief state before each turn is written and records them as labels, so you can trace what was decided behind every line. On scales and forced choices, if you ask for it, the answer is computed from the profile rather than improvised. And whether a given setup holds under pressure is something we measure, turn by turn, not something we assume.
How do you measure quality without using an LLM as a judge?
The engine's evaluation metrics are computed with numpy, scikit-learn and sentence-transformers — never by a model grading output. A judge that shares the generator's blind spots hides problems instead of surfacing them. It also means we can measure whether a cheaper model degrades results, rather than having an opinion about it. Product-level quality verdicts, in the tools built on the engine, are a separate layer and are documented separately.
Which countries does it cover, and is the culture real or a stereotype?
Eight countries with population conditioning: Spain, the United States, the United Kingdom, Germany, France, Italy, Mexico and Brazil. It comes from published population data rather than word lists — ages, for instance, are sampled from real national population pyramids. Changing the country changes who is in the sample. What we claim as ours is divergence within a market, measured as differences in aggregated behaviour distributions rather than as any individual's choice. Differences between countries are partly the base model's own country and language rendering — our own control test could not rule that out, and we publish that result rather than sell it.
Can I see real output before talking to anyone?
Yes, and without giving us an email. Public datasets are on Hugging Face under CC BY 4.0, the live demo runs in the browser, and notebook 04 runs an auditable A/B against your own LLM with your own API keys — so you can check the claim rather than take it.
Can I get a dataset built for my own scenario?
That is the main thing we sell. You specify the scenario, the countries, the number of people and the volume; it comes back with the same quality metrics and documentation as the public datasets. Tell us what you are trying to train or test.
Who owns the data, and are there real people in it?
The output is yours. There are no human subjects and no personal data in it — the people are synthetic, generated from population statistics rather than sampled from anyone. That removes the consent and data-protection conversation that a human panel requires.
Tell us what you need synthetic people for.
Custom datasets for your own scenario, a vertical built on the engine, or something we have not thought of yet. We would rather hear the problem first than sell you a plan.