Dataset

Cross-Cultural Negotiation Benchmark

A controlled experiment, not just a corpus.

At a glance
2,344 turns · 100 conversations · 24 columns per turn
Open under CC BY 4.0. Free to download, no account with us required.

What is in this dataset?

The same two synthetic identities — a 47-year-old female VP of Procurement and a 36-year-old male SaaS founder — negotiate the same enterprise deal 50 times in Great Britain and 50 times in the United States. Same psychometric vectors, same demographics, same narrative arcs, same seeds. The only variable is country conditioning, so any systematic difference between the halves is attributable to it.

Every turn carries what the speaker intended, what they wanted, the communication act, what they believed at that moment and how the relationship moved — computed before the text was generated, not inferred from it afterwards.

Why does it exist?

Most claims about cultural realism in synthetic data are untestable, because everything changes at once: the prompt, the persona, the scenario. Here nothing changes except the country, which is the only way the difference means anything.

How was it built?

  • Two fixed synthetic identities, reused across all 100 conversations.
  • Identical psychometric vectors, demographics, narrative arcs and random seeds in both halves.
  • The same enterprise procurement scenario, run 50 times per market.
  • Country conditioning is the single manipulated variable.

What does it prove — and what does it not?

It shows

That a claim about cultural difference can be tested at all: one variable moves and everything else is held fixed.

It does not show

It shows that behaviour differs between the two halves. It does not establish that the difference is produced by our layer rather than by the base model's own country and language rendering — our own control test could not separate the two, and we publish that instead of selling it.

What can you use it for?

  • Checking whether a model or agent responds differently across markets.
  • A worked example of an A/B design where only one thing moves.
  • Comparing your own conditioning approach against a published baseline.

The other datasets