Dataset

Agent Stress Test

Users who escalate, accuse and push back.

At a glance
2,068 rows · scenarios ROM-02 · PRO-02 · FAM-03
Open under CC BY 4.0. Free to download, no account with us required.

What is in this dataset?

Jealousy escalation, performance reviews that go wrong, attempts to cut someone off. The conversations a support or coaching agent handles worst, available before your customers run them for you.

Every turn carries what the speaker intended, what they wanted, the communication act, what they believed at that moment and how the relationship moved — computed before the text was generated, not inferred from it afterwards.

Why does it exist?

Conversational systems are usually evaluated on cooperative users. The failures that matter happen with the other kind, and those arrive in production rather than in the test set.

How was it built?

  • Three scenarios: jealousy escalation, difficult performance reviews, estrangement attempts.
  • The difficult speaker holds a coherent position instead of escalating at random.
  • Every turn ships with the intent and goal behind it, so a failure can be located.

What does it prove — and what does it not?

It shows

Where a system breaks — before production does the experiment.

It does not show

Adversarial here means socially difficult, not a security test. It is not a jailbreak or prompt-injection corpus and should not be used as one.

What can you use it for?

  • Red-teaming a support, coaching or sales agent before launch.
  • Building regression tests around the turns where a system previously gave way.
  • Measuring whether a model concedes a position under repeated pressure.

The other datasets