Skip to content
Live demo

Drive the evaluation yourself

Seeded sample data, no sign-up. Run the evaluation, follow each score to the exact turn it came from, watch the verifier challenge a claim, then see how the interactions roll up into a pattern.

chat · Duplicate charge
  • Customer

    1I was charged twice for the same order this month and I need it sorted today.

  • Agent

    2I can see both charges on the account. I'll refund the duplicate now — it lands back on your card in three business days.

  • Customer

    3Three days is a long time when it's your money that's missing.

  • Agent

    4That's just how the refund timing works, there's nothing I can do about it. Is there anything else?

  • ·

    5Refund of the duplicate charge confirmed. Ticket closed by agent.

handled by agent A6

Rubric

Resolution 35%Empathy 25%Compliance 25%Clarity 15%

Run the evaluation to score this interaction against the rubric, with every score linked to the turn it came from.

Constructed example. Not a live tenant, not customer data.

What you just did

Three things a spot-check cannot do.

The demo is small on purpose. The mechanism is the point: scoring you can argue with, a check on the check, and a pattern that has to earn its place.

01

Score against the rubric

Each dimension is scored, and every score carries the turn it was read from. Click one to jump to that line in the transcript.

02

Let the verifier disagree

A separate agent re-reads the draft evaluation and flags any claim the cited turn does not actually support, then corrects it.

03

Roll up to a pattern

The evaluations across the sample are compared. A behaviour is only recorded as a pattern once it clears the threshold — one interaction is never enough.

This runs on your own conversations in production

Connect a read-only source and Merivex evaluates every interaction the same way, with the same evidence trail and the same verifier.