Drive the evaluation yourself
Seeded sample data, no sign-up. Run the evaluation, follow each score to the exact turn it came from, watch the verifier challenge a claim, then see how the interactions roll up into a pattern.
- Customer
1I was charged twice for the same order this month and I need it sorted today.
- Agent
2I can see both charges on the account. I'll refund the duplicate now — it lands back on your card in three business days.
- Customer
3Three days is a long time when it's your money that's missing.
- Agent
4That's just how the refund timing works, there's nothing I can do about it. Is there anything else?
- ·
5Refund of the duplicate charge confirmed. Ticket closed by agent.
Rubric
Run the evaluation to score this interaction against the rubric, with every score linked to the turn it came from.
Constructed example. Not a live tenant, not customer data.
What you just did
Three things a spot-check cannot do.
The demo is small on purpose. The mechanism is the point: scoring you can argue with, a check on the check, and a pattern that has to earn its place.
Score against the rubric
Each dimension is scored, and every score carries the turn it was read from. Click one to jump to that line in the transcript.
Let the verifier disagree
A separate agent re-reads the draft evaluation and flags any claim the cited turn does not actually support, then corrects it.
Roll up to a pattern
The evaluations across the sample are compared. A behaviour is only recorded as a pattern once it clears the threshold — one interaction is never enough.
This runs on your own conversations in production
Connect a read-only source and Merivex evaluates every interaction the same way, with the same evidence trail and the same verifier.