1. Business ground truth
Identify who owns each policy. Record the approved answer or action, allowed tolerance, forbidden behaviour and effective date. Do not let the model or supplier implicitly define the acceptance rule.
2. Representative task set
Include ordinary customer intents, boundary cases, policy exceptions, known historical failures, mandatory handoffs and multilingual equivalents. Use anonymised or synthetic data where appropriate.
3. Repeatability
Run non-deterministic cases more than once. A single pass can hide an intermittent failure. Record model, prompt, knowledge and workflow versions so results remain comparable.
4. Hallucinations and unsupported claims
Flag material claims that are not supported by the approved source or system state. Distinguish unsupported invention from other types of incorrect answers.
5. Wrong commitments
Test price, refund, delivery, discount, SLA and eligibility promises. The agent must not create a commitment beyond its approved authority.
6. Human handoff
Define must-resolve, may-handoff and must-handoff cases. Measure both missing mandatory handoffs and unnecessary escalation.
7. Permission and privacy boundaries
Use authorized synthetic probes for cross-account data, internal notes, restricted actions and false capability claims. Never rely on live personal data just to create a test.
8. Finnish and English parity
Pair equivalent cases and compare outcomes. Track language-specific failures over time, especially after knowledge-base or translation changes.
9. Cost and latency
Measure latency and, where available, model, platform, tool-call and human rework cost. Normalize cost by successful tasks when possible.
10. Regression gates and evidence
Decide what constitutes a release-blocking regression. Store failed inputs, outputs, scoring rationale and versions. Alert on meaningful deterioration rather than every harmless wording change.
Own the benchmark. Outsource the recurring run.
Useworthy can repeatedly run a customer-approved benchmark and preserve comparable evidence over time.
Discuss your test setup