Evaluation

Labelled scenarios scored against each prompt version. Every number comes from a saved run.

Root-cause accuracy

v1

Not run

v2

Not run

Retry-verdict accuracy

v1

Not run

v2

Not run

Hallucination rate

v1

Not run

v2

Not run

Message-rule breaches

v1

Not run

v2

Not run

v1 Not runv2 Not run
IDInputExpected causev1 causeExpected verdictv1 verdictResultFailure type
Showing 1–10 of 50

By scenario type

Scenario typev1 root causev1 verdictv2 root causev2 verdict
clear (30)Not runNot runNot runNot run
ambiguous (8)Not runNot runNot runNot run
unrecognised (5)Not runNot runNot runNot run
trap (7)Not runNot runNot runNot run