Evaluation Ai Human Agreement
Ad-hoc AI-vs-human agreement over this evaluation’s EXISTING runs.
Pairs each conversation’s AI verdict against its human verdict per criterion and returns Cohen’s κ + percent agreement — the same κ math as a calibration set’s Gate 2, computed on-the-fly over the runs that already exist (no calibration set required).
DELIBERATE divergence from a formal calibration set: this takes the
single MOST-RECENT run per (conversation, assessor) rather than pooling
every assignment into a consensus, so the two can differ when a
conversation has multiple runs by the same assessor. na outcomes are
excluded from κ (an N/A is “doesn’t apply”, not a disagreement); for a
rigorous, multi-rater, applicability-aware measure, use a calibration set.

