Guides Knowledge AI
Analytics / Compare feedback

Compare feedback

The Compare feedback tab shows a side-by-side analysis of human ratings and AI Judge scores on the same answers. It requires both feedback sources to have rated at least some of the same answers. Key metrics include:

  • Matched turns — the number of answers that have been rated by both a human and the AI Judge.
  • Verdict agreement — how often the human and AI agree on whether an answer is good, needs improvement, or is poor.
  • Overall quality — a Human vs. AI Judge score comparison showing each side's average quality rating and evaluation count.
  • Per-dimension comparison — a table comparing Human and AI averages for each dimension (Relevance, Completeness, Faithfulness, Coherence) along with the MAE (Mean Absolute Error). An MAE near 0 means strong agreement; 1+ indicates significant disagreement.
Tip
Focus on dimensions with high MAE values — these are areas where the AI Judge and your users disagree most, and represent your highest-priority improvement opportunities.

FAQ

What is the Compare feedback tab?
The Compare feedback tab shows a side-by-side analysis of human ratings and AI Judge scores on the same answers. It requires both feedback sources to have rated at least some of the same answers so there is overlap to compare.
What does Matched turns mean in Compare feedback?
Matched turns is the number of answers that have been rated by both a human and the AI Judge. It reflects how many shared data points exist for the comparison.
What does Verdict agreement measure in Compare feedback?
Verdict agreement measures how often the human and AI agree on whether an answer is good, needs improvement, or is poor. It summarizes alignment between the two feedback sources at the verdict level.
What does Overall quality show in Compare feedback?
Overall quality is a Human vs. AI Judge score comparison. It shows each side's average quality rating along with the evaluation count.
What is the per-dimension comparison MAE, and how should I use it?
Per-dimension comparison shows Human and AI averages for each dimension (Relevance, Completeness, Faithfulness, Coherence) and includes MAE (Mean Absolute Error). An MAE near 0 means strong agreement, while 1+ indicates significant disagreement. Focus on dimensions with high MAE values because they highlight where the AI Judge and users disagree most and are high-priority improvement opportunities.