Refined outline ยท after evidence gathering

Evidence-based comparison and selection

Structure refined from interviews across evaluation, diagnosis, and engineering perspectives.

Why RAG evaluation must be decomposed

Parametric generation and retrieved non-parametric memory

Provenance, context use, and end-to-end quality

Shared evaluation dimensions

Context relevance

Answer faithfulness

Answer relevance and response quality

Retrieval and generation diagnostics

Three framework strategies

RAGAs: reference-free multidimensional metrics

ARES: trained judges with human-calibrated inference

RAGChecker: fine-grained component diagnosis

Choosing by evaluation need

Fast development feedback

System comparison with calibration

Failure localization and architecture analysis

A layered evaluation practice

Automated regression checks

Diagnostic escalation

Human validation of consequential judgments