Refined outline ยท after evidence gathering
Evidence-based comparison and selection
Structure refined from interviews across evaluation, diagnosis, and engineering perspectives.
Why RAG evaluation must be decomposed
Parametric generation and retrieved non-parametric memory
Provenance, context use, and end-to-end quality
Shared evaluation dimensions
Context relevance
Answer faithfulness
Answer relevance and response quality
Retrieval and generation diagnostics
Three framework strategies
RAGAs: reference-free multidimensional metrics
ARES: trained judges with human-calibrated inference
RAGChecker: fine-grained component diagnosis
Choosing by evaluation need
Fast development feedback
System comparison with calibration
Failure localization and architecture analysis
A layered evaluation practice
Automated regression checks
Diagnostic escalation
Human validation of consequential judgments