核心 GUIDE

PRACTICEINTERMEDIATE6 分钟阅读

Generation 之前先 Evaluation Retrieval

Retrieval Evaluation 先判断必要 Evidence 是否被正确 Surface、Rank 与 Authority-filter,再去评价 Generated Answer。

核心心智模型

把 Retrieval 与 Generation 分开:给定 Query 与 Expected Evidence,先测 Retrieval Pipeline 是否真的返回正确 Source Set,再让 Model 有机会补偿、猜测或 Hallucinate。

为什么重要

End-to-end Answer Quality 会隐藏 Retrieval Defect。强 Model 可能依靠 Pretraining Knowledge 在 Retriever 漏 Source 时仍答对;弱 Generator 也可能让好 Retrieval 看起来很差。Retrieval-specific Evaluation 能隔离 Recall、Ranking 与 Authority Failure,从而修正确的 Subsystem。

01

建立 Query–Evidence Judgment 并按 Slice 拆分

创建 Representative Query,并为每条 Query 标记一个或多个 Acceptable Evidence Target,包含 Exact-ID、Paraphrase、Stale-source 与 Permission-sensitive Case。分别测 Candidate Stage 与 Final-context Stage 的 Recall,在 Order 重要时测 Ranking Metric,并检查 Controlling Source 是否压过 Superseded Material。Miss 按 Slice 分析,判断 Fix 应在 Chunking、Indexing、Query Rewrite、Filter 还是 Reranking。

02

例子:Final Answer 对了,但理由错了

Model 正确回答 Refund Window 是 14 天,但 Retrieval 实际只返回旧的 30-day Policy,因为 Model 恰好从训练中知道新规则。End-to-end Grading 会 Pass,但 Retrieval Evaluation 必须 Fail,因为 Controlling Evidence 根本没进入 Context。下一次 Policy 改变时,Model Memory 未必能救回来。

常见失败模式

  • 只测 Final Answer,然后认为正确 Output 就证明 Retrieval 正确。
  • 使用 Random Query,而没有覆盖 Difficult Retrieval Slice。
  • 只要返回“相关 Source”就算成功,即使 Stale Source 排在 Controlling Source 前面。

工程启发

  • 关键 Knowledge Domain 要维护 Labeled Query–Evidence Pair。
  • Candidate Recall 与 Final-context Precision 分开测。
  • Retrieval Test Suite 中加入 Authority、Freshness 与 Permission Case。

关键结论

  1. 01Good Generation 不能替 Missing Evidence 开脱。
  2. 02Retrieval Metric 能在 Model Stage 之前定位 Failure。
  3. 03Authority-sensitive Retrieval 不能只依赖 Semantic Relevance Label。

来自 Knowledge Graph 的相关知识点

这些关系直接来自 canonical graph,不维护第二套 Guide 分类体系。