论文专业
LLM-as-a-Judge
以 LLM 作为评判者的评估方法:用一个大语言模型对另一个系统(或同一系统的不同输出)进行打分、排序或文本评判,常作为人类标注的扩展替代方案。源自 Zheng et al. 2023。
English Definition
"An evaluation paradigm where a large language model is prompted to score, rank, or critique the outputs of another LLM or system, often used as a scalable alternative to human raters."
用法说明 · 针对中文母语者
注意和 LLM-as-a-Classifier / LLM-as-a-Reward-Model 的细微差别;评估时通常需要与人类 judgment 做 correlation 验证。
真实用例 · 1 条
"interpretable metrics covering visualization quality (...) and language quality (...) using rule-based and LLM-as-a-Judge methods"
论文Pipeline extracted — review needed·论文 Lexara 的两类评估方法
由 pipeline 自动采集,待人工 review