DesignLex
/设计研究/LLM-as-a-Judge
论文专业

LLM-as-a-Judge

以 LLM 作为评判者的评估方法:用一个大语言模型对另一个系统(或同一系统的不同输出)进行打分、排序或文本评判,常作为人类标注的扩展替代方案。源自 Zheng et al. 2023。

English Definition

"An evaluation paradigm where a large language model is prompted to score, rank, or critique the outputs of another LLM or system, often used as a scalable alternative to human raters."

用法说明 · 针对中文母语者

注意和 LLM-as-a-Classifier / LLM-as-a-Reward-Model 的细微差别;评估时通常需要与人类 judgment 做 correlation 验证。

真实用例 · 1

"interpretable metrics covering visualization quality (...) and language quality (...) using rule-based and LLM-as-a-Judge methods"

论文Pipeline extracted — review needed·论文 Lexara 的两类评估方法
由 pipeline 自动采集,待人工 review