DesignLex
/设计研究/LLM-as-judge
论文专业

LLM-as-judge

以 LLM 作为评审,用一个大语言模型自动给另一个模型的输出打分或评判,常作为人工评估的可扩展替代方案。

English Definition

"An automated evaluation methodology in which one large language model scores or rates the outputs of another model, often used as a scalable alternative to human raters."

用法说明 · 针对中文母语者

常保留连字符写法;也可写作 'LLM-as-a-judge'。在 HCI/AI 评测论文中频繁出现。

真实用例 · 3

"We conduct a fine-grained evaluation of three LLMs using both human annotators and LLM-as-judge methods, utilizing rubrics created in consultation with CSOs."

论文Pipeline extracted — review needed·介绍评测方法时,把人工标注和 LLM-as-judge 并列

"In parallel, LLM-as-judge methods are employed to enhance efficiency, ensure consistency, and provide cross-validation of human ratings."

论文Pipeline extracted — review needed·解释为什么采用 LLM-as-judge

"We also analyze the evaluation behavior of humans and LLM-as-judges, finding that automated judges produce extremely stable yet compressed score distributions, while human raters exhibit broader variance and stronger sensitivity to linguistic nuance."

论文Pipeline extracted — review needed·对比人评与 LLM 评的得分分布
由 pipeline 自动采集,待人工 review