LLM-as-judge
以 LLM 作为评审,用一个大语言模型自动给另一个模型的输出打分或评判,常作为人工评估的可扩展替代方案。
English Definition
"An automated evaluation methodology in which one large language model scores or rates the outputs of another model, often used as a scalable alternative to human raters."
用法说明 · 针对中文母语者
常保留连字符写法;也可写作 'LLM-as-a-judge'。在 HCI/AI 评测论文中频繁出现。
真实用例 · 3 条
"We conduct a fine-grained evaluation of three LLMs using both human annotators and LLM-as-judge methods, utilizing rubrics created in consultation with CSOs."
"In parallel, LLM-as-judge methods are employed to enhance efficiency, ensure consistency, and provide cross-validation of human ratings."
"We also analyze the evaluation behavior of humans and LLM-as-judges, finding that automated judges produce extremely stable yet compressed score distributions, while human raters exhibit broader variance and stronger sensitivity to linguistic nuance."