论文进阶
Benchmarking
基准测试,通过标准化测试、数据集或任务对系统或模型性能进行系统化衡量与比较的评估范式。
English Definition
"A systematic evaluation paradigm that uses standardized tests, datasets, or tasks to measure and compare the performance of systems or models."
用法说明 · 针对中文母语者
在 HCI/UX 研究中常用来指代某种评测方法或评测集;中文常译作 '基准测试' 或 '基准评测'。
真实用例 · 3 条
"Benchmarking has become the dominant paradigm for evaluation, including efforts to assess cultural and domain-specific performance and cross-cultural representation."
论文Pipeline extracted — review needed·段落开头,把 benchmarking 定位为当前主流的评估范式
"Today's benchmarks fall short: they often use translated or artificially created data that overlook real user needs and elevate institutional priorities above those of the communities most affected."
论文Pipeline extracted — review needed·作者批评现有基准脱离真实用户需求
"Our approach enables scalable, automated benchmarking through a culturally aware, community-driven pipeline."
论文Pipeline extracted — review needed·提出新的自动化基准评测方法
由 pipeline 自动采集,待人工 review