论文进阶
red teaming
红队测试 / 红队演练:在 AI 安全语境下,由专人或团队以对抗性方式向 AI 系统输入有害、边缘或操纵性内容,以发现其安全漏洞与潜在风险。
English Definition
"Adversarial testing practice in which a team deliberately probes an AI system with harmful, edge-case, or manipulative inputs to surface safety vulnerabilities before deployment."
用法说明 · 针对中文母语者
源自军事/安全领域的「红蓝对抗」概念,近几年被广泛用于 AI safety 语境。中文常直译为「红队测试」或「红队攻防」。
真实用例 · 3 条
"Responsible AI (RAI) content work, such as annotation, moderation, or red teaming for AI safety, has become essential"
论文Pipeline extracted — review needed·介绍 RAI 内容工作的具体类型之一
"Their tasks – collectively described as Responsible AI (RAI) content work [70], including annotation, moderation, and red teaming for model safety"
论文Pipeline extracted — review needed·列举 RAI 工作的三类核心任务
"AI companies then leverage to scale annotation, moderation, and adversarial testing"
论文Pipeline extracted — review needed·说明平台规模化运作红队/对抗性测试等任务
由 pipeline 自动采集,待人工 review