ai.hackcv
论文精选 65arXiv

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors· AI 语音评估系统的偏见分析

Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attributes (`concepts') in feature-based graders, we extend CAV-based analysis to two neural speaking assessment systems: a text-based BERT grader and a speech-and-text multimodal grader based on Whisper. CAVs repre

AI 解读论文

使用概念激活向量分析AI语音评估系统的偏见问题

核心方法
基于Concept Activation Vectors (CAVs) 技术,扩展了对两个神经网络语音评估系统的偏见检测,包括基于文本的BERT评估系统和基于语音与文本的Whisper多模态评估系统
适合谁读
研究者、工程师、教育工作者
要解决的问题
自动语音评估系统在高风险场景中使用时,评估分数可能受到无关演讲者属性(如母语或年龄)的影响,而非真正反映口语水平
关键实验
对比分析了BERT和Whisper系统对不同L1和年龄的L2学习者口语评估中的偏见情况
主要贡献
首次将CAV技术应用于神经网络语音评估系统的偏见分析,提高了对这些系统公平性和解释性的理解
意义与局限
此研究有助于确保AI语音评估系统的公正性,减少因无关因素导致的评估偏差,但考虑到实际应用中的复杂性,进一步验证和优化仍需进行
领域:cs.AI作者:Arya Labroo、Mengjie Qian、Kate Knill
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考