ai.hackcv
论文精选 61arXiv

A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol· 教学反馈分类协议的耐用性与跨语言迁移基准

Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented annotation guide, an intra-annotator reliability measurement, stratified cross-validation, and a held-out evaluation on a Spanish institutional corpus with a frozen-encoder design. Two questions limit its reuse: whether a protocol fixed to 2019-era frozen embeddings stays competitive as representation methods advance, and whether it transfers to a second language. We re-run it on the original Spanish data across three representation generations, sparse lexical features, frozen transformer embeddings, and prompted large language models, and transfer its sentiment task to English with a balanced 45,000-comment corpus checked against an aspect-labeled education dataset. Treating paired comparisons as descriptive, we find the protocol durable: a 2026 frontier model posts the highest thematic F1 on the hardest Spanish task, yet shows no sentiment advantage over a cheap model and no descriptive separation from it on English, so model choice is a deployment decision, not a property of the method.

AI 解读论文

教学反馈分类协议的跨时代和跨语言性能评估

核心方法
使用三个不同表示方法(稀疏词汇特征、冻结的变换器嵌入、大型语言模型提示)重新运行分类协议,并将其情感任务迁移到英语数据集
适合谁读
研究者、工程师
要解决的问题
评估教学反馈分类协议是否在新的表示方法和不同语言中仍然有效
关键实验
在原始西班牙语数据上测试了三种表示方法,并在45,000条评论的英语数据集上进行了情感任务迁移测试
主要贡献
证明了该协议在新模型和跨语言任务中的耐用性,为实际部署提供了灵活性
意义与局限
为教学反馈分类提供了可靠的基准和方法,影响未来的相关研究和应用;但不同模型在情感任务上的表现差异不大,存在局限性
领域:cs.CL作者:Esteban U. Vega Barajas
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考