ai.hackcv
论文精选 85arXiv

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models· FriendBench:人类和多模态大模型的熟悉度推断基准测试

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward "stranger"---a difference in effective prior, not discrimination. Richer channels help both unequally, and on

领域:cs.CL作者:Jeffrey M. Girard、Jason Z. Zheng、Jacqueline R. Vertino
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考