ai.hackcv
论文精选 65arXiv

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models· 策略基准:评估大模型的显式策略归纳

As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstracting task rules, making it sensitive to example construction. In contrast, human learners often reduce such sensitivity by first summarizing task rules from examples and then applying them to new instances. To evaluate this ability, we propose StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility. We further analyze strategy induction from three perspectives: task variation, model configuration, and a

AI 解读论文

评估大语言模型的显式策略归纳能力

核心方法
构建StrategyBench基准,从BIG-Bench中选择可归纳策略的任务,构造参考策略,定义策略质量和下游效用的评估指标
适合谁读
研究者
要解决的问题
大语言模型在少样本场景中直接学习任务规则时敏感性高,需要评估其归纳策略的能力
关键实验
分析了任务变体、模型配置对策略归纳的影响,具体实验结果未提供
主要贡献
提供了评估大模型显式策略归纳能力的新方法和基准,有助于理解模型在少样本学习中的表现
意义与局限
有助于优化大模型的少样本学习能力,促进其在新任务中的应用;局限在于当前实验覆盖范围有限
领域:cs.AI作者:Jinghan Tan、Yuanzheng Wang、Lu Chen
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考