When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation· 当防护栏看起来有效时:LLM 代理商业评估中的构念效度失败
Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B--14B ladder. It also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the paired contrasts to +7.2, -13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% b
LLM代理商业评估中,构念效度失败导致防护措施效果被高估。
- 核心方法
- 通过一个配置酒店交易的多轮买家-卖家测试平台,对比不同防护措施下的市场福利变化,同时控制了提议模式和买家选择过程的一致性,以审计不同规模模型在防护及未防护场景下的表现差异。
- 适合谁读
- 适合研究者、政策制定者和关注LLM商业应用效果评估的工程师阅读。
- 要解决的问题
- 论文指出,在使用语言模型代理进行商业互动模拟时,即使输出的数据看起来合理,但所报告的市场防护措施效果可能缺乏构念效度,从而高估了实际效果。
- 关键实验
- 实验涉及不同规模的Qwen2.5模型(1.5B到14B参数),在控制提议模式和买家选择过程一致性的情况下,重新评估了四种主要防护措施的效果,结果显示多次生成后效果显著降低。
- 主要贡献
- 揭示了LLM代理商业评估中的构念效度问题,并通过实验展示了不同条件下防护措施的真实效果差异,为未来研究提供了重要参考。
- 意义与局限
- 意义:强调了在评估LLM代理商业效果时,确保构念效度的重要性。影响:可能影响当前基于LLM代理的市场研究和政策制定。局限:仅在一个特定的酒店交易测试平台进行了评估,结果的泛化性有待进一步研究。