ai.hackcv
论文精选 65arXiv

Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades· 廉价验证者,大盲点

Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natural extension closes the loop: fine-tune the cheap student on the verifier's rejections so the escalation rate, and cost, fall each round. We measure this loop on real LLMs and report four findings. First, the verifier's blind spot, the fraction of the student's wrong answers it accepts, is large and moves adversarially: it grows with student capability ($β$ from 0.12 to 0.55 as the student scales 0.5B to 32B) and shrinks with verifier capability, so it is worst in the cheap-student, cheap-verifier regime cascades exist to create. Second, buying it away returns the saving: a frontier verifier drives $β$ to about 0.05 but then escalates on 46%

AI 解读论文

研究了通过 cheap student 和 verifier 之间的闭合循环降低成本的可靠性和代价。

核心方法
通过将廉价模型的错误答案反馈给前沿验证者,不断优化廉价模型,减少升级到验证者的查询率。
适合谁读
研究者
要解决的问题
如何在保持可靠性的前提下,通过使用廉价模型和前沿验证者之间的级联来降低成本。
关键实验
使用真实的大规模语言模型进行实验,测量了廉价模型和验证者之间的盲点变化情况。
主要贡献
报告了廉价验证者存在较大盲点,且随着模型规模的变化而变化,提出了通过增强验证者能力来降低盲点的方法。
意义与局限
证明了级联模型在降低成本的同时,可靠性会受到严重影响,为模型优化提供了新方向。但盲点问题在廉价模型和廉价验证者组合中尤为突出。
领域:cs.AI作者:Dushyant Rajput
相关推荐

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考