SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control· SIRF:工业内容风险控制的规范内化风险基础模型
For industrial content risk control, the real deployment constraint is not average accuracy but how much risk can be auto-handled under high precision and second-level latency. We present SIRF (Spec-Internalized Risk Foundation Model), which internalizes a platform's complex policies, synthesized without additional human annotation via EntiGraph, MAGA rewriting and account-level chain-of-thought (CoT), into the weights via continued pretraining (CPT), so rules are applied at high precision under an ultra-low-latency, verdict-only deployment. A controlled same-source comparison (Qwen3-8B-SFT vs. SIRF-8B-SFT, identical policy injection and verdict-only output form, differing only in policy-grounded CPT) attributes the gain to internalization: SIRF-8B-SFT reaches 71.3% Black Recall@P95, +15.1
规范内化的风险基础模型提升工业内容风险控制的自动处理能力。
- 核心方法
- SIRF模型通过持续预训练将平台复杂政策内化到模型权重中,利用EntiGraph、MAGA重写和账户级思考链实现政策合成。
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 工业内容风险控制中的实际部署约束是高精度和低延迟下的自动处理能力,而非平均准确性。
- 关键实验
- 通过Qwen3-8B-SFT与SIRF-8B-SFT的对比实验,验证了政策内化的有效性。
- 主要贡献
- 在相同的源和部署方式下,SIRF-8B-SFT在P95黑色召回率上提升15.1%,达到71.3%。
- 意义与局限
- SIRF模型提高了内容风险控制的效率和精度,适用于大规模工业应用,但可能需要针对不同平台定制化调整。