Can Edge-Deployable Vision-Language Models Identify Species?· 边缘部署的视觉-语言模型能否识别物种?
Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relevant class to evaluate for species identification. We test whether models in this deployment-relevant 2--8B range carry genuine taxonomic knowledge, evaluating four such VLMs (Qwen3-VL 2B/4B/8B, Gemma3 4B) against the domain-specific specialist BioCLIP (300M parameters) on a 96-species task, comparing clean iNaturalist photographs against camera-trap imagery from 6 LILA.science collections, on two independently-sampled evaluation sets. All models identify species far above chance, but every model -- general-purpose or specialist -- degrades sharply on field imagery (domain gaps of 9.6--26.6 per
研究边缘硬件上部署的小型视觉-语言模型能否用于物种识别。
- 核心方法
- 在96个物种的任务中,评估四个2-8B参数范围的VLMs(Qwen3-VL 2B/4B/8B, Gemma3 4B)和一个300M参数的专业模型BioCLIP,比较干净的iNaturalist照片和来自6个LILA.science集合的野外影像。
- 适合谁读
- 研究者、工程师
- 要解决的问题
- 相机陷阱在野外使用时,由于边缘硬件的限制,需要评估小型、本地可部署的视觉-语言模型是否能有效识别物种。
- 关键实验
- 使用两个独立采样的评估集,比较了不同模型在干净iNaturalist照片和野外影像上的识别性能。
- 主要贡献
- 发现所有模型在物种识别上远超随机水平,但在野外影像上的表现显著下降,揭示了模型在不同数据域上的泛化能力问题。
- 意义与局限
- 该研究影响了边缘设备上的物种识别应用,强调了模型在实际部署中需要克服的数据域差距问题。模型性能的下降表明现有的VLMs在野外应用中的局限性。