ai.hackcv
论文精选 60arXiv

The Effects of Synthetic Data and Label Distribution on Canola Branch Counting

Collecting annotated plant images for automated phenotyping is often slow and expensive. Plant models simulating growth and development can generate unlimited synthetic images with exact labels. However, previous work has established that whether incorporating synthetic data improves performance depends on the ratio of synthetic to real images and the label distribution of the synthetic dataset. To systematically quantify both factors, we train ResNet-18 models on a canola branch-counting task using a calibrated L-system plant model. We vary each factor independently. Synthetic-to-real ratios of 1:5 to 1:22 broadly improve performance; the best ratio (1:7) reduces mean absolute difference by 7.6% over real-only training. For label distribution, a uniform synthetic distribution is strongly suboptimal (abs. diff. of approximately 1.70); interpolating 90% toward the real distribution yields abs. diff. 0.927, whereas Gaussian smoothing of the real label distribution yields the best overall result (abs. diff. 0.912, a 14.7% improvement over real-only). A minimum of 10 synthetic images per label offers a simpler alternative with modest gains, while 100 per label over-corrects and hurts performance.

AI 解读论文

合成数据与标签分布对油菜枝条计数的影响研究

核心方法
使用校准的L-系统植物模型生成合成图像,训练ResNet-18模型进行油菜枝条计数,独立改变合成与真实图像比例及标签分布
适合谁读
研究者、工程师
要解决的问题
解决在植物自动表型中,收集注释图像速度慢且成本高的问题
关键实验
测试了不同合成与真实图像比例(1:5至1:22)及标签分布(均匀、插值、高斯平滑)对模型性能的影响
主要贡献
系统量化了合成数据与真实数据比例及标签分布对模型性能的影响,找到了最佳的标签分布调整方法
意义与局限
本研究为植物图像自动表型中合成数据的有效利用提供了指导,但其应用范围可能受限于特定植物模型和任务
领域:cs.CV作者:Amirsalar Darvishpour、Mikolaj Cieslak、Adam Runions
相关推荐
论文精选 60已解读

Principia: Relational Physics Tests for Video Models· Principia:视频模型的物理测试

Evaluating physical reasoning in video models is difficult because absolute motion measure…

领域:cs.CV作者:Varun Varma Thozhiyoor、Shivam Tripathi、Venkatesh Babu Radhakrishnan
📎 arXiv🕒 09-04 01:59🔗 arxiv.org

本站内容由 LLM 精选聚合,原文版权归 arXiv 所有 · 摘录仅供参考