遇见数据集

Compact Vision-Language Models for Cross-Crop Plant-Disease Diagnosis at the Edge: A CPU-Only Study

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

Cloud vision–language models diagnose plant disease well but bill per query and need connectivity, anda conventional classifier has no output unit for an unseen crop. We ask what lets a small, frozen vision–language model diagnose crops it was never trained on. On the SAGE dataset, four frozen compactcontrastive encoders of 11.4–86.3 M parameters match leaf images to written disease descriptions, andwe compare seven ways of authoring those descriptions across nested held-out sets of 16, 34 and 51unseen classes. First, authoring outweighs the encoder upgrade: the best descriptions gain 10.3–16.2points over a bare class name, against 4.0–5.1 points between the smallest and largest encoder, and at51 classes it leads on both label sets and under either measure of the encoder effect. Second, per-classsentences describing each disease as it appears in a photograph reach 30.9% top-1 on 51 unseen classes,7.4 times the majority-class prior, and beat source-grounded descriptions on all four encoders, by 7.0points on average. Third, requiring a citable source costs no measurable accuracy, re-embedding thesource-grounded text as a sentence ensemble recovers only 4% of its deficit, and stripping its pathogenand taxonomy fields recovers none of it, so the gap lies in the text rather than in citation, embedding orthe schema’s non-visual fields. The smallest encoder runs at 17.4 ms per image on a laptop CPU in45.8 MB. For CPU-only deployment, effort spent on descriptions returns more than effort spent on parameters.

提供机构:
Zenodo
创建时间:
2026-09-23
二维码
社区交流群
二维码
科研交流群
商业服务