gmanolache/BioInteract
收藏资源简介:
BioInteract是一个用于评估视觉语言模型在语义变化下的可扩展基准数据集。它包含丰富注释的图像,描绘了生物体之间的相互作用(生物相互作用),为涉及图像和自由形式自然语言的任务提供了自然测试平台。具体来说,BioInteract是最大的公开可用的生物相互作用多模态数据集,专门为AI驱动的生态研究中的视觉和机器学习应用而策划。数据集包括256,000张图像,注释有15,400个独特的生物相互作用知识图谱,这些知识图谱将实体之间的语义关系表示为三元组(源分类群、相互作用类型、目标分类群),涵盖五个界(动物界、植物界、真菌界、色藻界和incertae sedis)和九种生态标准化的相互作用类型。关键贡献在于它可以直接从底层知识图谱生成语义受控的语言变体,通过利用结构化三元组,系统地构建意义保持和矛盾的查询变体,从而实现对语义相似性的显式控制,允许区分正确性和一致性,并在有针对性的语言转换下严格评估模型鲁棒性。
BioInteract is a scalable benchmark for evaluating vision-language models under semantic variation. It comprises richly annotated images depicting interactions between organisms, or biotic interactions, providing a natural testbed for tasks involving images and unconstrained, free-form natural language. Specifically, BioInteract is the largest publicly available multimodal dataset of biotic interaction, curated for vision and machine learning applications in the context of AI-driven ecological research. It includes 256K images annotated with 15.4K unique biotic interactions knowledge graphs, which represent the semantic relationship between entities as triplets—source taxon, interaction type, target taxon—across five kingdoms (Animalia, Plantae, Fungi, Chromista, and incertae sedis) and nine ecologically standardized interaction types. A key contribution is that it can generate semantically controlled linguistic variations directly from the underlying knowledge graph, enabling systematic construction of both meaning-preserving and contradictory query variants for explicit control over semantic similarity, allowing to disentangle correctness from consistency and rigorously evaluate model robustness under targeted linguistic transformations.



