遇见数据集

BioInteract/BioInteract

收藏
Hugging Face2026-05-22 更新2026-05-31 收录
官方服务:

资源简介:

BioInteract是一个用于评估视觉语言模型在语义变化下性能的可扩展基准数据集,包含丰富标注的生物相互作用图像。该数据集是最大的公开可用多模态生物相互作用数据集,专门为AI驱动的生态研究中的视觉和机器学习应用而设计。它包括25.6万张图像,标注有1.54万个独特的生物相互作用知识图谱,这些图谱以三元组形式表示实体间的语义关系——源分类群、相互作用类型、目标分类群,涵盖五个界(动物界、植物界、真菌界、色藻界和未定类)和九种生态标准化相互作用类型。BioInteract的一个关键贡献是能够直接从底层知识图谱生成语义控制的语言变体。通过利用结构化三元组,可以系统地构建意义保留和矛盾查询变体,从而实现对语义相似性的显式控制。这有助于区分正确性与一致性,并在有针对性的语言转换下严格评估模型的鲁棒性。

BioInteract comprising richly annotated images depicting interactions between organisms, or biotic interactions, provides a natural testbed for tasks involving images and unconstrained, free-form natural language, as interacting organisms are discerned from images alone and their relationship can be expressed through multiple linguistic forms. BioInteract, the largest publicly available multimodal dataset of biotic interaction, specifically curated for vision and machine learning application in the context of AI-driven ecological research. BioInteract includes 256K images annotated with 15.4K unique biotic interactions knowledge graphs which represent the semantic relationship between entities as triplets—source taxon, interaction type, target taxon—across five kingdoms (Animalia, Plantae, Fungi, Chromista, and incertae sedis) and nine ecologically standardized interaction types. A key contribution of BioInteract is that it can generate semantically controlled linguistic variations directly from the underlying knowledge graph. By leveraging structured triplets, we can systematically construct both meaning-preserving and contradictory query variants, enabling explicit control over semantic similarity. This allows us to disentangle correctness from consistency and to rigorously evaluate model robustness under targeted linguistic transformations.

提供机构:
BioInteract
二维码
社区交流群
二维码
科研交流群
商业服务