Viet Medi Species 2026
收藏资源简介:
🌿 Viet Medi Species 2026 A Multilingual Dataset for Vietnamese Medicinal Biodiversity 📌 Overview Viet Medi Species 2026 is a large-scale, multilingual image dataset designed to support research in biodiversity informatics, medicinal plant identification, and AI-based classification. The dataset integrates Vietnamese traditional medicinal knowledge with global taxonomic standards, addressing the lack of Vietnamese-language resources in existing biodiversity platforms. The dataset contains 310,647 images across 4,799 species spanning four biological kingdoms: Plantae, Fungi, Chromista, and Bacteria. Each species is linked to the Global Biodiversity Information Facility (GBIF) taxonomy and enriched with curated Vietnamese vernacular names. 📊 Dataset Statistics Total species: 4,799 Total images: 310,647 Kingdoms covered: 4 (Plantae, Fungi, Chromista, Bacteria) Species with Vietnamese names: 4,031 (~84%) Taxonomic coverage: 355 families 1,896 genera Image Distribution The dataset follows a long-tail distribution typical of real-world biodiversity data: Average: 87 images/species Median: 35 images/species Range: 1 to 130 images per species 🧬 Data Composition Kingdom Species Images Vietnamese Name Coverage Plantae 4,667 297,320 85% Fungi 120 12,675 60% Chromista 9 476 67% Bacteria 3 176 100% This distribution reflects real-world biodiversity observations, where plant species dominate available data. 🏷️ Metadata Each species is associated with rich taxonomic and linguistic metadata following Darwin Core standards, including: speciesKey (GBIF identifier) scientificName canonicalName vietnameseName (semicolon-separated) kingdom, phylum, class, order, family, genus taxonomicStatus Vietnamese names are encoded in UTF-8 and may include multiple regional variations. 🛠️ Data Collection & Curation The dataset was built using a fully reproducible pipeline consisting of: 1. Taxonomic Normalization Species were sourced from the Vietnamese Medicinal Plant Catalogue and matched to GBIF using the Species Match API to ensure global taxonomic consistency. 2. Image Acquisition Images were downloaded from GBIF using: Only occurrences with images Up to 130 images per species Original resolution (no thumbnails) An asynchronous pipeline enabled efficient large-scale downloading. 3. Vietnamese Name Curation Vietnamese names were manually curated through 320+ hours of validation using multiple authoritative sources. Quality criteria included: Cross-verification from at least two sources Linguistic correctness Evidence of medicinal or botanical use 4. Metadata Integration All taxonomic and linguistic data were merged into a unified bilingual dataset. (See pipeline diagram on page 3 of the paper for workflow overview.) 🌏 Key Features 🌿 Largest Vietnamese medicinal plant dataset to date 🌐 Multilingual (Vietnamese + scientific taxonomy) 🔗 Integrated with GBIF global infrastructure 📈 Real-world long-tail distribution ♻️ Fully reproducible pipeline and open-source code 🚀 Use Cases This dataset enables a wide range of applications: 🌱 Medicinal plant identification (mobile/web apps) 🤖 Deep learning for species classification 📚 Vietnamese-language biodiversity education 🌍 Citizen science and conservation projects 🔍 Cross-lingual biodiversity research ⚠️ Limitations Long-tail imbalance across species Limited representation for minority kingdoms (e.g., Bacteria) Some species (16%) lack Vietnamese names Dependent on GBIF data availability and geographic bias License This dataset contains two kinds of material, licensed separately. 1. Curated metadata and annotations: CC BY 4.0 The CSV metadata files, species list, taxonomic normalisation, Vietnamese vernacular names, class list, and data splits were created by the authors and are released under the Creative Commons Attribution 4.0 International licence (https://creativecommons.org/licenses/by/4.0/). You may share and adapt them for any purpose, including commercially, provided you give appropriate credit. 2. Images: original GBIF licences apply All images were retrieved from occurrence records published through the Global Biodiversity Information Facility (GBIF). Each image remains under the licence chosen by its original publisher, typically CC0 1.0, CC BY 4.0, or CC BY-NC 4.0. The authors do not own these images and do not relicense them. Users are responsible for complying with the licence of each image, including the non-commercial restriction where it applies. The licence, publisher, and GBIF occurrence ID for each image can be looked up through the GBIF record linked in the metadata. Source records: GBIF.org (2026). 📜 Citation If you use this dataset, please cite: Tran, T. P., Din, F. U., Brankovic, L., Sanin, C., and Hester, S. M., Viet Medi Species 2026: A Web-Accessible Multilingual Dataset for Vietnamese Medicinal Biodiversity., In Companion Proceedings of the ACM Web Conference 2026 (WWW Companion ’26), June 29 - July 03, 2026, Dubai, United Arab Emirates. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3774905.3795608 Tran, T. P., Ud Din, F., Brankovic, L., Sanin, C., & Hester, S. M. (2026). Viet Medi Species 2026 [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.22928927



