遇见数据集

RECODE: Relational Ecological COrpus for Data Extraction

收藏
Zenodo2025-10-28 更新2026-05-26 收录
官方服务:

资源简介:

RECODE is a manually annotated corpus of ecological and taxonomic literature, aimed at training and fine-tuning models for automated extraction of occurrence and trait data from unstructured text. Documents present have been annotated and validated by experts familiar with the traits of the included taxa (currently spiders and insects). Furthermore, this dataset is goal-oriented and published simultaneously with a complementary R package (arete) and a practical case of its usage for both finetuning and validation. Table of Contents The following are all the elements contained in the archive ./recode.zip: recode, directory. Contains metadata.csv, containing the metadata for each entry in RECODE. Also subdivides into all available taxa: currently ./insecta and ./araneae. These directories subdivide by annotator if available. If not, .tsv files are placed under a ./all directory. Finally, annotation files are named by the focus taxa they belong to and their document ID. metadata.csv ./araneae ./insecta r, directory. Contains the script used to calculate the values and plots present in the describing manuscript as well as all necessary files. These include two .csv tables containing taxonomic information on the species annotated in the dataset. These are supplied instead of being generated as part of the script as: 1) this data is prone to becoming outdated; 2) the method used to extract this information requires internet access. script_publish.R taxa_table_insects.csv taxa_table_spiders.csv plots, directory. Contains all plots created in R that appear on the release version of the describing manuscript. FIG_1.png FIG_3.png FIG_4.png FIG_5.png FIG_6.png

提供机构:
Zenodo
创建时间:
2025-10-28
二维码
社区交流群
二维码
科研交流群
商业服务