遇见数据集

CoDiet corpus

收藏
Zenodo2025-11-18 更新2026-05-26 收录
官方服务:

资源简介:

The CoDiet corpora created as part of a multi-institution collaboration to annotate biomedical literature. The repository contains 5 datasets (zipped) in BioC-JSON formats: CoDiet-Gold-public, 450 full-text documents independently annotated by 2 individuals and any disagreements adjudicated by a third independent person) CoDiet-Gold-private, 50 full-text documents unannotated to be used for benchmarking on Codabench (see link on GitHub) CoDiet-Electrum, 4,445 full-text documents machine-annotated (with post-processing rules) using the Gold-public labels/vocabulary only CoDiet-Silver, 4,445 full-text documents machine-annotated (with post-processing rules) using existing deep learning, rule-based, and dictionary-based methods CoDiet-Bronze, 4,445 full-text documents machine-annotated (without post-processing rules) using existing deep learning, rule-based, and dictionary-based methods If you use this resource, please cite the related works below.

提供机构:
Zenodo
创建时间:
2025-11-18
二维码
社区交流群
二维码
科研交流群
商业服务