遇见数据集

Relation Extraction for Diet, Non-Communicable Disease and Biomarker Associations (RECoDe): A CoDiet study

收藏
Zenodo2026-03-16 更新2026-05-26 收录
官方服务:

资源简介:

This is the data accompanying the submitted conference proceeding to ECCB (pre-print available on bioRxiv). This submission contains 3 JSON lines files used for training (train), optimising (val), and evaluating (test) a multi-entity relation extraction model: train.jsonl val.jsonl test_public.jsonl The train and val datasets contain the following fields (definition): pmc_id (publication ID contained in the redistributable CoDiet corpus) original_text (sentence in which an entity pair has a relation type) original_text_offset (sentence offset in the corpus document) type (one of 10 relation types) Node1_str (CoDiet-Silver corpus annotated entity text) Node1_type (CoDiet-Silver corpus entity type) Node1_offset (Entity offset in original_text) Node1_length (Entity length) Node2_str (CoDiet-Silver corpus annotated entity text) Node2_type (CoDiet-Silver corpus entity type) Node2_offset(Entity offset in original_text) Node2_length (Entity length) Where the type is the relation between Node1_str and Node2_str. The test_public dataset does not contain the type field (as it is used for evaluation), but does include: rel_id (unique relation number in test set used for evaluation purposes) The test labels are kept separate and will be hosted on Codabench (see GitHub for updates) to allow benchmarking without test set leakage into models.

提供机构:
Zenodo
创建时间:
2026-03-16
二维码
社区交流群
二维码
科研交流群
商业服务