遇见数据集

A2H_Preclinical_Translation_Data

收藏
Zenodo2026-04-02 更新2026-05-26 收录
官方服务:

资源简介:

🧾 Description This dataset accompanies a computational framework for large-scale analysis of animal-to-human translation in drug development, with a focus on the neuroscience domain. It integrates data extracted from biomedical literature with clinical trial and regulatory information using natural language processing (NLP) methods. The dataset is organized as a multi-stage processing pipeline, including raw PubMed queries, filtered animal studies, named entity recognition (NER) outputs, entity normalization, and final translation analyses linking preclinical evidence to clinical outcomes. Key components include: PubMed full Neuro data (01_pubmed_query_neuro/): This directory contains raw literature records retrieved from PubMed across multiple query rounds. data/pubmed_queries/union_all_queries_pmids_21704840.txt: the PMIDs of all articles retrieved from the neuropsychiatric queries. data/full_pubmed_raw: Raw literature retrieved from PubMed in multiple query rounds. Each file consists of line-based entries, where each line represents one publication record. Fields are separated using a triple-pipe delimiter. The records follow the schema: <PMID> ||| <Year> ||| <Journal> ||| <Title> ||| <Abstract> ||| <Publication Types> Animal study classification (02_animal_study_classification/): This directory contains data and model outputs for identifying and curating animal studies from PubMed. data/animal_studies/ Filtered PubMed metadata for animal studies. model_predictions/ Cleaned and aggregated prediction outputs for animal study classification, including intermediate and final datasets. pubmed_results/, pubmed_results_r2/, pubmed_results_r3/...Results from multiple PubMed query rounds. Later rounds (e.g., `r3`) include chunked prediction outputs. Information extraction (NER) (03_IE_ner/): This directory contains datasets and model outputs for named entity recognition (NER) on animal studies. data/animal_studies_with_drug_disease/ Metadata for animal studies enriched with drug and disease information. data/finetuning_ner/ Datasets used for NER model training and evaluation, including: animals_nr/, drug_disease/, strain/: domain-specific annotated datasets (JSON format) k_fold_splits/: cross-validation splits for each dataset model_predictions/ Model outputs for NER tasks, organized by entity type (e.g., drug–disease), stored as chunked CSV files. Entity normalization (04_normalization/): This directory contains normalized entity mappings for diseases and drugs extracted from animal studies. data/mapped_all/entities_drug_disease_preclin.csv Combined dataset of extracted drug and disease entities with normalization, including:- PMID: document identifier - disease_*: disease mentions mapped to MONDO terms and IDs - drug_*: intervention mentions mapped to UMLS terms and IDs - merged_*: consolidated normalized labels and identifiers - nearest_dataset_parent_*: hierarchy-based parent mappings (if available) Model predictions (08_IE_full_text/model_predictions/) This directory contains document-level predictions and metadata extracted from full-text articles. animals_nr/Extracted total counts of animals and converted to numeric where possible. regex/Rule-based (regex) document-level predictions for key reporting features, including:assay, blinding, randomization, sample, sample size, sex, species, and welfare. strain/Extracted and normalized animal strain information. author_affiliations_pubmed_mapped.csv Maps PubMed records to author affiliations (e.g., institutions, countries). full_text_combined_all_annotations_metadata.csvAggregated dataset combining full-text annotations, predictions, and metadata across pipelines. Translation analysis (10_drug_disease_translation_analysis/): out/ Final translation association datasets, including full and subset-specific outputs (e.g., NEURO). interm_drug_disease_pairs/ Intermediate processed drug–disease pairs, including: aggregated (agg) and flat annotation formats translation tables linking drug–disease pairs across datasets Key Files translation_table_drug_disease_*.csv: Core mapping table linking normalized drug–disease pairs across datasets. drug_disease_pairs_annotations_*.csv: Annotated drug–disease pairs in both flattened (each PMID a seaparate row) formats, and aggregated. Aggregated preclinical evidence per drug–disease pair includes:- animal characteristics (sex, species, strain)- assay types- country of first author- supporting PMIDs- counts of articles and reporting quality indicators (blinding, randomization, welfare) df_translation_associations_*.csv: Feature table for modeling (e.g., logistic regression), where each row represents a drug–disease pair. Includes aggregated features from preclinical studies, such as:- experimental design signals (blinding, randomization, welfare)- diversity metrics (species, strain, assay, country)- article counts - target: outcome variable for translation analysis 🎯 Potential uses This dataset can be used to: Study large-scale patterns of animal-to-human translation Analyze experimental design factors associated with translational success Support automated evidence synthesis in biomedical research Develop and evaluate NLP methods for biomedical information extraction Investigate trends in preclinical neuroscience research

提供机构:
Zenodo
创建时间:
2026-04-02
二维码
社区交流群
二维码
科研交流群
商业服务