A2H_Preclinical_Translation_Data
收藏资源简介:
🧾 Description This dataset accompanies a computational framework for large-scale analysis of animal-to-human translation in drug development, with a focus on the neuroscience domain. It integrates data extracted from biomedical literature with clinical trial and regulatory information using natural language processing (NLP) methods. The dataset is organized as a multi-stage processing pipeline, including raw PubMed queries, filtered animal studies, named entity recognition (NER) outputs, entity normalization, and final translation analyses linking preclinical evidence to clinical outcomes. Key components include: PubMed data collection (01_pubmed_query_neuro/):Raw literature retrieved from PubMed in multiple query rounds, along with query definitions. Animal study classification (02_animal_study_classification/):Filtered corpus of animal studies, including metadata for over 6 million records. Information extraction (NER) (03_IE_ner/, 08_IE_full_text/):NLP-based extraction of drug and disease entities from abstracts and full texts, including model predictions and rule-based (regex) outputs. Entity normalization (04_normalization/):Harmonized drug and disease entities to reduce terminology heterogeneity across studies. Translation analysis (10_drug_disease_translation_analysis/):Structured datasets linking animal studies to clinical trials and regulatory approvals, including: drug–disease pair mappings translation tables association metrics for translational success These resources enable systematic investigation of how findings from animal studies translate into clinical development. 🎯 Potential uses This dataset can be used to: Study large-scale patterns of animal-to-human translation Analyze experimental design factors associated with translational success Support automated evidence synthesis in biomedical research Develop and evaluate NLP methods for biomedical information extraction Investigate trends in preclinical neuroscience research



