PMOA-TTS
收藏资源简介:
PMOA-TTS数据集由卡内基梅隆大学机器学习系、信息系統与公共政策学院、美国国立卫生研究院国家医学图书馆的研究人员创建,包含124,699份来自PubMed Open Access的病例报告,每份报告都通过可扩展的基于大型语言模型(LLM)的管道转换为结构化的(事件,时间)时间序列。该数据集通过启发式过滤和Llama 3.3识别单个患者的病例报告,并使用Llama 3.3和DeepSeek R1进行提示驱动提取,最终生成了超过560万个带时间戳的临床事件。该数据集在临床和人口统计覆盖范围广泛,并在下游生存预测任务中表现出色,嵌入从提取的时间序列中获得的预测性能可达0.82 ± 0.01。PMOA–TTS为时间线提取、时间推理和纵向建模提供了可扩展的基础,可用于生物医学自然语言处理。数据集可在Hugging Face平台上获取。
PMOA-TTS was developed by researchers from the Department of Machine Learning, Heinz College of Information Systems and Public Policy, Carnegie Mellon University, and the National Library of Medicine, National Institutes of Health. It contains 124,699 case reports sourced from PubMed Open Access. Each report was converted into a structured (event, time) time series via a scalable large language model (LLM)-based pipeline. This dataset identifies single-patient case reports through heuristic filtering and Llama 3.3, and uses Llama 3.3 and DeepSeek R1 for prompt-driven extraction, ultimately generating over 5.6 million timestamped clinical events. The dataset features broad clinical and demographic coverage, and performs excellently on downstream survival prediction tasks: the predictive performance derived from embeddings of the extracted time series reaches 0.82 ± 0.01. PMOA-TTS provides a scalable foundation for timeline extraction, temporal reasoning, and longitudinal modeling for biomedical natural language processing, and is available on the Hugging Face platform.

- 1PMOA-TTS: Introducing the PubMed Open Access Textual Times Series Corpus卡内基梅隆大学机器学习系、信息系統与公共政策学院、美国国立卫生研究院国家医学图书馆 · 2025年



