Textual Time Series Corpus for Sepsis-3 (T2S2)
收藏资源简介:
该数据集是由卡内基梅隆大学和 NIH 的国家图书馆医学部合作构建的,包含2139份开放获取的病例报告,这些报告来源于Pubmed开放获取子集。数据集通过使用大型语言模型构建,旨在定位临床事件的时间,并生成与时间相关的文本时间序列。该数据集的构建旨在解决临床事件时间序列分析的问题,为更精确的疾病预测和治疗提供支持。
This dataset was collaboratively constructed by Carnegie Mellon University and the National Library of Medicine (NLM) under the National Institutes of Health (NIH). It contains 2,139 open-access case reports sourced from the PubMed Open Access Subset. Developed using large language models (LLMs), this dataset is designed to localize the timestamps of clinical events and generate time-related textual temporal sequences. The construction of this dataset aims to address the challenge of clinical event temporal sequence analysis, so as to support more precise disease prediction and treatment.

- 1Reconstructing Sepsis Trajectories from Clinical Case Reports using LLMs: the Textual Time Series Corpus for Sepsis卡内基梅隆大学 · 2025年



