遇见数据集

BART-large-CNN hyperparameters.

收藏
Figshare2026-02-12 更新2026-04-28 收录
官方服务:

资源简介:

Timely identification of patients who meet clinical trial eligibility criteria is a persistent bottleneck in trial recruitment because the criteria are written in flexible natural language, while hospital EHRs are stored in structured schemas. To bridge this gap, we propose EC2Seq2Sql, an end-to-end, two-stage framework that automatically converts narrative eligibility criteria into executable SQL queries for EHR-based patient screening. In the first stage, a BART-based semantic parser transforms free-text trial criteria into lightweight structured pattern sequences defined over seven common clinical domains. In the second stage, an LLM-based agent, guided by system- and human-designed prompts, grounds these structured patterns to the target database schema and generates syntactically valid and logically coherent SQL statements. We evaluated the framework on the ClinicalTrials.gov eligibility-criteria dataset and further validated it on a de-identified real-world hepatocellular carcinoma EHR cohort from Zhongshan Hospital, Fudan University. The BART parser outperformed representative Seq2Seq baselines, achieving ROUGE_L 0.8067 and BLEU 0.8427, while the SQL generation stage reached an exact-match accuracy of 0.84 and an execution accuracy of 0.91 after SQL normalization. On the real-world cohort, the generated queries achieved a clinical match accuracy of 0.88 after expert review, indicating that the proposed pipeline can retrieve trial-eligible patients from operational EHR data. These results suggest that EC2Seq2Sql can substantially reduce manual screening effort and provide a reproducible path from narrative criteria to database-level cohort identification, although broader multi-center validation and ontology-based normalization will be needed for large-scale deployment.

及时识别符合临床试验入组标准的患者始终是临床试验招募的核心瓶颈,因为入组标准以灵活的自然语言撰写,而医院电子病历(Electronic Health Record, EHR)均存储于结构化数据库模式中。为弥合这一差距,我们提出EC2Seq2Sql——一款端到端的两阶段框架,可自动将叙述性临床试验入组标准转换为基于电子病历的患者筛选所需的可执行SQL查询语句。第一阶段,基于BART的语义解析器将自由文本形式的试验入组标准,转换为针对七个常见临床领域定义的轻量级结构化模式序列。第二阶段,基于大语言模型(Large Language Model, LLM)的AI智能体(AI Agent)在系统与人工设计的提示词引导下,将这些结构化模式映射至目标数据库模式,并生成语法合规、逻辑自洽的SQL语句。我们在ClinicalTrials.gov的临床试验入组标准数据集上对该框架进行了评估,并在复旦大学附属中山医院的去标识化真实肝细胞癌(hepatocellular carcinoma)电子病历队列中完成了进一步验证。BART解析器的性能优于代表性Seq2Seq基线模型,取得了ROUGE_L 0.8067与BLEU 0.8427的评估结果;SQL生成阶段在完成SQL标准化后,精确匹配准确率达0.84,执行准确率达0.91。在真实世界队列中,经专家评审后,生成的查询语句临床匹配准确率达0.88,表明所提出的流程可从运营中的电子病历数据中检索出符合临床试验入组标准的患者。上述结果显示,EC2Seq2Sql可大幅缩减人工筛选的工作量,并为从叙述性标准到数据库级队列识别提供了一条可复现的路径;不过若要实现大规模部署,仍需开展更广泛的多中心验证以及基于本体的标准化工作。

创建时间:
2026-02-12
二维码
社区交流群
二维码
科研交流群
商业服务