DART (Drug Annotation from Regulatory Texts)
收藏资源简介:
DART数据集是由意大利药品管理局(AIFA)的官方数据构建的,它是一个包含超过16029个意大利药品特性摘要的结构化语料库,这些摘要来源于意大利药品管理局的官方存储库。该数据集提供了关于药品的关键药理学领域,如适应症、不良反应和药物相互作用的结构化信息。DART数据集是通过一个可复现的流程构建的,包括从网络规模文档检索、监管部分的语义分割,以及使用低温度解码的少量样本调整的大型语言模型进行临床摘要。该数据集为意大利临床自然语言处理社区和更广泛的健康数据科学生态系统提供了一个宝贵的资源,支持大规模语言模型在监管和临床环境中的训练、评估和部署。
The DART dataset is constructed using official data from the Italian Medicines Agency (AIFA). It is a structured corpus containing over 16,029 Italian drug characteristic summaries sourced from the official repository of the Italian Medicines Agency. This dataset provides structured information on key pharmacological domains of medicinal products, including indications, adverse reactions, and drug-drug interactions. The DART dataset is built via a reproducible workflow, which encompasses web-scale document retrieval, semantic segmentation of regulatory sections, and clinical summarization powered by few-shot fine-tuned large language models with low-temperature decoding. This dataset serves as a valuable resource for the Italian clinical natural language processing community and the broader health data science ecosystem, supporting the training, evaluation, and deployment of large language models in regulatory and clinical settings.

- 1通过意大利那不勒斯费德里科二世大学电子工程与信息技术系(DIETI),意大利那不勒斯,意大利 · 2025年



