遇见数据集

Billy6310/guidelines

收藏
Hugging Face2026-03-18 更新2026-03-29 收录
官方服务:

资源简介:

--- license: other license_name: common-crawl license_link: LICENSE task_categories: - text-generation language: - en pretty_name: Clinical Guidelines size_categories: - 10K<n<100K configs: - config_name: default data_files: - split: train path: open_guidelines.jsonl tags: - medical - health dataset_info: features: - name: id dtype: string - name: source dtype: string - name: title dtype: string - name: clean_text dtype: string - name: raw_text dtype: string - name: url dtype: string - name: overview dtype: string --- ### 🎉 **NEW DROP** 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! # Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the [Meditron](https://huggingface.co/epfl-llm/meditron-70b) Large Language Model (LLM). We publicly release a subset of 37K articles from our Guidelines corpus, extracted from 9 of 17 sources that allow content redistribution, namely CCO, CDC, CMA, ICRC, NICE, PubMed, SPOR, WHO and WikiDoc. You can scrape and clean all 17 guideline sources using our code in [epfLLM/meditron](https://github.com/epfLLM/meditron). <img width=75% src="sources.png" alt="Sources of Clinical Practice Guidelines" title="CPG sources"> ## Dataset Details <!-- Provide a longer summary of what this dataset is. --> - **Curated by:** [EPFL LLM Team](https://huggingface.co/epfl-llm) - **Language(s):** English only - **License:** [Common Crawl Foundation Terms of Use](https://commoncrawl.org/terms-of-use) - **Repository:** [epfLLM/meditron](https://github.com/epfLLM/meditron) - **Paper:** *[MediTron-70B: Scaling Medical Pretraining for Large Language Models](https://arxiv.org/abs/2311.16079)* - **Knowledge Cutoff**: August 2023 ## Dataset Creation ### Curation Rationale <!-- Motivation for the creation of this dataset. --> The dataset was curated to provide a high-quality collection of clinical practice guidelines (CPGs) for the medical training of LLMs. Our Clinical Guidelines corpus comprises 48,096 articles from 17 globally recognized sources for clinician and patient-directed guidance across high and low-resource settings, multiple medical domains (internal medicine, pediatrics, oncology, infectious disease, etc.) and multiple geographical locations. ### Source Data <!-- This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...). --> Clinical practice guidelines are rigorously researched frameworks designed to guide healthcare practitioners and patients in making evidence-based decisions regarding diagnosis, treatment, and management. They are compiled through a systematic process of collaborative consensus between experts to establish recommendations from the latest evidence on best practices that would maximize benefit in light of practical concerns such as available resources and context. As a super-synthesis of meta-analyses, they sit atop the *evidence pyramid* and form the basis of actionable evidence-based practice. Clinical guidelines differ based on several factors: - **Organizational level**: CPGs are produced at various organizational granularities, ranging from global to hospital-level initiatives directed by international professional medical associations to informal consortia, regional or national governmental bodies to individual NGOs and hospitals. - **Geographic scope**: The geographic scope ranges from global (WHO) to national (CDC, NICE) and regional (Ontario, Melbourne) to institutional (ICRC, Mayo Clinic). This corpus is biased towards English-speaking regions due to its exclusive focus on English content. - **Resource level**: The corpus also represents health care concerns from high- (Ontario, Melbourne), low- (WHO), and volatile- (ICRC) resource settings. - **Audience level**: Guidelines also contains a range of technical and conversational vocabulary with target audiences of clinicians or patients (or both), and is sometimes highly specialized within a theme (cancer, pediatrics, infectious disease). - **Peer-review**: The peer review processes also ranged from UN bodies (WHO), institutional review boards (ICRC), professional associations (AAFP) to publicly crowdsourced knowledge bases (WikiDoc). - **Document size**: Article length varies widely from very short statements to 100+ page guides. #### Who are the source data producers? <!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. --> The dataset is sourced from 17 globally recognized medical entities, covering a wide range of healthcare contexts and audiences. We employed pragmatic selection criteria over medical sources, seeking CPGs that were: - (1) open-access - (2) systematically formatted with homogenous textual structure (i.e., in a format in which automated processes could be deployed without excessive risk of misaligning textual sequences) - (3) in the language predominantly represented by the pre-training corpus of Llama (i.e., English) - (4) covering a breadth of medical sub-domains, audiences (clinician, nurse, patient), and resource settings (high, low, and humanitarian response settings) | Source | Full Name | Tag | Guidelines | Words | Audience | Country | Released | |-|-|-|-|-|-|-|-| | **[AAFP](https://www.aafp.org)** | American Academy of Family Physicians | `aafp` | 50 | 9.4K | Doctor | USA | No | | **[CCO](https://www.cancercareontario.ca/en/guidelines-advice)** | Cancer Care Ontario | `cco` | 87 | 199K | Doctor | Canada | **Yes** | | **[CDC](https://www.cdc.gov/)** | Center for Disease Control and Prevention | `cdc` | 621 | 6.7M | Doctor | USA | **Yes** | | **[CMA](https://joulecma.ca/)** | Canadian Medical Association | `cma` | 431 | 1.7M | Doctor | Canada | **Yes** | | **[CPS](https://cps.ca)** | Canadian Paediatric Society | `cps` | 54 | 133K | Doctor | Canada | No | | **[drugs.com](https://www.drugs.com/)** | Drugs.com | `drugs` | 6548 | 4.1M | Both | International | No | | **[GuidelineCentral](https://www.guidelinecentral.com/)** | GuidelineCentral | `gc` | 1029 | 1M | Doctor | Mix | No | | **[ICRC](http://icrc.org/)** | International Committee of the Red Cross | `icrc` | 49 | 1.2M | Doctor | International | **Yes** | | **[IDSA](https://www.idsociety.org/)** | Infectious Diseases Society of America | `idsa` | 47 | 646K | Doctor | USA | No | | **[MAGIC](https://magicevidence.org/)** | Making GRADE The Irresistible Choice | `magic` | 52 | 415K | Doctor | Mix | No | | **[MayoClinic](https://www.mayoclinic.org/)** | MayoClinic | `mayo` | 1100 | 2.2M | Patient | USA | No | | **[NICE](https://www.nice.org.uk/guidance)** | National Institute for Health and Care Excellence | `nice` | 1656 | 8.1M | Doctor | UK | **Yes** | | **[PubMed](https://pubmed.ncbi.nlm.nih.gov)** | PubMed | `pubmed` | 1627 | 10.8M | Doctor | Mix | **Yes** | | **[RCH](https://www.rch.org.au/clinicalguide/about_rch_cpgs/welcome_to_the_clinical_practice_guidelines/)** | Royal Children's Hospital Melbourne | `rch` | 384 | 410K | Doctor | Australia | No | | **[SPOR](https://sporevidencealliance.ca/key-activities/cpg-asset-map/cpg-database/)** | Strategy for Patient-Oriented Research | `spor` | 217 | 1.1M | Doctor | Canada | **Yes** | | **[WHO](https://www.who.int/publications/who-guidelines)** | World Health Organization | `who` | 223 | 3.1M | Both | International | **Yes** | | **[WikiDoc](https://www.wikidoc.org/)** | WikiDoc | `wikidoc` | 33058 | 34M | Both | International | **Yes** | #### Data Collection and Processing <!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. --> PDF documents were converted to text using [GROBID](https://github.com/kermitt2/grobid). After extracting the raw text from each source, we cleaned data with an ad-hoc process to exclude irrelevant or repetitive content that did not contribute to the textual content, such as URLs, references, figures, table delimiters, and ill-formatted characters. This filtering procedure was performed differently for each source using a sample of 50 articles. Please note that this procedure is not perfect, as it may have removed useful information or kept superfluous content. We provide the `raw_text` for each article if you would like to perform your own cleaning step. Additionally, the text was standardized to a unified format with hierarchical section headers indicated by `'#'`, homogenous spacing `'\n\n'` separating paragraphs, and normalized lists formatted with `'- '` bullet points. Finally, all samples were deduplicated using title matching, and articles that were too short or not English were filtered out. #### Personal and Sensitive Information <!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. --> As the articles are publicly accessible, no personal or sensitive information is included. ## Dataset Structure <!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. --> Each row of the dataset represents one clinical practice guideline article, and consists of the following dataset fields (all strings): | Field | Description | Sources with field | |-------------|-------------------------------------------|------------------------------| | `id` | Unique identifier for each article | All | | `source` | Source tag (`cco`, `cdc`, `cma`, `icrc`, `nice`, `spor`, `who` or `wikidoc`)| All | | `title` | Title of the article | CMA, NICE & WikiDoc | | `url` | URL of the article | NICE, WikiDoc & PubMed | | `raw_text` | Unprocessed scraped article text | All | | `clean_text`| Cleaned and formatted article text | All | | `overview` | Short summary or abstract of the article | NICE & Pubmed | ## Uses <!-- Address questions around how the dataset is intended to be used. --> The dataset is intended for use in tasks related to text generation, specifically in the context of clinical practice guidelines. It can be employed for training language models and other natural language processing applications within the healthcare domain. ### Out-of-Scope Use <!-- This section addresses misuse, malicious use, and uses that the dataset will not work well for. --> - **Redistribution**: Please always check redistribution licenses before using the content as these may also evolve over time. To the best of our knowledge, we are following the redistribution licensing of each source and we invite users to inform us if that is not the case. - **Malicious use**: We do not support any use of this corpus that may be harmful. Creating tools that provide clinical advice is commendable, but extremely dangerous if not done with the appropriate care. Such tools need to be validated for safety and utility by medical professionals in randomized controlled trials. i.e. please do not create cowboy health apps that fool vulnerable users into thinking they are receiving validated advice. ## Bias, Risks, and Limitations <!-- This section is meant to convey both technical and sociotechnical limitations. --> - **Peer-Review Quality**: It is important to understand that while most sources are validated by internationally endorsed professional associations, a large proportion of articles are from Wikidoc which contains crowdsourced content. While edits in Wikidoc are generally restricted to expert review, the process of consensus and oversight is different from the traditional rigor of clinical guidelines. - **Representation**: This corpus is in English, and over-represents English-speaking regions. While we have included WHO and ICRC guidelines for low-resource settings, further work needs to be done to scrape sources from diverse contexts. - **Temporal scope**: Guidelines are constantly updated and these represent a snapshot of each in August 2023. Please re-scrape for updated content. ### Recommendations <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. --> We warmly invite users to help us build a more representative corpus with high-quality peer-reviewed clinical practice guidelines in various languages and representing the full scope of clinical specialties and geographic regions. We encourage users of this content to be mindful of its current limitations in temporal and geographic scope and we repeat our warning: creating tools that provide clinical advice is commendable, but extremely dangerous if not done with the appropriate care. Such tools need to be validated for safety and utility by medical professionals in randomized controlled trials. i.e. Please don’t create cowboy health apps that fool vulnerable users into thinking they are receiving validated advice. ## Acknowledgments The availability of open-access clinical practice guidelines (CPG) was critical to this work, and we thank all the societies listed above. A broader representation of geography, medical specialties, and contexts (especially low-resource settings) could be achieved through more standardized CPG formatting practices to ensure reliable textual extraction (e.g., releasing `.txt` or `.html` versions with structured content). We encourage the CPG community to continue to make these documents available (open-access with permissive licenses for incorporation into large language models) and easily usable. ## Authors - **Curation**: Mary-Anne Hartley - **Scraping**: Antoine Bonnet, Alexandre Sallinen, Igor Krawczuk, Kyle Matoba - **Cleaning**: Antoine Bonnet, Alexandre Sallinen ## Citation <!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. --> If you use the Clinical Guidelines corpus, please cite out work: ``` @misc{chen2023meditron70b, title={MEDITRON-70B: Scaling Medical Pretraining for Large Language Models}, author={Zeming Chen and Alejandro Hernández-Cano and Angelika Romanou and Antoine Bonnet and Kyle Matoba and Francesco Salvi and Matteo Pagliardini and Simin Fan and Andreas Köpf and Amirkeivan Mohtashami and Alexandre Sallinen and Alireza Sakhaeirad and Vinitra Swamy and Igor Krawczuk and Deniz Bayazit and Axel Marmet and Syrielle Montariol and Mary-Anne Hartley and Martin Jaggi and Antoine Bosselut}, year={2023}, eprint={2311.16079}, archivePrefix={arXiv}, primaryClass={cs.CL} } @software{epfmedtrn, author = {Zeming Chen and Alejandro Hernández-Cano and Angelika Romanou and Antoine Bonnet and Kyle Matoba and Francesco Salvi and Matteo Pagliardini and Simin Fan and Andreas Köpf and Amirkeivan Mohtashami and Alexandre Sallinen and Alireza Sakhaeirad and Vinitra Swamy and Igor Krawczuk and Deniz Bayazit and Axel Marmet and Syrielle Montariol and Mary-Anne Hartley and Martin Jaggi and Antoine Bosselut}, title = {MediTron-70B: Scaling Medical Pretraining for Large Language Models}, month = November, year = 2023, url = {https://github.com/epfLLM/meditron} } ```

license: 其他 license_name: Common Crawl license_link: LICENSE task_categories: - 文本生成 language: - 英语 pretty_name: 临床指南 size_categories: - 10K<n<100K configs: - config_name: default data_files: - split: train path: open_guidelines.jsonl tags: - 医学 - 健康 dataset_info: features: - name: id dtype: 字符串 - name: source dtype: 字符串 - name: title dtype: 字符串 - name: clean_text dtype: 字符串 - name: raw_text dtype: 字符串 - name: url dtype: 字符串 - name: overview dtype: 字符串 🎉 **全新上线** 🎉 PubMed 临床指南 我们于2023年12月23日将收录于PubMed及PubMed Central的1627条临床指南新增至本数据集。圣诞快乐! # 临床指南数据集 本临床指南语料库是一个包含来自17个高质量在线医学来源的47000条临床实践指南(Clinical Practice Guidelines, CPG)的全新数据集。该数据集是[Meditron](https://huggingface.co/epfl-llm/meditron-70b)大语言模型(Large Language Model, LLM)原始训练语料库的核心组成部分。我们从17个来源中选取了9个允许内容再分发的来源,从中提取了37000篇文章并公开发布,这9个来源分别为CCO、CDC、CMA、ICRC、NICE、PubMed、SPOR、WHO及WikiDoc。 你可以使用[epfLLM/meditron](https://github.com/epfLLM/meditron)中的代码对全部17个指南来源进行爬取与清洗。 ![临床实践指南来源](sources.png) ## 数据集详情 - **整理方**:[EPFL LLM团队](https://huggingface.co/epfl-llm) - **语言**:仅英语 - **许可证**:[Common Crawl基金会使用条款](https://commoncrawl.org/terms-of-use) - **代码仓库**:[epfLLM/meditron](https://github.com/epfLLM/meditron) - **相关论文**:*[MediTron-70B:面向大语言模型的医学预训练扩展](https://arxiv.org/abs/2311.16079)* - **知识截止日期**:2023年8月 ## 数据集构建 ### 整理动因 本数据集的整理旨在为大语言模型的医学训练提供高质量的临床实践指南集合。我们的临床指南语料库包含来自17个全球公认来源的48096篇文章,这些来源可为不同资源配置场景、多个医学领域(内科、儿科、肿瘤学、传染病学等)及多个地理区域的临床医师与患者提供指导性内容。 ### 源数据 临床实践指南是经过严谨研究的框架,用于指导医疗从业者与患者基于证据做出诊断、治疗与管理相关决策。 它们通过专家间系统性的协作共识流程编制而成,旨在基于最新的最佳实践证据制定推荐方案,同时兼顾可用资源与实际场景等现实考量以最大化收益。作为元分析的超级综述,临床指南位于*证据金字塔*的顶端,是可落地的循证实践的基础。 临床指南可根据多项维度进行区分: - **组织层级**:临床实践指南的产出组织粒度各异,从由国际专业医学协会主导的全球至医院级项目,到非正式联盟、区域或国家政府机构,再到单个非政府组织与医院。 - **地理范围**:地理范围覆盖全球(世界卫生组织, WHO)、国家(美国疾病控制与预防中心, CDC、英国国家卫生与临床优化研究所, NICE)、区域(安大略省、墨尔本)至机构(红十字国际委员会, ICRC、梅奥诊所)。由于本数据集仅聚焦英语内容,因此其分布偏向英语使用区域。 - **资源配置水平**:本语料库涵盖了高资源配置(安大略省、墨尔本)、低资源配置(WHO)及动荡环境(ICRC)下的医疗关切议题。 - **受众层级**:指南包含技术与会话两类词汇,目标受众为临床医师或患者(或两者兼具),部分指南还会针对特定主题(癌症、儿科、传染病)进行高度专业化的内容设计。 - **同行评审流程**:同行评审流程涵盖联合国机构(WHO)、机构审查委员会(ICRC)、专业协会(美国家庭医师学会, AAFP)至公众众包知识库(WikiDoc)。 - **文档规模**:文章长度差异巨大,从简短声明到百余页的指南不等。 #### 源数据生产者是谁? 本数据集的来源为17个全球公认的医学机构,覆盖了广泛的医疗场景与受众群体。 我们针对医学来源采用了务实的筛选标准,所选取的临床实践指南需满足: 1. 开放获取 2. 具有系统化的格式与统一的文本结构(即可以部署自动化流程且不会出现文本序列错位的过高风险) 3. 语言与Llama预训练语料库的主要语言一致(即英语) 4. 覆盖广泛的医学亚领域、受众(临床医师、护士、患者)及资源配置场景(高、低及人道主义响应场景) | 来源 | 全称 | 标签 | 指南数量 | 单词数 | 受众 | 国家/地区 | 是否可分发 | |-|-|-|-|-|-|-|-| | **[AAFP](https://www.aafp.org)** | 美国家庭医师学会(American Academy of Family Physicians) | `aafp` | 50 | 9.4K | 医师 | 美国 | 否 | | **[CCO](https://www.cancercareontario.ca/en/guidelines-advice)** | 安大略省癌症护理中心(Cancer Care Ontario) | `cco` | 87 | 199K | 医师 | 加拿大 | 是 | | **[CDC](https://www.cdc.gov/)** | 美国疾病控制与预防中心(Center for Disease Control and Prevention) | `cdc` | 621 | 6.7M | 医师 | 美国 | 是 | | **[CMA](https://joulecma.ca/)** | 加拿大医学协会(Canadian Medical Association) | `cma` | 431 | 1.7M | 医师 | 加拿大 | 是 | | **[CPS](https://cps.ca)** | 加拿大儿科协会(Canadian Paediatric Society) | `cps` | 54 | 133K | 医师 | 加拿大 | 否 | | **[drugs.com](https://www.drugs.com/)** | Drugs.com | `drugs` | 6548 | 4.1M | 两者兼具 | 国际 | 否 | | **[GuidelineCentral](https://www.guidelinecentral.com/)** | GuidelineCentral | `gc` | 1029 | 1M | 医师 | 混合 | 否 | | **[ICRC](http://icrc.org/)** | 红十字国际委员会(International Committee of the Red Cross) | `icrc` | 49 | 1.2M | 医师 | 国际 | 是 | | **[IDSA](https://www.idsociety.org/)** | 美国传染病学会(Infectious Diseases Society of America) | `idsa` | 47 | 646K | 医师 | 美国 | 否 | | **[MAGIC](https://magicevidence.org/)** | 推广GRADE标准联盟(Making GRADE The Irresistible Choice) | `magic` | 52 | 415K | 医师 | 混合 | 否 | | **[MayoClinic](https://www.mayoclinic.org/)** | 梅奥诊所(MayoClinic) | `mayo` | 1100 | 2.2M | 患者 | 美国 | 否 | | **[NICE](https://www.nice.org.uk/guidance)** | 英国国家卫生与临床优化研究所(National Institute for Health and Care Excellence) | `nice` | 1656 | 8.1M | 医师 | 英国 | 是 | | **[PubMed](https://pubmed.ncbi.nlm.nih.gov)** | PubMed | `pubmed` | 1627 | 10.8M | 医师 | 混合 | 是 | | **[RCH](https://www.rch.org.au/clinicalguide/about_rch_cpgs/welcome_to_the_clinical_practice_guidelines/)** | 墨尔本皇家儿童医院(Royal Children's Hospital Melbourne) | `rch` | 384 | 410K | 医师 | 澳大利亚 | 否 | | **[SPOR](https://sporevidencealliance.ca/key-activities/cpg-asset-map/cpg-database/)** | 患者导向研究战略联盟(Strategy for Patient-Oriented Research) | `spor` | 217 | 1.1M | 医师 | 加拿大 | 是 | | **[WHO](https://www.who.int/publications/who-guidelines)** | 世界卫生组织(World Health Organization) | `who` | 223 | 3.1M | 两者兼具 | 国际 | 是 | | **[WikiDoc](https://www.wikidoc.org/)** | WikiDoc | `wikidoc` | 33058 | 34M | 两者兼具 | 国际 | 是 | #### 数据收集与处理 我们使用[GROBID](https://github.com/kermitt2/grobid)将PDF文档转换为文本。 从每个来源提取原始文本后,我们通过临时流程对数据进行清洗,剔除与文本内容无关或重复的内容,例如URL、参考文献、图表、表格分隔符及格式错误的字符。 我们针对每个来源使用50篇文章的样本进行了差异化的过滤流程。请注意,该流程并非完美,可能会误删有用信息或保留冗余内容。若你希望自行执行清洗步骤,我们为每篇文章提供了`raw_text`字段。 此外,我们将文本标准化为统一格式:使用`'#'`表示层级章节标题,使用统一的换行符`' '`分隔段落,并将列表格式化为以`'- '`开头的项目符号。 最后,我们通过标题匹配对所有样本进行去重,并过滤掉过短或非英语的文章。 #### 个人与敏感信息 由于所有文章均可公开获取,本数据集未包含任何个人或敏感信息。 ## 数据集结构 本数据集的每一行对应一篇临床实践指南文章,包含以下全部为字符串类型的数据集字段: | 字段名 | 描述 | 包含该字段的来源 | |-------------|-------------------------------------------|------------------------------| | `id` | 每篇文章的唯一标识符 | 所有来源 | | `source` | 来源标签(`cco`、`cdc`、`cma`、`icrc`、`nice`、`spor`、`who` 或 `wikidoc`)| 所有来源 | | `title` | 文章标题 | CMA、NICE 及 WikiDoc | | `url` | 文章链接 | NICE、WikiDoc 及 PubMed | | `raw_text` | 未经过处理的爬取文章文本 | 所有来源 | | `clean_text`| 经过清洗与格式化的文章文本 | 所有来源 | | `overview` | 文章的简短摘要或概述 | NICE 及 PubMed | ## 数据集用途 本数据集旨在用于与临床实践指南相关的文本生成任务,可应用于医疗领域内的大语言模型训练及其他自然语言处理应用。 ### 超出范围的使用场景 - **再分发**:使用本数据集内容前,请务必核查其再分发许可证,因为许可证条款可能随时间变化。据我们所知,我们已遵循每个来源的再分发许可要求,若用户发现有不符之处,欢迎告知我们。 - **恶意使用**:我们不支持任何可能造成危害的本语料库使用方式。开发提供临床建议的工具是值得肯定的,但如果缺乏适当的严谨性则会极具危险性。此类工具必须由医疗专业人员通过随机对照试验验证其安全性与实用性。也就是说,请不要开发所谓“野生”医疗应用,误导脆弱用户以为其获得了经过验证的医疗建议。 ## 偏差、风险与局限性 - **同行评审质量**:需要注意的是,尽管大多数来源均经过国际认可的专业协会验证,但其中很大一部分文章来自WikiDoc的众包内容。尽管WikiDoc的编辑通常仅限专家审核,但其共识与监督流程与传统临床指南的严谨性有所差异。 - **代表性偏差**:本语料库仅包含英语内容,且过度代表英语使用区域。尽管我们纳入了针对低资源场景的WHO及ICRC指南,但仍需进一步爬取来自多元场景的来源以完善代表性。 - **时间范围局限**:临床指南会持续更新,本数据集仅收录了2023年8月的快照版本。若需获取更新内容,请重新进行爬取。 ### 建议 我们诚挚邀请用户协助我们构建更具代表性的语料库,涵盖多语言、覆盖全部临床专科与地理区域的高质量同行评审临床实践指南。 我们鼓励本数据集的使用者注意其当前在时间与地理范围上的局限性,并再次重申我们的警告:开发提供临床建议的工具是值得肯定的,但如果缺乏适当的严谨性则会极具危险性。此类工具必须由医疗专业人员通过随机对照试验验证其安全性与实用性。也就是说,请不要开发所谓“野生”医疗应用,误导脆弱用户以为其获得了经过验证的医疗建议。 ## 致谢 开放获取的临床实践指南对本工作至关重要,我们感谢上述所有机构。若能采用更标准化的临床实践指南格式以确保可靠的文本提取(例如发布包含结构化内容的`.txt`或`.html`版本),则可以进一步提升地理、医学专科及场景(尤其是低资源场景)的代表性。我们鼓励临床实践指南社区持续开放这些文档(以允许纳入大语言模型的宽松许可证实现开放获取)并确保其易于使用。 ## 作者 - **整理工作**:Mary-Anne Hartley - **爬取工作**:Antoine Bonnet、Alexandre Sallinen、Igor Krawczuk、Kyle Matoba - **清洗工作**:Antoine Bonnet、Alexandre Sallinen ## 引用 若你使用本临床指南语料库,请引用以下论文: @misc{chen2023meditron70b, title={MEDITRON-70B: Scaling Medical Pretraining for Large Language Models}, author={Zeming Chen and Alejandro Hernández-Cano and Angelika Romanou and Antoine Bonnet and Kyle Matoba and Francesco Salvi and Matteo Pagliardini and Simin Fan and Andreas Köpf and Amirkeivan Mohtashami and Alexandre Sallinen and Alireza Sakhaeirad and Vinitra Swamy and Igor Krawczuk and Deniz Bayazit and Axel Marmet and Syrielle Montariol and Mary-Anne Hartley and Martin Jaggi and Antoine Bosselut}, year={2023}, eprint={2311.16079}, archivePrefix={arXiv}, primaryClass={cs.CL} } @software{epfmedtrn, author = {Zeming Chen and Alejandro Hernández-Cano and Angelika Romanou and Antoine Bonnet and Kyle Matoba and Francesco Salvi and Matteo Pagliardini and Simin Fan and Andreas Köpf and Amirkeivan Mohtashami and Alexandre Sallinen and Alireza Sakhaeirad and Vinitra Swamy and Igor Krawczuk and Deniz Bayazit and Axel Marmet and Syrielle Montariol and Mary-Anne Hartley and Martin Jaggi and Antoine Bosselut}, title = {MediTron-70B: Scaling Medical Pretraining for Large Language Models}, month = November, year = 2023, url = {https://github.com/epfLLM/meditron} }

提供机构:
Billy6310
二维码
社区交流群
二维码
科研交流群
商业服务