遇见数据集

nicholas-ugbala-hf/chatdoctor-cleaned-10k

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

ChatDoctor HealthCareMagic — Cleaned 10k 是一个经过清洗和格式化的数据集子集,源自原始的 ChatDoctor HealthCareMagic 100k 数据集,专门设计用于 Llama 3.2 3B 模型的指令微调。原始数据集包含 112,165 行数据,经过清洗后保留 45,205 行,最终采样 10,000 行用于训练。清洗步骤包括移除包含平台伪影(如输入中的 ChatDoctor 引用)的行、移除输出少于 30 词的样本、移除总词数超过 600 的序列、去除输出中的填充开头(如问候语、平台签名)和结尾签名(如Best wishes、Hope this helps等),并进行二次过滤以移除任何残留的平台名称引用。数据集采用 Llama 3.2 聊天模板格式,包含系统、用户和助手轮次。数据分为训练集(9,000 个样本)和评估集(1,000 个样本),用于医疗领域语言模型的微调项目。

ChatDoctor HealthCareMagic — Cleaned 10k is a cleaned and formatted subset of the ChatDoctor HealthCareMagic 100k dataset, prepared for instruction fine-tuning of Llama 3.2 3B. The original dataset contains 112,165 raw rows, which were reduced to 45,205 after cleaning, and then sampled to 10,000 for training. Cleaning steps included removing rows with platform artifacts (such as ChatDoctor references in inputs), outputs under 30 words, sequences exceeding 600 combined words, stripping filler openings (e.g., greetings, platform sign-offs) and trailing sign-offs (e.g., Best wishes, Hope this helps) from outputs, and a second pass filter to remove any surviving platform name references. The dataset is formatted using the Llama 3.2 chat template with system, user, and assistant turns. It is split into train (9,000 samples) and eval (1,000 samples) sets, and is used in a healthcare LLM fine-tuning project.

提供机构:
nicholas-ugbala-hf
二维码
社区交流群
二维码
科研交流群
商业服务