JoelMba/PARHAF_gen_CASM2_guidelines_v4
收藏资源简介:
该数据集是一个用于命名实体识别(NER)任务的数据集,专注于医疗或健康领域的疾病实体标注。数据包含文本标记(tokens)和对应的命名实体标签(ner_tags),标签采用BIO标注方案,包括B-Disorders(疾病开始)、I-Disorders(疾病内部)和O(非疾病实体)。数据集划分为训练集(12828个示例)和验证集(4466个示例),适用于自然语言处理模型训练和评估,特别是在医疗文本分析场景中。
This dataset is designed for Named Entity Recognition (NER) tasks, focusing on disease entity annotation in medical or health-related contexts. It includes text tokens and corresponding named entity labels (ner_tags) using the BIO tagging scheme, with labels for B-Disorders (beginning of disease), I-Disorders (inside disease), and O (non-disease entity). The dataset is split into training (12,828 examples) and validation (4,466 examples) sets, suitable for training and evaluating natural language processing models, particularly in medical text analysis scenarios.




