MCSCSet
收藏资源简介:
MCSCSet是由清华大学深圳国际研究生院和腾讯Jarvis实验室合作创建的大型汉语拼写校正数据集,专注于医学领域。该数据集包含约200,000条真实医疗查询样本,每条样本均由医学专家手动标注,确保了数据的高质量和专业性。数据集不仅提供了详细的错误类型和位置信息,还包含了一个医学混淆集,用于自动生成新的医学领域拼写校正数据。MCSCSet的应用旨在解决医学文本中复杂和罕见医学实体的拼写错误问题,提高医疗信息检索和处理的准确性。
MCSCSet is a large-scale Chinese spelling correction dataset jointly created by Tsinghua University Shenzhen International Graduate School and Tencent Jarvis Lab, focusing on the medical field. This dataset contains approximately 200,000 real medical query samples, each manually annotated by medical experts to ensure the high quality and professionalism of the data. In addition to providing detailed error types and location information, the dataset also includes a medical confusion set for automatically generating new medical-domain spelling correction data. The application of MCSCSet aims to solve the problem of spelling errors of complex and rare medical entities in medical texts, and improve the accuracy of medical information retrieval and processing.

- 1MCSCSet: A Specialist-annotated Dataset for Medical-domain Chinese Spelling Correction清华大学深圳国际研究生院 · 2022年



