Chinese Polyphones with Pinyin (CPP)
收藏资源简介:
Chinese Polyphones with Pinyin (CPP) 数据集是由韩国科学技术院创建,旨在解决汉语拼音转换中的多音字问题。该数据集包含超过99,000个句子,专门用于训练和测试多音字发音的识别模型。数据集通过从维基百科提取中文文本,并由两名母语为中文的标注者进行人工标注,确保每个多音字的发音准确无误。CPP数据集的应用领域主要集中在汉语语音合成系统中,以提高多音字发音的准确性和自然度。
The Chinese Polyphones with Pinyin (CPP) dataset was developed by the Korea Advanced Institute of Science and Technology (KAIST) to address the challenge of polyphonic character pronunciation in Chinese pinyin conversion. This dataset contains over 99,000 sentences, and is specifically designed for training and testing models for polyphonic character pronunciation recognition. The CPP dataset is constructed by extracting Chinese texts from Wikipedia, followed by manual annotation conducted by two native Chinese speakers, ensuring the accurate pronunciation of every polyphonic character. The primary application scope of the CPP dataset lies in Chinese speech synthesis systems, aiming to enhance the accuracy and naturalness of polyphonic character pronunciation.




