neurlang/ipa-lexicon-4v0-7M
收藏资源简介:
IPA音标词典(690万单词)是一个包含690多万个单词发音数据的国际音标(IPA)词典,涵盖350多种语言。数据来源于语音文件,通过neurlang/ipa-whisper-medium模型转换为IPA音标,并对多词IPA进行了后处理以生成单词语音记录。数据格式为JSONL/ZST,包含语言、词条、查询词、投票数、来源、下载链接、是否为ogg格式、ID和IPA音标等字段。质量评分基于投票系统,投票数代表发音准确性的社区置信度,建议使用投票数≥1的数据作为高质量子集。数据集存在一些局限性,如说话者分布不均、某些语言(如英语、西班牙语)过度代表、国家/语言标签可能不准确、音频质量参差不齐、IPA音标为模型生成而非人工验证,以及地区方言差异可能被模糊处理。
The IPA Phonetic Lexicon (6.9M words) is an International Phonetic Alphabet (IPA) lexicon containing over 6.9 million word pronunciations across more than 350 languages. The data is derived from speech files converted into IPA using the neurlang/ipa-whisper-medium model, with postprocessing to join multiword IPA into single-word records for appropriate languages. The data format is JSONL/ZST, including fields such as language, headword, query word, votes, origin, download URL, is_ogg flag, ID, and IPA transcription. Quality is scored via a voting system, where votes represent community confidence in pronunciation accuracy, and a filter of votes ≥ 1 is recommended for high-quality subsets. Limitations include uneven speaker distribution across countries, overrepresentation of some languages (e.g., English, Spanish), potential inaccuracies in country/language labels, variable audio quality, model-derived IPA (not manually verified), and possible blurring of regional dialect distinctions.



