Calvin-Xu/FLFL-Aozora-Speech-Train
收藏资源简介:
该数据集是[Calvin-Xu/Furigana-Aozora-Speech](https://huggingface.co/datasets/Calvin-Xu/Furigana-Aozora-Speech/blob/main/README.md)的一个更严格清理版本,包含2,536,041条数据,原始数据来源于青空文庫及サピエ的音声デイジーデータ。数据集用于训练[Calvin-Xu/FLFL](https://huggingface.co/Calvin-Xu/FLFL)模型,该模型是基于[stockmark/gpt-neox-japanese-1.4b](https://huggingface.co/stockmark/gpt-neox-japanese-1.4b)微调的。数据集还包含了对Whisper生成的转录错误的过滤,以及对连浊问题的修正。
This is a cleaned-up version containing 2,536,041 entries, which are generated from the raw data Aozora Bunko and SAPIERs audio Daisy data to create a furigana-annotated speech corpus. The dataset is used to train the Calvin-Xu/FLFL model, which is fine-tuned from stockmark/gpt-neox-japanese-1.4b. Additional sanity-checking is implemented to filter on the reading of common kanji and eliminate wildly inaccurate entries. Furthermore, the dataset corrects many words consonant voicing issues, especially for numbers and counters.



