遇见数据集

Calvin-Xu/FLFL-Aozora-Speech-Train

收藏
Hugging Face2024-08-22 更新2024-12-14 收录
官方服务:

资源简介:

该数据集是[Calvin-Xu/Furigana-Aozora-Speech](https://huggingface.co/datasets/Calvin-Xu/Furigana-Aozora-Speech/blob/main/README.md)的一个更严格清理版本,包含2,536,041条数据,原始数据来源于青空文庫及サピエ的音声デイジーデータ。数据集用于训练[Calvin-Xu/FLFL](https://huggingface.co/Calvin-Xu/FLFL)模型,该模型是基于[stockmark/gpt-neox-japanese-1.4b](https://huggingface.co/stockmark/gpt-neox-japanese-1.4b)微调的。数据集还包含了对Whisper生成的转录错误的过滤,以及对连浊问题的修正。

This is a cleaned-up version containing 2,536,041 entries, which are generated from the raw data Aozora Bunko and SAPIERs audio Daisy data to create a furigana-annotated speech corpus. The dataset is used to train the Calvin-Xu/FLFL model, which is fine-tuned from stockmark/gpt-neox-japanese-1.4b. Additional sanity-checking is implemented to filter on the reading of common kanji and eliminate wildly inaccurate entries. Furthermore, the dataset corrects many words consonant voicing issues, especially for numbers and counters.

提供机构:
Calvin-Xu
二维码
社区交流群
二维码
科研交流群
商业服务