遇见数据集

Steveeeeeeen/yodas-granary-it-neucodec-300k-5s30s

收藏
Hugging Face2026-05-15 更新2026-05-31 收录
官方服务:

资源简介:

Granary Italian VoxPopuli NeuCodec是一个用于NeuTTS微调的数据集,包含经过NeuCodec编码的意大利语语音数据。数据来源于`espnet/yodas-granary`数据集的`Italian`配置,流式处理并存储为可恢复的Parquet分片。数据集包含文本、编码后的音频代码、时长、源数据集信息等列,并应用了文本过滤(仅保留意大利语内容)和音频/文本筛选条件,如音频时长在5到30秒之间,文本字符数在8到260之间,字符每秒速率在3到32之间,最大训练样本数为300,000。

Granary Italian VoxPopuli NeuCodec is a dataset of NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from `espnet/yodas-granary` / `Italian` and uploaded as resumable Parquet shards. It includes columns such as text, codes, duration, source metadata, and applies text filtering (requiring Italian language), with constraints on audio duration (5.0 to 30.0 seconds), text length (8 to 260 characters), characters per second rate (3.0 to 32.0), and a maximum of 300,000 training samples.

提供机构:
Steveeeeeeen
二维码
社区交流群
二维码
科研交流群
商业服务