librivox-mirror
收藏资源简介:
LibriVox Mirror 是一个快速、结构化且持续更新的 LibriVox 音频镜像数据集。它汇总了 LibriVox 项目中公开领域的朗读音频,并提供标准化的元数据索引。截至最新快照,数据集包含 6,247 本已出版书籍、133,969 个章节,总计约 39,324.5 小时的音频,覆盖 68 种语言,其中英语占绝大部分(38,900.7 小时),其他语言包括德语、法语、西班牙语、葡萄牙语、荷兰语、丹麦语、意大利语、波兰语、拉丁语、日语、阿拉伯语等。数据集提供三种配置:默认的 preview 配置包含受限的、可在浏览器中播放的原始 MP3 文件;sections 和 books 配置则提供完整的 Parquet 格式索引,音频数据通过 WebDataset 分片流式传输。每个样本链接到原始 LibriVox 项目和互联网档案馆的源文件,并保留上游校验和及镜像 SHA-256,音频未经转码。该数据集适用于自动语音识别(ASR)和文本转语音(TTS)任务,用户可利用 language 和 hash_partition 字段构建稳定的下游子集或评估分割。镜像的编译、整理、标准化元数据、索引和文档采用 CC BY 4.0 许可,原始 LibriVox 音频在美国属于公共领域。
LibriVox Mirror is a fast, structured, and continuously updated mirror dataset of LibriVox audio. It aggregates public domain audiobook recordings from the LibriVox project and provides standardized metadata indexing. As of the latest snapshot, the dataset contains 6,247 published books, 133,969 chapters, totaling approximately 39,324.5 hours of audio, covering 68 languages, with English dominating (38,900.7 hours) and other languages including German, French, Spanish, Portuguese, Dutch, Danish, Italian, Polish, Latin, Japanese, Arabic, etc. The dataset offers three configurations: the default preview configuration includes restricted, browser-playable original MP3 files; the sections and books configurations provide full Parquet-format indices with audio data streamed via WebDataset shards. Each sample links to the original LibriVox project and Internet Archive source files, preserving upstream checksums and mirror SHA-256, with audio untranscoded. The dataset is suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks; users can leverage the language and hash_partition fields to build stable downstream subsets or evaluation splits. The compilation, organization, standardized metadata, indexing, and documentation of the mirror are licensed under CC BY 4.0, while the original LibriVox audio is in the public domain in the United States.
数据集概述:LibriVox Mirror
LibriVox Mirror 是一个快速、结构化且持续更新的 LibriVox 音频镜像数据集,主要面向自动语音识别(ASR)和文本转语音(TTS)任务。数据集采用多语言音频内容,并遵循 CC BY 4.0 许可证,原始 LibriVox 音频在美国属于公共领域。
1. 数据集规模与快照
| 指标 | 数值 |
|---|---|
| 已发布书籍 | 6,744 本 |
| 已发布章节 | 144,558 个 |
| 音频总时长 | 42,276.3 小时 |
| 音频语言数量 | 68 种 |
| 隔离书籍(问题书籍) | 34 本 |
| 最后更新时间 (UTC) | 2026-08-29T17:46:54.698025Z |
2. 音频语言分布(部分)
| 语言 | 时长(小时) |
|---|---|
| 英语 | 41,829.0 |
| 德语 | 235.6 |
| 法语 | 46.0 |
| 西班牙语 | 33.0 |
| 葡萄牙语 | 23.3 |
| 荷兰语 | 21.7 |
| 丹麦语 | 11.2 |
| 意大利语 | 10.5 |
| 日语 | 10.4 |
| 波兰语 | 9.3 |
| 拉丁语 | 5.1 |
| 阿拉伯语 | 3.7 |
| 世界语 | 3.6 |
| 希腊语 | 2.4 |
| 中文 | 1.9 |
| 乌克兰语 | 1.9 |
| 俄语 | 1.4 |
| 韩语 | 0.2 |
| 其他语言(如泰米尔语、希伯来语、芬兰语等) | 各小于 1 小时 |
3. 数据集结构与配置
- 默认配置(preview):包含一组有限的、可直接在浏览器播放的原始 MP3 文件。
- 完整配置(sections 和 books):提供类型化的 Parquet 元数据文件,包括章节级和书籍级的索引。
- 权威音频文件通过 WebDataset 分片(位于
data/目录下)进行流式访问。 - 用户可通过
language和hash_partition字段构建稳定的下游子集或评估划分。 - 元数据中保留了 LibriVox 和 Internet Archive 的原始字段,以确保来源可追溯。
4. 数据集配置详情
| 配置名称 | 说明 | 数据文件路径 |
|---|---|---|
preview |
默认配置,包含可播放的示例音频 | metadata/preview/*.parquet |
sections |
章节级元数据 | metadata/sections/*.parquet |
books |
书籍级元数据 | metadata/books/*.parquet |
5. 许可与引用
- 镜像特定内容(如编译、策划、标准化元数据、索引和文档)采用 CC BY 4.0 许可证。
- 原始 LibriVox 音频在美国属于公共领域,未被本镜像重新授权,其他司法管辖区的规则可能有所不同。
- 推荐引用方式:
bibtex @misc{ding2026librivoxmirror, author = {James Ding}, title = {LibriVox Mirror}, year = {2026}, publisher = {Hugging Face}, howpublished = {url{https://huggingface.co/datasets/twangodev/librivox-mirror}}, note = {Continuously updated dataset} }
6. 数据溯源与完整性
- 每个样本均链接到对应的 LibriVox 项目和 Internet Archive 源文件。
- 保留上游校验和及镜像端的 SHA-256 哈希值。
- 音频文件未经转码,确保原始质量。




