遇见数据集

mllp/LHCP-ASR

收藏
Hugging Face2026-04-14 更新2026-04-12 收录
官方服务:

资源简介:

--- license: cc-by-nc-nd-4.0 task_categories: - automatic-speech-recognition language: - en tags: - particle-physics - speech-corpus - domain-adaptation - longform - segments - manual-transcription - cern size_categories: - 100K<n<1M dataset_info: - config_name: longform features: - name: audio dtype: audio - name: transcription dtype: string splits: - name: train num_bytes: 48517243701.0 num_examples: 560 - name: dev_2020 num_bytes: 1365959644.0 num_examples: 14 - name: dev_2022 num_bytes: 1104577536.0 num_examples: 11 - name: test_2020 num_bytes: 1581799930.0 num_examples: 15 - name: test_2022 num_bytes: 3155188963.0 num_examples: 32 download_size: 52181447847 dataset_size: 55724769774.0 - config_name: segments features: - name: audio dtype: audio - name: transcription dtype: string splits: - name: train num_bytes: 40009686180.696 num_examples: 174376 - name: dev_2020 num_bytes: 1318912292.128 num_examples: 3896 - name: dev_2022 num_bytes: 754694121.774 num_examples: 3622 - name: test_2020 num_bytes: 1672421485.652 num_examples: 4017 - name: test_2022 num_bytes: 2277429244.504 num_examples: 9738 download_size: 46780473298 dataset_size: 46033143324.754 configs: - config_name: longform data_files: - split: train path: longform/train-* - split: dev_2020 path: longform/dev_2020-* - split: dev_2022 path: longform/dev_2022-* - split: test_2020 path: longform/test_2020-* - split: test_2022 path: longform/test_2022-* - config_name: segments data_files: - split: train path: segments/train-* - split: dev_2020 path: segments/dev_2020-* - split: dev_2022 path: segments/dev_2022-* - split: test_2020 path: segments/test_2020-* - split: test_2022 path: segments/test_2022-* --- # LHCP-ASR This dataset is another version of the [LHCP-ASR](https://github.com/mllpresearch/LHCP-ASR) corpus, an English speech dataset for narrow-domain ASR benchmarking in high-energy physics. Unlike the original distribution, which includes video, slides and text data, this version focuses entirely on audio-text pairs DESCRIPTION ----------- The speech data are **30 hours** of LHCP plenary conference talks (2020, 2022) with manual (human) verbatim transcriptions and **205 hours** of LHCP conference talks (2020-2022) with automatic verbatim transcriptions (pseudo-labels) for training/adaptation. This version has been released in two formats: `segments` of the talks (less than 30s each) and `longform`, which provides the full talk. ### Data structure Each sample in the dataset contains: * `audio` * `transcription` USAGE ----- You can load this dataset directly using the `datasets` library: ```python from datasets import load_dataset # Segmented version segmented_dataset = load_dataset("mllp/LHCP-ASR", "segments") # Longform version longform_dataset = load_dataset("mllp/LHCP-ASR", "longform") ```` Both configurations (`longform` and `segments`) include the following splits: - `train` - `dev_2020` - `dev_2022` - `test_2020` - `test_2022` RESULTS ------- We report here, for the development and test sets, the **longform** WER% of some Whisper models using normalised references (lowercased, no punctuation). The base models WER% are from the [original paper](https://www.isca-archive.org/interspeech_2025/santamariajorda25_interspeech.html) and the fine-tuned models WER% can be checked in this [final degree project](https://riunet.upv.es/entities/publication/914a3a30-6a42-404c-8fa8-805b70f1317c). | Model | 2020 dev | 2020 test | 2022 dev | 2022 test | |---|---|---|---|---| | **whisper medium** | 13.2 | 15.9 | 17.3 | 17.7 | | **whisper turbo** | 13.8 | 15.4 | 16.7 | 16.7 | | **whisper medium FT** | 12.1 | 13.2 | 14.4 | 14.7 | | **whisper turbo FT** | **12.0** | **12.9** | **14.3** | **14.5** | CITATION -------- If you use this dataset, please cite the original work: ```bibtex @inproceedings{santamariajorda25_interspeech, title = {{LHCP-ASR: An English Speech Corpus of High-Energy Particle Physics Talks for Narrow-Domain ASR Benchmarking}}, author = {Jaume Santamaría-Jordà and Pablo Segovia-Martínez and Gonçal V. {Garcés Díaz-Munío} and Joan Albert Silvestre-Cerdà and Adrià Giménez and Rubén {Gaspar Aparicio} and René {Fernández Sánchez} and Jorge Civera and Albert Sanchis and Alfons Juan}, year = {2025}, booktitle = {{Interspeech 2025}}, pages = {4033--4037}, doi = {10.21437/Interspeech.2025-2630}, issn = {2958-1796}, } ``` For more details on the original dataset, visit [https://github.com/mllpresearch/LHCP-ASR](https://github.com/mllpresearch/LHCP-ASR). LEGAL DISCLAIMER --------------- Speech and text data were provided by the [European Organization for Nuclear Research (CERN)](https://home.cern/) under PO OV9177345. The following disclaimers are those available in the [CERN Document Server (CDS)](https://cds.cern.ch/) repository on May 30th, 2025: #### CERN Document Server - Terms and Conditions [![DOI](https://repository.cern/badge/DOI/10.17181/s2cm2-jaj10.svg)](https://doi.org/10.17181/s2cm2-jaj10) Use of the CERN Document Server service (hereafter "CDS") denotes agreement with the following terms of use: * CDS is provided free of charge. It serves as a comprehensive institutional repository and dissemination platform for the research and historical output produced by CERN, the European Organization for Nuclear Research, and its members of personnel. See Content Policy [1] for more details. * By uploading content to CDS, the content provider affirms that such content complies with all applicable laws, licence conditions and third party rights, and shall hold CERN free and harmless from any related liability. * All content is provided "as is" and without warranty of any kind. The user shall hold CERN and individual content providers free and harmless from any related liability in connection with its use of such content. * Users shall respect copyright and all applicable licence conditions. The download and use of content from CDS does not amount to a transfer of intellectual property. * CERN reserves the right, without notice or liability, and at its sole discretion, to restrict or remove a user's access or remove any uploaded content, where it considers that use of CDS interferes with its operations or violates these Terms and Conditions, and/or applicable laws. * CERN bases CDS on leading technologies and architectures, operated within the limits of its financial and human resources and made available by CERN on an "as is" and "best efforts" basis. Access to, availability and use of CDS is not guaranteed nor can be expected. * CERN excludes and disclaims all liability for damage resulting from users' access, or inability to access, or use of CDS. * These terms and conditions of use are subject to change by CERN at any time and without notice, other than through posting the updated terms on the CDS website. Any revised terms and conditions of use shall become effective immediately upon posting. If you have any questions or comments with respect to CDS, or if you are unsure whether your intended use is in line with these Terms and Conditions, or if you seek permission for a use that does not fall within these Terms and Conditions, please contact CDS support. [1] CDS Content Policy [![DOI](https://repository.cern/badge/DOI/10.17181/8sm4v-js382.svg)](https://doi.org/10.17181/8sm4v-js382) LICENSE ------- This dataset is licenced under CC-BY-NC-ND 4.0. To view a copy of this licence, visit https://creativecommons.org/licenses/by-nc-nd/4.0/

许可协议:CC BY-NC-ND 4.0 任务类别:自动语音识别(Automatic Speech Recognition, ASR) 语言:英语(en) 标签: - 粒子物理(particle-physics) - 语音语料库(speech-corpus) - 领域自适应(domain-adaptation) - 长文本格式(longform) - 分段格式(segments) - 人工转录(manual-transcription) - 欧洲核子研究组织(CERN) 样本规模:10万 < 样本数量 < 100万 数据集信息: - 配置名称:longform(长文本格式) 字段信息: - 字段名:audio,数据类型:音频 - 字段名:transcription,数据类型:字符串 拆分子集: - 训练集(train):数据量48517243701.0字节,样本数量560 - 2020年开发集(dev_2020):数据量1365959644.0字节,样本数量14 - 2022年开发集(dev_2022):数据量1104577536.0字节,样本数量11 - 2020年测试集(test_2020):数据量1581799930.0字节,样本数量15 - 2022年测试集(test_2022):数据量3155188963.0字节,样本数量32 下载总大小:52181447847字节,数据集总大小:55724769774.0字节 - 配置名称:segments(分段格式) 字段信息: - 字段名:audio,数据类型:音频 - 字段名:transcription,数据类型:字符串 拆分子集: - 训练集(train):数据量40009686180.696字节,样本数量174376 - 2020年开发集(dev_2020):数据量1318912292.128字节,样本数量3896 - 2022年开发集(dev_2022):数据量754694121.774字节,样本数量3622 - 2020年测试集(test_2020):数据量1672421485.652字节,样本数量4017 - 2022年测试集(test_2022):数据量2277429244.504字节,样本数量9738 下载总大小:46780473298字节,数据集总大小:46033143324.754字节 配置项: - 配置名称:longform(长文本格式) 数据文件路径: - 训练集(train):longform/train-* - 2020年开发集(dev_2020):longform/dev_2020-* - 2022年开发集(dev_2022):longform/dev_2022-* - 2020年测试集(test_2020):longform/test_2020-* - 2022年测试集(test_2022):longform/test_2022-* - 配置名称:segments(分段格式) 数据文件路径: - 训练集(train):segments/train-* - 2020年开发集(dev_2020):segments/dev_2020-* - 2022年开发集(dev_2022):segments/dev_2022-* - 2020年测试集(test_2020):segments/test_2020-* - 2022年测试集(test_2022):segments/test_2022-* # LHCP-ASR 本数据集为[LHCP-ASR](https://github.com/mllpresearch/LHCP-ASR)语料库的衍生版本,是一款面向高能物理领域窄域自动语音识别评测的英语语音数据集。与包含视频、幻灯片与文本数据的原始分发版本不同,本版本仅聚焦于音频-文本配对数据。 ## 数据集描述 本数据集的语音数据包含30小时带有**人工逐字转录**的LHCP全体大会演讲音频(2020、2022年),以及205小时带有自动逐字转录(伪标签)的LHCP大会演讲音频(2020-2022年),用于模型训练与领域自适应。本版本提供两种格式:演讲分段(segments,单段时长小于30秒)与完整演讲(longform)。 ### 数据结构 数据集中的每条样本包含以下字段: * `audio`:音频数据 * `transcription`:转录文本 ## 使用方法 可通过`datasets`库直接加载本数据集: python from datasets import load_dataset # 分段格式版本 segmented_dataset = load_dataset("mllp/LHCP-ASR", "segments") # 长文本格式版本 longform_dataset = load_dataset("mllp/LHCP-ASR", "longform") 两种配置(`longform`与`segments`)均包含以下拆分子集: - `train`:训练集 - `dev_2020`:2020年开发集 - `dev_2022`:2022年开发集 - `test_2020`:2020年测试集 - `test_2022`:2022年测试集 ## 评测结果 本文针对开发集与测试集,报告了使用标准化参考文本(小写化、无标点)时,多款Whisper模型在`longform`配置下的词错误率(Word Error Rate, WER%)。基础模型的WER%数据来自[原论文](https://www.isca-archive.org/interspeech_2025/santamariajorda25_interspeech.html),微调模型的WER%可参阅该[本科毕业设计](https://riunet.upv.es/entities/publication/914a3a30-6a42-404c-8fa8-805b70f1317c)。 | 模型 | 2020开发集 | 2020测试集 | 2022开发集 | 2022测试集 | |---|---|---|---|---| | **whisper medium** | 13.2 | 15.9 | 17.3 | 17.7 | | **whisper turbo** | 13.8 | 15.4 | 16.7 | 16.7 | | **whisper medium FT** | 12.1 | 13.2 | 14.4 | 14.7 | | **whisper turbo FT** | **12.0** | **12.9** | **14.3** | **14.5** | ## 引用说明 若使用本数据集,请引用以下原始文献: bibtex @inproceedings{santamariajorda25_interspeech, title = {{LHCP-ASR: An English Speech Corpus of High-Energy Particle Physics Talks for Narrow-Domain ASR Benchmarking}}, author = {Jaume Santamaría-Jordà and Pablo Segovia-Martínez and Gonçal V. {Garcés Díaz-Munío} and Joan Albert Silvestre-Cerdà and Adrià Giménez and Rubén {Gaspar Aparicio} and René {Fernández Sánchez} and Jorge Civera and Albert Sanchis and Alfons Juan}, year = {2025}, booktitle = {{Interspeech 2025}}, pages = {4033--4037}, doi = {10.21437/Interspeech.2025-2630}, issn = {2958-1796}, } 如需了解原始数据集的更多细节,请访问[https://github.com/mllpresearch/LHCP-ASR](https://github.com/mllpresearch/LHCP-ASR)。 ## 法律声明 语音与文本数据由**欧洲核子研究组织(CERN, European Organization for Nuclear Research)**依据PO OV9177345协议提供。以下免责声明取自2025年5月30日欧洲核子研究中心文档服务器(CDS, CERN Document Server)仓库中的内容: ### 欧洲核子研究中心文档服务器使用条款 [![DOI](https://repository.cern/badge/DOI/10.17181/s2cm2-jaj10.svg)](https://doi.org/10.17181/s2cm2-jaj10) 使用欧洲核子研究中心文档服务器(以下简称“CDS”)服务即表示您同意以下使用条款: 1. CDS免费向用户开放,作为欧洲核子研究组织及其工作人员产出的研究成果与历史资料的综合性机构知识库与传播平台。更多详情请参阅[内容政策[1]](https://doi.org/10.17181/8sm4v-js382)。 2. 向CDS上传内容时,内容提供者应确保该内容符合所有适用法律、许可协议条款与第三方权利要求,并保证欧洲核子研究组织不会因此承担任何相关责任。 3. 所有内容均按“现状”提供,不附带任何形式的担保。用户因使用该内容引发的任何相关责任,均由用户本人与内容提供者承担,欧洲核子研究组织不承担任何责任。 4. 用户应尊重版权与所有适用的许可协议条款。从CDS下载并使用内容不代表知识产权的转让。 5. 欧洲核子研究组织保留随时、无需通知且自行决定限制或移除用户访问权限,或移除已上传内容的权利,若其认为使用CDS会干扰其正常运营,或违反本使用条款及/或适用法律。 6. 欧洲核子研究组织基于领先技术与架构运营CDS,运营范围受限于其财务与人力资源,并按“现状”与“最大努力”原则向用户提供服务。CDS的访问、可用性与使用不提供任何保证,也无法预期。 7. 欧洲核子研究组织不承担因用户访问、无法访问或使用CDS所导致的任何损害赔偿责任。 8. 欧洲核子研究组织可随时修改本使用条款,无需提前通知,仅需在CDS网站上发布更新后的条款即可。任何修订后的使用条款自发布之日起立即生效。 若您对CDS有任何疑问或意见,或不确定您的预期使用是否符合本使用条款,或寻求超出本条款范围的使用许可,请联系CDS支持团队。 [1] CDS内容政策 [![DOI](https://repository.cern/badge/DOI/10.17181/8sm4v-js382.svg)](https://doi.org/10.17181/8sm4v-js382) ## 许可协议 本数据集采用CC BY-NC-ND 4.0许可协议进行授权。如需查看许可协议副本,请访问https://creativecommons.org/licenses/by-nc-nd/4.0/

提供机构:
mllp
二维码
社区交流群
二维码
科研交流群
商业服务