FluidInference/ami-corpus-mirror
收藏资源简介:
该数据集是AMI会议语料库的一个镜像子集,专门用于FluidAudio的说话人日志化基准测试。它包含了AMI公开手动注释v1.6.2(重新打包自官方存档,内容相同,包括segments/、words/、corpusResources/meetings.xml)和官方16个会议AMI-SDM评估分割的音频文件(EN2002、ES2004、IS1009、TS3003 × a–d),这些音频文件从上游AMI语料库镜像获取。数据集由AMI联盟/爱丁堡大学根据知识共享署名4.0国际许可证分发,本仓库在相同许可证下未更改地重新分发并注明归属。使用此数据时需引用相关文献。
This dataset is a mirrored subset of the AMI Meeting Corpus, specifically designed for speaker diarization benchmarking of FluidAudio. It includes the publicly available manual annotations of AMI v1.6.2 (repackaged from the official archive with identical content, including directories segments/, words/, and the file corpusResources/meetings.xml) and the audio files from 16 official AMI-SDM evaluation split meetings (EN2002, ES2004, IS1009, TS3003 × a–d), which are mirrored from the upstream AMI Corpus. The dataset is distributed by the AMI Consortium / The University of Edinburgh under the Creative Commons Attribution 4.0 International License. This repository redistributes it unchanged under the same license with proper attribution. Relevant literature should be cited when using this dataset.




