AliMeeting4MUG Corpus
收藏资源简介:
AliMeeting4MUG Corpus是由阿里巴巴达摩院语音实验室和浙江大学联合创建的大型中文会议数据集,包含654个多样话题的会议记录,旨在推动长篇口语语言处理(SLP)技术的发展。该数据集通过手动转录和标注,支持多种SLP任务,如话题分割、摘要生成和关键短语提取等。数据集的创建过程严格遵循数据收集和标注的标准流程,确保数据质量。应用领域广泛,主要用于提高会议信息处理的效率和准确性,解决长篇口语文档处理中的关键技术问题。
AliMeeting4MUG Corpus is a large-scale Chinese meeting corpus jointly created by the Speech Lab of Alibaba DAMO Academy and Zhejiang University. It contains 654 meeting recordings covering diverse topics, and is aimed at advancing the development of long-form spoken language processing (SLP) technologies. The corpus is manually transcribed and annotated, supporting multiple SLP tasks including topic segmentation, summary generation, key phrase extraction and other related tasks. The construction of the dataset strictly follows standard data collection and annotation procedures to ensure high data quality. It has a wide range of application scenarios, mainly used to improve the efficiency and accuracy of meeting information processing and address key technical challenges in long-form spoken document processing.




