遇见数据集

MM-Office Dataset: multi-view and multi-modal dataset in an office environment

收藏
Zenodo2022-02-21 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

MM-office is a multi-view and multi-modal dataset in an office environment (MM-Office) that records events, e.g., 'enter' to the office room, 'sit down' on the chair, and 'take out' something from a shelf, in the room assuming the daily work. These events are recorded simultaneously using eight non-directional microphones and four cameras. The audio and video clips are divided into scenes, each about 30 to 90 seconds. The amount of data was 880 clips per point and sensor. The labels available for training are given as multi-labels that indicate which each clip contains what event. Only the test data is annotated with a strong label containing the onset/offset time of each event. License: see the file named LICENSE.pdf Further information is available at [1] and Github: https://github.com/nttrd-mdlab/mm-office [1] Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito, Noboru Harada “Multi-view and Multi-modal Event Detection Utilizing Transformer-based Multi-sensor fusion,” in IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP), 2022.

MM-office(MM-Office)是一款面向办公场景的多视角多模态数据集,用于记录日常办公环境中的各类事件,例如进入办公室、在座椅上就座、从货架取出物品等。该数据集通过8台无指向性麦克风与4台摄像头同步采集上述事件的相关数据。采集得到的音视频片段被划分为场景单元,单段时长介于30至90秒之间。每个采集点位与传感器对应880段数据片段。用于模型训练的标签为多标签形式,用于标注每段片段所包含的事件类型;仅测试集数据配有强标注信息,包含各事件的起始与结束时间戳。授权协议详见名为LICENSE.pdf的文件。更多详细信息可查阅参考文献[1]及GitHub仓库:https://github.com/nttrd-mdlab/mm-office。[1] 正田雅弘(Masahiro Yasuda)、大西泰宪(Yasunori Ohishi)、斋藤翔一(Shoichiro Saito)、原田昇(Noboru Harada):《基于Transformer的多传感器融合多视角多模态事件检测》,收录于IEEE国际声学、语音与信号处理会议(ICASSP),2022年。

提供机构:
Zenodo
创建时间:
2022-02-17
二维码
社区交流群
二维码
科研交流群
商业服务