遇见数据集

OpenAMD: A Multilingual Voicemail and Answering Machine Dataset

收藏
Zenodo2026-07-07 更新2026-08-01 收录
官方服务:

资源简介:

OpenAMD: A Multilingual Voicemail and Answering Machine Dataset Dataset accompanying the paper: [paper link to be added]. Overview The OpenAMD dataset is a publicly available dataset for non-commercial use. It contains audio recordings of humans and voicemails in several languages and under various augmentations in telephony-like settings. The dataset contains in total 4,264 recordings of which 533 are unique recordings: 233 humans and 300 machines. Basic audio information All audio is sampled at 8kHz and encoded to one of the variants of the standard telephony codec G.711 [1]; either PCMA or PCMU. Recordings are usually a few seconds long Dataset structure The dataset contains files grouped into several directories with the following hierarchical layout: open_amd/├── open_amd.csv/├── original/│ ├── human/ # 233 files│ ├── machine/ # 300 files├── noise_0dB/│ ├── human/ # 233 files│ ├── machine/ # 300 files├── noise_5dB/│ ├── human/ # 233 files│ ├── machine/ # 300 files├── noise_15dB/│ ├── human/ # 233 files│ ├── machine/ # 300 files├── noise_20dB/│ ├── human/ # 233 files│ ├── machine/ # 300 files ├── packet_loss_low/│ ├── human/ # 233 files│ ├── machine/ # 300 files ├── packet_loss_high/│ ├── human/ # 233 files│ ├── machine/ # 300 files The metadata regarding each file is located inside open_amd.csv. It contains the following information: Column Description Example filename Relative path to file, incldues both the channel condition and label of the file. noise_0dB/human/00f39e8c-7756-475c-a166-be0164564402_noise_0dB.wav folder(augmentation) Files are grouped into folders, depending on the channel condition: whether packet loss is present or whether static noise at various dB levels is present. Can take only one of the following values: original noise_0dB noise_5dB noise_10dB noise_15dB noise_20dB packet_loss_low packet_loss_high Each of the folders contains 533 files. Packet loss was simulated using the Gilbert-Elliot model [2]. Noise is static, additive white Gaussian noise. packet_loss_low language Language of the file. Can be one of: English: en (128) Bosnian/Croatian: bs-hr (112) German: de (77) Dutch: nl (56) Italian: it (53) Chinese (Mandarin): zh (43) Arabic: ar (43) Turkish: tr (20) bs-hr voice Unique name (string identifier) of the speaker in case of human recordings, or name of Azure TTS neural voice [3] in case of voicemails. There are 8 unique human speakers and 44 synthetic voices. zh-CN-YunjianNeural gender Associated speaker gender. Either M or F. M text Transcript/spoken text in the recording. Evet, Kerem. Kim arıyor? augmentation Contains information about any additional augmentation applied to the recording. This was mostly done to voicemails to prevent models from overfitting. The augmentations are pretty self-explanatory and simulate common behavior in telephone conversations. Possible values: no-aug click beep ring pause+beep pause ring+beep click+beep click+beep codec All audio is encoded to the telephony standard G.711 [1], either PCMA or PCMU variants. PCMU ground_truth True label of the file. Either HUMAN and MACHINE MACHINE Additionally, the csv file contains the following columns (one for each model analyzed in the paper): RNN-YAMNet [4] CNN6-LogReg CNN6-GBM-Balanced CNN6-GBM wav2vec-finetune [5] whisper-GPT-5.4 These columns represent the predictions of those models for each given file. Consult the paper for more information about the models. License This dataset is released under Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) You are free to share and adapt the material for any non-commercial use, provided appropriate credit is given. Commercial use: contact authors for a separate license. Contact Feel free to contact any of the authors via: [name.surname]@infobip.com References [1] International Telecommunication Union: Pulse code modulation (PCM) of voice frequencies. ITU-T Recommendation G.711, International Telecommunication Union, Geneva, Switzerland (Nov 1988), https://www.itu.int/rec/T-REC-G.711-198811-I/en [2] Pieper, J.: Relationships between gilbert-elliot burst error model parameters and error statistics. Tech. rep., Institute for Telecommunication Sciences (2023) [3] Microsoft: Azure Speech Studio. https://speech.microsoft.com/ (2026) [4] Altwlkany, K., Delalić, S., Selmanović, E., Alihodžić, A., Lovrić, I.: A recurrent neural network approach to the answering machine detection problem. In: 2024 47th MIPRO ICT and Electronics Convention (MIPRO). pp. 85–90. IEEE (2024) [5] Downie, J.: wav2vec-vm-finetune. https://huggingface.co/jakeBland/wav2vec-vm-finetune (2025)

提供机构:
Zenodo
创建时间:
2026-07-07
二维码
社区交流群
二维码
科研交流群
商业服务