遇见数据集

Dataset for Continuous Piano Sustain Pedal Depth Estimation Under Varying Acoustic Conditions

收藏
Zenodo2025-06-27 更新2026-05-26 收录
官方服务:

资源简介:

This dataset supports the research presented in our ISMIR paper High-Resolution Sustain Pedal Depth Estimation from Piano Audio across Room Acoustics. Unlike traditional binary classification approaches, our work estimates continuous pedal depth values, capturing fine-grained expressiveness in classical piano performance. The dataset consists of extracted features from real piano performance recordings and synthetic versions rendered under various acoustic conditions, enabling further investigation into the influence of room acoustics on sustain pedal estimation. Source and Licensing For the synthetic data, we use MIDI files from MAESTRO v3.0.0 [1], following its original official train/validation/test split. Audio synthesis was produced using Pianoteq 8 Stage developed by Modartt. We have obtained explicit permission to distribute the processed audio features under a license that aligns with MAESTRO’s distribution guidelines. Dataset Structure The data is stored in three HDF5 (.h5) files, each containing "features", "average_labels", "instant_values", "metadata" (and can be retrieved using these keys). The synthesized audio files are large (~1.5 TB). If you would like access to these materials, please contact us directly. For each example (track), the corresponding row includes: Features:Combined log-Mel spectrograms (229 Mel bands) and MFCCs (20 coefficients), extracted from resampled 16kHz audio using a 2048-point FFT and a 10ms hop size. Features are normalized for each individual audio file (i.e., music piece) (per position across time frames) and stored as a [249, n_frames] array. Justifications of the parameter choices can be found in our paper. Average Labels:Frame-level labels indicating a distribution over pedal states (0 = off, 1 = partial, 2 = full), derived from MIDI CC64 values aligned to audio frame timing. Instant Values:Raw pedal depth values at the start, middle, and end of each frame, stored as a [3, n_frames] array. Metadata:Each track includes the following metadata fields (no field name, just retrieved using row indices): room_id: 0 for real audio, 1-4 for synthesized rooms) midi_id: The id for the original performance in MAESTRO pedal_multi_factor: Custom scaling factor for sustain pedal depth values (Note: We synthesized a version without sustain pedal usage and stored as pedal_multi_factor=0, which was not used in our project. Therefore, in this dataset, pedal_multi_factor=1 all the time.) split_encoded: MAESTRO split (0 = train, 1 = validation, 2 = test) Although we are unable to provide all audio files due to the size limit, we have included a few audio samples synthesized using Pianoteq. The files are named as [midi_id]+[room_id]. For example, 118iv.wav is the song with MIDI ID 118 under room setting 4. Here are the samples (all from the test set to include room 4): 118 Domenico Scarlatti: Sonata in D Minor, K. 213 MAESTRO path: 2018/MIDI-Unprocessed_Recital9-11_MID--AUDIO_09_R1_2018_wav--5.wav 297 Franz Schubert: "Gretchen am Spinnrade" MAESTRO path: 2009/MIDI-Unprocessed_08_R1_2009_01-04_ORIG_MID--AUDIO_08_R1_2009_08_R1_2009_03_WAV.wav 959 Ludwig van Beethoven: Sonata No. 12 in A-flat Major, Op. 26, First Movement MAESTRO path: 2008/MIDI-Unprocessed_06_R1_2008_01-04_ORIG_MID--AUDIO_06_R1_2008_wav--2.wav * The original MIDI files share the same names as the MAESTRO audio files, excluding the file extensions. * According to Modartt, synthesized audio files cannot be reused as a sample library or virtual instrument. Synthesis Parameters All audio files are rendered in uncompressed WAV format, 16-bit depth, and 44,100 Hz sampling rate using Pianoteq 8 Stage. Room Name Mix Dur Size PreD D.Mix D.Amt D.FB Piano Model 1 Dry Room - - - - - - - NY Steinway D Classical 2 Clean Studio +10dB 0.4s 12m 0s 6% 60ms 0% NY Steinway D Classical 3 Concert Hall +50dB 4s 50m 0.01s 15% 60ms 5% NY Steinway D Classical 4 Church +10dB 2.5s 18m 0s 25% 30ms 0% Bösendorfer 280VC Classical Notes: Mix: Reverb mix level Dur: Reverb duration Size: Room size PreD: Pre-delay before reverb onset D.Mix: Delay mix percentage D.Amt: Delay amount D.FB: Delay feedback percentage No reverberation is applied in the Dry Room (Room 1). Most parameters are default values in the preset. Other unlisted settings are all unmodified. Pianoteq Stage only supports the default microphone position, so we did not modify that part. For Room 4, we only synthesized the test part for our out-of-domain test. Data Access and Usage During model training and inference, individual examples are retrieved via a .json file that contains metadata entries in the following format: { "file_path": "test.h5", "example_index": 0, "num_frames": 16776, "room_id": 1, "midi_id": 7, "pedal_factor": 1, "split": 2 } References [1] C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.-Z. A. Huang, S. Dieleman, E. Elsen, J. Engel, and D. Eck,“Enabling factorized piano music modeling and generation with the MAESTRO dataset,”Proceedings of the 7th International Conference on Learning Representations (ICLR), New Orleans, Louisiana, United States, 2019. Paper If you end up using this dataset, please cite our ISMIR paper instead of the auto-generated Zenodo citation. @inproceedings{KZ25pedal, title={High-Resolution Sustain Pedal Depth Estimation from Piano Audio across Room Acoustics}, author={Kun Fang and Hanwen Zhang and Ziyu Wang and Ichiro Fujinaga}, booktitle={Proceedings of the 26th International Society for Music Information Retrieval Conference (ISMIR)}, year={2025}, address={Daejeon, Korea}, month={September}, day={21--25} }

提供机构:
Zenodo
创建时间:
2025-06-27
二维码
社区交流群
二维码
科研交流群
商业服务