Pairwise Piano Spectrogram Dataset for Diffusion-Based Music Composition
收藏资源简介:
This dataset consists of 8,346 paired audio segments derived from public-domain classical piano recordings collected from MuseScore. Each pair contains a simplified (beginner) and an advanced arrangement of the same piece, aligned in key, time signature, and overall structure. All recordings were segmented into 1-second clips and converted into invertible grayscale spectrograms using magnitude-squared short-time Fourier transform (STFT) normalization and power-law compression (exponent 0.125), enabling reconstruction via the Griffin–Lim algorithm. The dataset is split into training and validation subsets. It was constructed to train a conditional diffusion model for AI-assisted music composition. Spectrogram parameters: FFT size 2048, hop length 512, library: Librosa.



