A Replication Dataset for Fundamental Frequency Estimation
收藏资源简介:
Part of the dissertation Pitch of Voiced Speech in the Short-Time Fourier Transform: Algorithms, Ground Truths, and Evaluation Methods. © 2020, Bastian Bechtold. All rights reserved. Estimating the fundamental frequency of speech remains an active area of research, with varied applications in speech recognition, speaker identification, and speech compression. A vast number of algorithms for estimatimating this quantity have been proposed over the years, and a number of speech and noise corpora have been developed for evaluating their performance. The present dataset contains estimated fundamental frequency tracks of 25 algorithms, six speech corpora, two noise corpora, at nine signal-to-noise ratios between -20 and 20 dB SNR, as well as an additional evaluation of synthetic harmonic tone complexes in white noise. The dataset also contains pre-calculated performance measures both novel and traditional, in reference to each speech corpus’ ground truth, the algorithms’ own clean-speech estimate, and our own consensus truth. It can thus serve as the basis for a comparison study, or to replicate existing studies from a larger dataset, or as a reference for developing new fundamental frequency estimation algorithms. All source code and data is available to download, and entirely reproducible, albeit requiring about one year of processor-time. Included Code and Data ground truth data.zip is a JBOF dataset of fundamental frequency estimates and ground truths of all speech files in the following corpora: CMU-ARCTIC (consensus truth) [1] FDA (corpus truth and consensus truth) [2] KEELE (corpus truth and consensus truth) [3] MOCHA-TIMIT (consensus truth) [4] PTDB-TUG (corpus truth and consensus truth) [5] TIMIT (consensus truth) [6] noisy speech data.zip is a JBOF datasets of fundamental frequency estimates of speech files mixed with noise from the following corpora: NOISEX [7] QUT-NOISE [8] synthetic speech data.zip is a JBOF dataset of fundamental frequency estimates of synthetic harmonic tone complexes in white noise. noisy_speech.pkl and synthetic_speech.pkl are pickled Pandas dataframes of performance metrics derived from the above data for the following list of fundamental frequency estimation algorithms: AUTOC [9] AMDF [10] BANA [11] CEP [12] CREPE [13] DIO [14] DNN [15] KALDI [16] MAPS MBSC [17] NLS [18] PEFAC [19] PRAAT [20] RAPT [21] SACC [22] SAFE [23] SHR [24] SIFT [25] SRH [26] STRAIGHT [27] SWIPE [28] YAAPT [29] YIN [30] noisy speech evaluation.py and synthetic speech evaluation.py are Python programs to calculate the above Pandas dataframes from the above JBOF datasets. They calculate the following performance measures: Gross Pitch Error (GPE), the percentage of pitches where the estimated pitch deviates from the true pitch by more than 20%. Fine Pitch Error (FPE), the mean error of grossly correct estimates. High/Low Octave Pitch Error (OPE), the percentage pitches that are GPEs and happens to be at an integer multiple of the true pitch. Gross Remaining Error (GRE), the percentage of pitches that are GPEs but not OPEs. Fine Remaining Bias (FRB), the median error of GREs. True Positive Rate (TPR), the percentage of true positive voicing estimates. False Positive Rate (FPR), the percentage of false positive voicing estimates. False Negative Rate (FNR), the percentage of false negative voicing estimates. F₁, the harmonic mean of precision and recall of the voicing decision. Pipfile is a pipenv-compatible pipfile for installing all prerequisites necessary for running the above Python programs. The Python programs take about an hour to compute on a fast 2019 computer, and require at least 32 Gb of memory. References: John Kominek and Alan W Black. CMU ARCTIC database for speech synthesis, 2003. Paul C Bagshaw, Steven Hiller, and Mervyn A Jack. Enhanced Pitch Tracking and the Processing of F0 Contours for Computer Aided Intonation Teaching. In EUROSPEECH, 1993. F Plante, Georg F Meyer, and William A Ainsworth. A Pitch Extraction Reference Database. In Fourth European Conference on Speech Communication and Technology, pages 837–840, Madrid, Spain, 1995. Alan Wrench. MOCHA MultiCHannel Articulatory database: English, November 1999. Gregor Pirker, Michael Wohlmayr, Stefan Petrik, and Franz Pernkopf. A Pitch Tracking Corpus with Evaluation on Multipitch Tracking Scenario. page 4, 2011. John S. Garofolo, Lori F. Lamel, William M. Fisher, Jonathan G. Fiscus, David S. Pallett, Nancy L. Dahlgren, and Victor Zue. TIMIT Acoustic-Phonetic Continuous Speech Corpus, 1993. Andrew Varga and Herman J.M. Steeneken. Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recog- nition systems. Speech Communication, 12(3):247–251, July 1993. David B. Dean, Sridha Sridharan, Robert J. Vogt, and Michael W. Mason. The QUT-NOISE-TIMIT corpus for the evaluation of voice activity detection algorithms. Proceedings of Interspeech 2010, 2010. Man Mohan Sondhi. New methods of pitch extraction. Audio and Electroacoustics, IEEE Transactions on, 16(2):262—266, 1968. Myron J. Ross, Harry L. Shaffer, Asaf Cohen, Richard Freudberg, and Harold J. Manley. Average magnitude difference function pitch extractor. Acoustics, Speech and Signal Processing, IEEE Transactions on, 22(5):353—362, 1974. Na Yang, He Ba, Weiyang Cai, Ilker Demirkol, and Wendi Heinzelman. BaNa: A Noise Resilient Fundamental Frequency Detection Algorithm for Speech and Music. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 22(12):1833–1848, December 2014. Michael Noll. Cepstrum Pitch Determination. The Journal of the Acoustical Society of America, 41(2):293–309, 1967. Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello. CREPE: A Convolutional Representation for Pitch Estimation. arXiv:1802.06182 [cs, eess, stat], February 2018. arXiv: 1802.06182. Masanori Morise, Fumiya Yokomori, and Kenji Ozawa. WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications. IEICE Transactions on Information and Systems, E99.D(7):1877–1884, 2016. Kun Han and DeLiang Wang. Neural Network Based Pitch Tracking in Very Noisy Speech. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 22(12):2158–2168, Decem- ber 2014. Pegah Ghahremani, Bagher BabaAli, Daniel Povey, Korbinian Riedhammer, Jan Trmal, and Sanjeev Khudanpur. A pitch extraction algorithm tuned for automatic speech recognition. In Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on, pages 2494–2498. IEEE, 2014. Lee Ngee Tan and Abeer Alwan. Multi-band summary correlogram-based pitch detection for noisy speech. Speech Communication, 55(7-8):841–856, September 2013. Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen, and Søren Holdt Jensen. Fast fundamental frequency estimation: Making a statistically efficient estimator computationally efficient. Signal Processing, 135:188–197, June 2017. Sira Gonzalez and Mike Brookes. PEFAC - A Pitch Estimation Algorithm Robust to High Levels of Noise. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 22(2):518—530, February 2014. Paul Boersma. Accurate short-term analysis of the fundamental frequency and the harmonics-to-noise ratio of a sampled sound. In Proceedings of the institute of phonetic sciences, volume 17, page 97—110. Amsterdam, 1993. David Talkin. A robust algorithm for pitch tracking (RAPT). Speech coding and synthesis, 495:518, 1995. Byung Suk Lee and Daniel PW Ellis. Noise robust pitch tracking by subband autocorrelation classification. In Interspeech, pages 707–710, 2012. Wei Chu and Abeer Alwan. SAFE: a statistical algorithm for F0 estimation for both clean and noisy speech. In INTERSPEECH, pages 2590–2593, 2010. Xuejing Sun. Pitch determination and voice quality analysis using subharmonic-to-harmonic ratio. In Acoustics, Speech, and Signal Processing (ICASSP), 2002 IEEE International Conference on, volume 1, page I—333. IEEE, 2002. Markel. The SIFT algorithm for fundamental frequency estimation. IEEE Transactions on Audio and Electroacoustics, 20(5):367—377, December 1972. Thomas Drugman and Abeer Alwan. Joint Robust Voicing Detection and Pitch Estimation Based on Residual Harmonics. In Interspeech, page 1973—1976, 2011. Hideki Kawahara, Masanori Morise, Toru Takahashi, Ryuichi Nisimura, Toshio Irino, and Hideki Banno. TANDEM-STRAIGHT: A temporally stable power spectral representation for periodic signals and applications to interference-free spectrum, F0, and aperiodicity estimation. In Acous- tics, Speech and Signal Processing, 2008. ICASSP 2008. IEEE International Conference on, pages 3933–3936. IEEE, 2008. Arturo Camacho. SWIPE: A sawtooth waveform inspired pitch estimator for speech and music. PhD thesis, University of Florida, 2007. Kavita Kasi and Stephen A. Zahorian. Yet Another Algorithm for Pitch Tracking. In IEEE International Conference on Acoustics Speech and Signal Processing, pages I–361–I–364, Orlando, FL, USA, May 2002. IEEE. Alain de Cheveigné and Hideki Kawahara. YIN, a fundamental frequency estimator for speech and music. The Journal of the Acoustical Society of America, 111(4):1917, 2002.
本数据集源自论文《短时傅里叶变换(Short-Time Fourier Transform)中浊语音基频:算法、基准真值(ground truth)与评估方法》,© 2020 Bastian Bechtold,保留所有权利。 语音基频估计仍是活跃的研究领域,在语音识别、说话人辨识与语音压缩等场景中拥有广泛应用价值。多年来,学界已提出大量针对该任务的算法,并开发了若干语音与噪声语料库用于算法性能评估。本数据集包含25种算法在6个语音语料库、2个噪声语料库,以及-20 dB至20 dB共9种信噪比(signal-to-noise ratio, SNR)下的基频估计轨迹,同时还包含对白噪声中合成谐波复合音的额外评估结果。 数据集还包含针对各语音语料库基准真值、算法自身的洁净语音估计结果,以及本研究的共识真值(consensus truth)所计算得到的新型与传统性能指标。因此,本数据集可作为对比研究的基础,或用于从更大规模数据集中复现已有研究,也可作为开发新型基频估计算法的参考。所有源代码与数据均可下载,且具备完全可复现性,不过其计算耗时约需一年的处理器运行时间。 ### 附带代码与数据文件说明 1. `ground truth data.zip`:为JBOF格式数据集,包含以下语料库中所有语音文件的基频估计结果与基准真值:CMU-ARCTIC(共识真值)[1]、FDA(语料库真值与共识真值)[2]、KEELE(语料库真值与共识真值)[3]、MOCHA-TIMIT(共识真值)[4]、PTDB-TUG(语料库真值与共识真值)[5]、TIMIT(共识真值)[6]。 2. `noisy speech data.zip`:为JBOF格式数据集,包含以下语料库噪声混合后的语音文件的基频估计结果:NOISEX [7]、QUT-NOISE [8]。 3. `synthetic speech data.zip`:为JBOF格式数据集,包含白噪声中合成谐波复合音的基频估计结果。 4. `noisy_speech.pkl`与`synthetic_speech.pkl`:为序列化(pickle)后的Pandas数据框(Pandas dataframe),包含基于上述JBOF数据集计算得到的以下25种基频估计算法的性能指标:AUTOC [9]、AMDF [10]、BANA [11]、CEP [12]、CREPE [13]、DIO [14]、DNN [15]、KALDI [16]、MAPS MBSC [17]、NLS [18]、PEFAC [19]、PRAAT [20]、RAPT [21]、SACC [22]、SAFE [23]、SHR [24]、SIFT [25]、SRH [26]、STRAIGHT [27]、SWIPE [28]、YAAPT [29]、YIN [30]。 5. `noisy speech evaluation.py`与`synthetic speech evaluation.py`:为Python程序,可基于上述JBOF数据集计算得到上述Pandas数据框。其计算的性能指标包括: - 总基频误差(Gross Pitch Error, GPE):估计基频与真实基频偏差超过20%的音高占比。 - 细基频误差(Fine Pitch Error, FPE):误差在允许范围内的估计结果的平均误差。 - 高低八度基频误差(High/Low Octave Pitch Error, OPE):属于GPE范畴且恰好为真实基频整数倍的音高占比。 - 剩余总误差(Gross Remaining Error, GRE):属于GPE范畴但不属于OPE范畴的音高占比。 - 剩余细偏差(Fine Remaining Bias, FRB):GRE类估计结果的中位数误差。 - 真阳性率(True Positive Rate, TPR):语音活性检测为真阳性的估计结果占比。 - 假阳性率(False Positive Rate, FPR):语音活性检测为假阳性的估计结果占比。 - 假阴性率(False Negative Rate, FNR):语音活性检测为假阴性的估计结果占比。 - F₁值:语音活性检测的精确率与召回率的调和均值。 6. `Pipfile`:为兼容Pipenv的依赖配置文件,用于安装运行上述Python程序所需的全部前置依赖。 该Python程序在2019年款高性能计算机上运行耗时约1小时,且至少需要32 GB内存。 ### 参考文献 [1] John Kominek与Alan W Black. CMU ARCTIC语音合成数据库, 2003. [2] Paul C Bagshaw, Steven Hiller与Mervyn A Jack. 增强型基频跟踪与F0轮廓处理用于计算机辅助语调教学. 见:EUROSPEECH, 1993. [3] F Plante, Georg F Meyer与William A Ainsworth. 基频提取参考语料库. 见:第四届欧洲语音通信与技术会议, 西班牙马德里, 1995: 837–840. [4] Alan Wrench. MOCHA多通道发音数据库:英语版, 1999年11月. [5] Gregor Pirker, Michael Wohlmayr, Stefan Petrik与Franz Pernkopf. 适用于多基频跟踪场景的基频跟踪语料库与评估. 2011: 4. [6] John S. Garofolo, Lori F. Lamel, William M. Fisher, Jonathan G. Fiscus, David S. Pallett, Nancy L. Dahlgren与Victor Zue. TIMIT声学-语音连续语音语料库, 1993. [7] Andrew Varga与Herman J.M. Steeneken. 自动语音识别评估:II. NOISEX-92:用于研究加性噪声对语音识别系统影响的数据库与实验. 语音通信, 1993, 12(3):247–251. [8] David B. Dean, Sridha Sridharan, Robert J. Vogt与Michael W. Mason. QUT-NOISE-TIMIT语料库用于语音活动检测算法评估. 见:Interspeech 2010会议论文集, 2010. [9] Man Mohan Sondhi. 新型基频提取方法. IEEE音频与电声学汇刊, 1968, 16(2):262–266. [10] Myron J. Ross, Harry L. Shaffer, Asaf Cohen, Richard Freudberg与Harold J. Manley. 基于平均幅度差函数的基频提取器. IEEE声学、语音与信号处理汇刊, 1974, 22(5):353–362. [11] Na Yang, He Ba, Weiyang Cai, Ilker Demirkol与Wendi Heinzelman. BANA:适用于语音与音乐的抗噪基频检测算法. IEEE/ACM音频、语音与语言处理汇刊, 2014, 22(12):1833–1848. [12] Michael Noll. 倒频谱基频测定. 美国声学学会期刊, 1967, 41(2):293–309. [13] Jong Wook Kim, Justin Salamon, Peter Li与Juan Pablo Bello. CREPE:用于基频估计的卷积表征方法. arXiv:1802.06182 [cs, eess, stat], 2018年2月. [14] Masanori Morise, Fumiya Yokomori与Kenji Ozawa. WORLD:基于声码器的高质量实时语音合成系统. IEICE信息与系统汇刊, 2016, E99.D(7):1877–1884. [15] Kun Han与DeLiang Wang. 高噪声环境下基于神经网络的基频跟踪. IEEE/ACM音频、语音与语言处理汇刊, 2014, 22(12):2158–2168. [16] Pegah Ghahremani, Bagher BabaAli, Daniel Povey, Korbinian Riedhammer, Jan Trmal与Sanjeev Khudanpur. 针对自动语音识别优化的基频提取算法. 见:2014年IEEE国际声学、语音与信号处理会议(ICASSP), 2014:2494–2498. [17] Lee Ngee Tan与Abeer Alwan. 基于多带相关图的噪声语音基频检测. 语音通信, 2013, 55(7-8):841–856. [18] Jesper Kjær Nielsen, Tobias Lindstrøm Jensen, Jesper Rindom Jensen, Mads Græsbøll Christensen与Søren Holdt Jensen. 快速基频估计:使统计高效估计器计算效率提升. 信号处理, 2017, 135:188–197. [19] Sira Gonzalez与Mike Brookes. PEFAC:抗高噪声的基频估计算法. IEEE/ACM音频、语音与语言处理汇刊, 2014, 22(2):518–530. [20] Paul Boersma. 采样声音的基频与谐波-噪声比的精确短时分析. 见:语音科学研究所论文集, 第17卷, 1993:97–110. [21] David Talkin. 鲁棒基频跟踪算法(RAPT). 语音编码与合成, 1995, 495:518. [22] Byung Suk Lee与Daniel PW Ellis. 基于子带自相关分类的抗噪基频跟踪. 见:Interspeech, 2012:707–710. [23] Wei Chu与Abeer Alwan. SAFE:适用于洁净与噪声语音的F0估计算法. 见:INTERSPEECH, 2010:2590–2593. [24] Xuejing Sun. 基于次谐波-谐波比的基频测定与语音质量分析. 见:2002年IEEE国际声学、语音与信号处理会议(ICASSP), 第1卷, 2002:I-333. [25] Markel. 用于基频估计的SIFT算法. IEEE音频与电声学汇刊, 1972, 20(5):367–377. [26] Thomas Drugman与Abeer Alwan. 基于残差谐波的联合鲁棒语音活性检测与基频估计. 见:Interspeech, 2011:1973–1976. [27] Hideki Kawahara, Masanori Morise, Toru Takahashi, Ryuichi Nisimura, Toshio Irino与Hideki Banno. TANDEM-STRAIGHT:周期性信号的时域稳定功率谱表征及其在无干扰频谱、F0与非周期性估计中的应用. 见:2008年IEEE国际声学、语音与信号处理会议(ICASSP), 2008:3933–3936. [28] Arturo Camacho. SWIPE:基于锯齿波波形的语音与音乐基频估计器. 佛罗里达大学博士论文, 2007. [29] Kavita Kasi与Stephen A. Zahorian. 另一种基频跟踪算法(YAAPT). 见:2002年IEEE国际声学、语音与信号处理会议, 美国奥兰多, 2002:I-361–I-364. [30] Alain de Cheveigné与Hideki Kawahara. YIN:适用于语音与音乐的基频估计器. 美国声学学会期刊, 2002, 111(4):1917.



