Subjective ratings and objective metric predictions of generic and personalized own voice reconstruction systems
收藏资源简介:
This upload contains subjective ratings, objective metric predictions, and audio files for the preprint "Subjective quality evaluation of personalized own voice reconstruction systems". Abstract Own voice pickup technology for hearable devices facilitates communication in noisy environments. Own voice reconstruction (OVR) systems enhance the quality and intelligibility of the recorded noisy own voice signals. Since disturbances affecting the recorded own voice signals depend on individual factors, personalized OVR systems have the potential to outperform generic OVR systems. In this paper, we propose personalizing OVR systems through data augmentation and fine-tuning, comparing them to their generic counterparts. We investigate the influence of personalization on speech quality assessed by objective metrics and conduct a subjective listening test to evaluate quality under various conditions. In addition, we assess the prediction accuracy of the objective metrics by comparing predicted quality with subjectively measured quality. Our findings suggest that personalized OVR provides benefits over generic OVR for some talkers only. Our results also indicate that performance comparisons between systems are not always accurately predicted by objective metrics. In particular, certain disturbances lead to a consistent overestimation of quality compared to actual subjective ratings. Quality predictions Quality predictions are stored in the folder quality_predictions/. Each CSV file has these columns: - rating_stimulus (processing condition)- sentence (which sentence index, ranging from 0 to 2)- base_trial_id - this is the noise type in the noisy microphone signals- SNR the signal-to-noise ratio in dB at the outer microphone (here, always 0)- predicted_score the score predicted by the metric.- talker the talker in the recorded microphone signals. good talker corresponds to the high predicted benefit case, bad talker to the low predicted benefit case. Some of the metrics have their own CSV file, others are put together.Metrics:- LEAP (leap_predictions.csv)- eMoBi-Q (emobiq_predictions.csv)- PEMO-Q PSM (pemoq_predictions.csv)- WV-MOS (wvmos_predictions.csv)- PESQ, ESTOI, DNSMOS (SIG, BAK, OVRL, P808), SCORE-Q (MOS, distance) -> (pesq_estoi_dnsmos_scoreq_predictions.csv) This CSV file has an extra column metric containing the information which metric the predicted_score in each row belongs to Subjective ratings The subjective ratings are stored in subjective_ratings_mushra.csv. The CSV file has these columns: - rating_stimulus (processing condition)- sentence (which sentence index, ranging from 0 to 2)- base_trial_id - this is the noise type in the noisy microphone signals- rating_score the subjective MUSHRA rating, on a scale from 0 to 100- talker the talker in the recorded microphone signals. good talker corresponds to the high predicted benefit case, bad talker to the low predicted benefit case.- ID the subject/participant identifier Audio files The unprocessed and processed stimuli / audio signals used in the listening experiments are stored in the folder audio_files.Each subfolder fold0 and fold6 contains signals from the talkers with index 0 (low predicted benefit case) and 6 (high predicted benefit case). In each of these, there are subfolders containing signals for each noise type (e.g., factory_diffuse_0dBSNR).The filenames for the stimuli are as follows:- clean_reference_0.wav are the clean speech signals at the outer microphone without noise that were used as reference for the MUSHRA-like test. Here, the _0 indicates the sentence index (0,1,2).- noisy_anchor refers to the noisy outer microphone signal (unprocessed)- noisy_inear refers to the noisy in-ear microphone signal (unprocessed)- output_mwf refers to signals processed by the MWF- output_cv_eben_nonindiv_sim_nonindiv_finetune refers to signals processed by the EBEN generator DNN (trained with generic data augmentation and generic fine-tuning)- output_cv_nonindiv_sim refers to signals processed by the FT-JNF DNN (trained with generic data augmentation)- output_cv_nonindiv_sim_nonindiv_finetune refers to signals processed by the FT-JNF DNN (trained with generic data augmentation and generic fine-tuning)- output_cv_nonindiv_sim_indiv_finetune refers to signals processed by the FT-JNF DNN (trained with generic data augmentation and personalized fine-tuning)- output_cv_indiv_sim_indiv_finetune refers to signals processed by the FT-JNF DNN (trained with personalized data augmentation and personalized fine-tuning) Listening examples (online) Selected examples can also be played directly in the browser here: https://m-ohlenbusch.github.io/subjective_ovr_personalized/ Change history: 2025-11-20: Changed folder names to be consistent.



