Fair-Speech
收藏资源简介:
Fair-Speech数据集由Meta AI创建,旨在评估语音识别系统在不同人口统计群体中的公平性。该数据集包含约26,500条由593名美国参与者录制的语音命令,涵盖年龄、性别、种族、地理和语言背景等多个维度。数据集的创建过程中,参与者自我报告了他们的社会人口统计信息,并提供了自然语言的语音命令。该数据集主要用于评估和改进语音识别模型在不同群体中的表现,以促进技术的公平性和包容性。
The Fair-Speech dataset was created by Meta AI to evaluate the fairness of speech recognition systems across diverse demographic groups. It contains approximately 26,500 speech commands recorded by 593 U.S. participants, covering multiple dimensions such as age, gender, race, geographic background, and linguistic background. During the dataset creation process, participants self-reported their socio-demographic information and provided natural language speech commands. This dataset is primarily used to evaluate and improve the performance of speech recognition models across different groups, so as to promote technological fairness and inclusivity.




