遇见数据集

FaiST: Fairness in Speech Technology

收藏
Zenodo2026-02-20 更新2026-05-29 收录
官方服务:

资源简介:

The FaiST dataset is an English speech corpus composed of natural, conversational audio gathered from speakers in the United States who represent several racial and ethnic groups. Speakers either explicitly self-identify or are identified by other speakers as belonging to a particular demographic group during the recorded interaction. This work curates these signals to support two primary goals: (i) enabling the community to assess the fairness of speech and language technologies under realistic conversational conditions, and (ii) facilitating the development of more equitable systems through training on data whose demographic labels are grounded in how people describe themselves in context. The label includes self-identified race, ethnicity, national origin, type of identification, and the span. Span is the exact line in the conversation where a speaker identifies themselves or another speaker as belonging to a certain demographic group. Type is either self-identification (speaker identifies themselves) or other-person in-interaction (speaker identifies someone else).

提供机构:
Zenodo
创建时间:
2026-02-20
二维码
社区交流群
二维码
科研交流群
商业服务