遇见数据集

CFPR accent metadata

收藏
Zenodo2026-03-12 更新2026-06-05 收录
官方服务:

资源简介:

To address the current lack of dataset of diverse regional variations of French, we have extended the Corpus du Français Parlé de nos Régions (CFPR) with explicit accent labels. This extension provides a new metadata to develop and assess speech technology and study phonetic diversity of the French-speaking world.CFPR is an open-access resource, developed at Sorbonne Université (Avanzi et al., 2019), designed to study regional and social linguistic variation. It consists of audio recordings from 186 speakers, each providing a single recording. The original dataset is rich in metadata, documenting: Demographics (Gender and year of birth), Geography (Country/region of origin, birthplace, and current residence) and other context (Recording date, location, and the speaker’s level of French). These recordings, conducted as interviews, span a vast global reach—from Ivory Coast to Algeria and mainland France—making the corpus an ideal foundation for accent analysis. The gold-standard labels were established through a concerted labeling process. Two persons, representing different regional backgrounds (Central and Southern France) listened and annotated all recordings collaboratively, achieving consensus after multiple passes. They had a posteriori access to the full speaker metadata (childhood and current locations at time of recording) to resolve ambiguous cases and contacted a third person in a couple of cases.The resulting dataset features 186 annotated recordings categorized into eight distinct regional classes: Canadian French, Caribbean French, Central France, Eastern & Northern France, North Africa, Pacific and Indian Area, Southern France, Sub-Saharan Africa.

提供机构:
Zenodo
创建时间:
2026-03-11
二维码
社区交流群
二维码
科研交流群
商业服务