遇见数据集

Dataset to publication: "Neighbors and relatives: How do speech embeddings reflect linguistic connections across the world?"

收藏
Zenodo2025-08-04 更新2026-05-26 收录
官方服务:

资源简介:

Dataset of voxlingua107-xls-r-300m-wav2vec (Alumäe & Kukk, 2022) language identification model embeddings extracted from utterances from the Common Voice 16.1 (Ardila et al., 2020) dataset. Used in "Neighbors and relatives: How do speech embeddings reflect linguistic connections across the world?". Preprint available at https://arxiv.org/abs/2506.08564 Code available at https://github.com/TuukkaOT/speech_embedding_analyzer

提供机构:
Zenodo
创建时间:
2025-08-04
二维码
社区交流群
二维码
科研交流群
商业服务