BUT-FIT/FLiP-data
收藏资源简介:
FLiP-data 是FLiP项目的预处理数据集,用于通过因子化线性投影解释多模态多语言句子嵌入。该数据集基于Mozilla Common Voice v15英语数据,包含SONAR语音嵌入、文本嵌入、语音与文本嵌入之间的余弦相似度分数、参考转录文本以及使用Gemini 2.5 Flash Lite提取的命名实体。数据分为训练集(约100万条语句)、开发集(约1.6万条)和测试集(约1.6万条)。嵌入使用SONAR编码器从Common Voice音频和转录计算,源数据遵循CC BY 4.0许可。数据集支持句子相似性和特征提取任务,旨在促进多模态嵌入的可解释性研究。
FLiP-data is a preprocessed dataset for the FLiP project, which focuses on interpreting multimodal multilingual sentence embeddings via factorized linear projection. It includes SONAR speech embeddings, text embeddings, cosine similarity scores between paired speech and text embeddings, reference transcripts, and named entities extracted with Gemini 2.5 Flash Lite, all based on Mozilla Common Voice v15 English data. The dataset is split into train (~1M utterances), dev (~16k), and test (~16k) sets. Embeddings are computed using the SONAR encoder from Common Voice audio and transcripts, with source data licensed under CC BY 4.0. It is designed for tasks such as sentence similarity and feature extraction to advance interpretability research in multimodal embeddings.




