遇见数据集

istiaqfuad/bangla-english-banglish-pairs

收藏
Hugging Face2026-05-19 更新2026-05-31 收录
官方服务:

资源简介:

该数据集提供孟加拉语、英语和Banglish(孟加拉语-英语混合语)三语种的对比训练对,用于微调句子嵌入模型(如BGE-M3),旨在增强模型对Banglish拼写变体的鲁棒性。数据集结合了两个主要来源:1. LLM生成的Banglish拼写变体;2. OPUS-100 EN-BN平行语料库。它包括三个文件:合并的训练集(包含Banglish和EN↔BN对,带有source标签)、清理后的EN↔BN对(从OPUS-100生成)和清理后的Banglish对(通过LLM生成的拼写变体)。数据集结构包括锚点句、正例句和负例句,支持对比学习,适用于特征提取和句子相似性任务。

This dataset provides contrastive training pairs across three languages: Bengali, English, and Banglish (a mixed Bengali-English language), for fine-tuning sentence embedding models such as BGE-M3, with the aim of enhancing the model's robustness against Banglish spelling variations. The dataset combines two primary sources: 1. Banglish spelling variants generated by LLMs; 2. the OPUS-100 EN-BN parallel corpus. It includes three files: the merged training set (containing Banglish and EN↔BN pairs with source tags), the cleaned EN↔BN pairs (generated from OPUS-100), and the cleaned Banglish pairs (spelling variants generated via LLMs). The dataset structure consists of anchor sentences, positive examples and negative examples, which supports contrastive learning and is suitable for feature extraction and sentence similarity tasks.

提供机构:
istiaqfuad
二维码
社区交流群
二维码
科研交流群
商业服务