sts
收藏资源简介:
Luganda STS Benchmark 是原始语义文本相似度(STS)基准的卢干达语改编版本,旨在评估语言模型对卢干达语句子对之间语义相似性的理解能力。该数据集通过手工改编原始STS基准的测试集构建,当前版本包含约169个卢干达语句子对,均来自原始STS基准的测试分割。开发集(dev)正在扩展中,训练集(train)尚未包含。每个样本包含两个卢干达语句子,模型需要判断它们在语义上的相似程度。该基准强调语义理解而非词汇或句法匹配,两个句子可能使用不同的词汇或语法结构但表达相同含义,也可能共享相似词汇但表达不同含义。该数据集可用于评估卢干达语句子嵌入模型、语义相似度模型、跨语言表示模型、信息检索系统、多语言语言模型、自然语言理解系统以及卢干达语特定语言模型。改编过程注重保留句子对之间的语义关系,而非逐字翻译,因此可能会在措辞、语法结构或词汇选择上进行调整。当前版本规模较小,应视为早期开发版本,评估结果需结合基准大小进行解读。未来版本将提供更广泛的语义关系和语言变体覆盖。
The Luganda STS Benchmark is an adaptation of the original Semantic Textual Similarity (STS) benchmark into Luganda, designed to evaluate the ability of language models to understand semantic similarity between sentence pairs in Luganda. The dataset is constructed by manually adapting the test set of the original STS benchmark, currently containing approximately 169 Luganda sentence pairs, all from the test split of the original STS benchmark. The development set (dev) is being expanded, and the training set (train) is not yet included. Each sample consists of two Luganda sentences, and the model needs to judge their semantic similarity. This benchmark emphasizes semantic understanding rather than lexical or syntactic matching; two sentences may use different vocabulary or grammar structures but express the same meaning, or share similar vocabulary but express different meanings. The dataset can be used to evaluate Luganda sentence embedding models, semantic similarity models, cross-lingual representation models, information retrieval systems, multilingual language models, natural language understanding systems, and Luganda-specific language models. The adaptation process focuses on preserving the semantic relationships between sentence pairs rather than word-for-word translation, so adjustments may be made in wording, grammatical structure, or vocabulary choice. The current version is small in scale and should be considered an early development version; evaluation results should be interpreted in the context of the benchmark size. Future versions will provide broader coverage of semantic relationships and language variants.
Luganda STS Benchmark 数据集概述
基本信息
- 许可证:Apache-2.0
- 任务类别:句子相似度(sentence-similarity)
- 语言:Luganda(卢干达语)
- 数据集规模:少于1,000个样本(n<1K)
数据集简介
该数据集是语义文本相似度(STS)基准的Luganda语改编版本,旨在评估语言模型在理解Luganda语句对之间语义相似性方面的能力。数据集源自原始的STS Benchmark,并正在人工改编和扩充中,其目标是评估语义理解能力,而非简单的词汇或句法匹配。
当前状态
- 初始版本:包含约200对来自原始STS Benchmark测试集的句子对
- 开发状态:数据集正在积极扩充中,后续将增加更多测试集样本,并逐步扩充开发集
| 数据划分 | 当前状态 |
|---|---|
| 测试集 | 初始Luganda改编版本,约200个样本(README后文提及约130个样本) |
| 开发集 | 待扩充 |
| 训练集 | 暂未包含 |
任务说明
每个样本包含一对Luganda语句,模型需要判断这两个句子在语义上传达相同含义的程度。该基准旨在测试语义理解能力,而非简单的词汇重叠——两个句子可能使用不同的词汇或语法结构却表达实质上相同的语义,而词汇相似的句子也可能表达不同的含义。
预期用途
该基准可用于评估以下模型或系统:
- Luganda语句嵌入模型
- 语义相似度模型
- 跨语言表示模型
- 信息检索系统
- 多语言语言模型
- 自然语言理解系统
- Luganda专用语言模型
此外,它也可用于评估Luganda语言建模的改进是否能带来更好的语义表征。
数据集开发方法
当前测试集是基准的初始发布版本,后续将从原始STS Benchmark中逐步改编并扩充更多样本。改编的重点是保留语句对之间的语义关系,而非逐字翻译。因此,忠实的语义改编可能需要在Luganda语中改变措辞、语法结构或词汇选择。
局限性
- 当前发布版本规模较小,应视为早期开发版本,而非完整的Luganda STS基准
- 测试集目前包含约130个样本,开发集仍在扩展中
- 在此版本上的评估结果应结合基准规模进行解读
- 未来版本预计将提供更广泛的语义关系和Luganda语言变异覆盖
归属与引用
该数据集是STS Benchmark的Luganda改编版本,原始基准源自SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation。
原始STS Benchmark引用
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., & Specia, L. (2017). SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation. Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), 1–14.
Luganda改编版本引用
Matovu Caleb (2026). Luganda STS Benchmark. SAGE POND. Luganda adaptation and expansion of the STS Benchmark.
致谢
该基准建立在STS Benchmark及相关研究者对原始语义文本相似度评估资源贡献的基础之上。Luganda改编和持续扩展工作由SAGE POND的Caleb Matovu开发。




