遇见数据集

Arabic Semantic Textual Similarity Benchmark

收藏
Zenodo2026-01-30 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is the Arabic version of the Semantic Textual Similarity Benchmark (Cer et al., 2017). It consists of sentence pairs collected from diverse sources, including news headlines, video and image captions, and natural language inference data. Each pair is originally annotated by human judges with a similarity score ranging from 1 to 5; in this variant, these scores are normalized to a continuous scale between 0 and 1, making the dataset suitable for training and evaluating semantic similarity and sentence embedding models.

本数据集为语义文本相似度基准测试集(Semantic Textual Similarity Benchmark,Cer等人,2017)的阿拉伯语版本。其包含从多样来源采集的句子对,涵盖新闻标题、视频与图像字幕以及自然语言推理数据。每一组句子对最初由人类标注者给出1至5分的相似度评分;在本变体数据集中,这些评分被归一化至0至1的连续区间,使得该数据集适用于语义相似度与句子嵌入模型的训练与评估。

提供机构:
Zenodo
创建时间:
2026-01-07
二维码
社区交流群
二维码
科研交流群
商业服务