Arabic Semantic Textual Similarity Benchmark
收藏官方服务:
资源简介:
This dataset is the Arabic version of the Semantic Textual Similarity Benchmark (Cer et al., 2017). It consists of sentence pairs collected from diverse sources, including news headlines, video and image captions, and natural language inference data. Each pair is originally annotated by human judges with a similarity score ranging from 1 to 5; in this variant, these scores are normalized to a continuous scale between 0 and 1, making the dataset suitable for training and evaluating semantic similarity and sentence embedding models.
提供机构:
Zenodo创建时间:
2026-01-07



