遇见数据集

Arabic Semantic Textual Similarity Benchmark

收藏
Zenodo2026-01-30 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is the Arabic version of the Semantic Textual Similarity Benchmark (Cer et al., 2017). It consists of sentence pairs collected from diverse sources, including news headlines, video and image captions, and natural language inference data. Each pair is originally annotated by human judges with a similarity score ranging from 1 to 5; in this variant, these scores are normalized to a continuous scale between 0 and 1, making the dataset suitable for training and evaluating semantic similarity and sentence embedding models.

提供机构:
Zenodo
创建时间:
2026-01-07
二维码
社区交流群
二维码
科研交流群
商业服务