SkMTEB
收藏资源简介:
SkMTEB是由夸美纽斯大学·布拉迪斯拉发等机构联合创建的首个斯洛伐克语大规模文本嵌入基准,旨在填补低资源西斯拉夫语言在语义表示评估领域的空白。该基准包含31个数据集,涵盖检索、重排序、分类等七类任务,数据来源包括新闻、政府文件、社交媒体及百科全书等多领域文本,时间跨度从2000年至2025年,其中包含七个全新构建的本地化数据集。基准通过整合现有语料库与人工标注流程构建,特别针对斯洛伐克语的语言特性进行优化,主要应用于评估嵌入模型在语义搜索、文本聚类及跨语言检索等任务中的性能,为低资源语言的模型适配与效率优化提供关键基础设施。
SkMTEB is the first large-scale Slovak text embedding benchmark jointly created by Comenius University Bratislava and other institutions, aiming to fill the gap in semantic representation evaluation for low-resource West Slavic languages. This benchmark includes 31 datasets covering seven types of tasks such as retrieval, reranking, classification and more. Its data sources cover multi-domain texts including news, government documents, social media, encyclopedias and other materials, with a time span from 2000 to 2025, and it contains seven newly constructed localized datasets. Built by integrating existing corpora and manual annotation workflows, the benchmark is specially optimized for the linguistic characteristics of Slovak. It is mainly used to evaluate the performance of embedding models on tasks like semantic search, text clustering and cross-lingual retrieval, providing key infrastructure for model adaptation and efficiency optimization of low-resource languages.

- 1SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation夸美纽斯大学·布拉迪斯拉发; 思科系统; 科希策技术大学; 肯佩伦智能技术研究所 · 2026年



