mteb/MultiEupSlovakGenderClassification
收藏资源简介:
MultiEupSlovakGenderClassification是一个用于文本分类的数据集,属于MTEB(大规模文本嵌入基准)的一部分。该数据集专注于二元分类任务,旨在根据欧洲议会议员在Multi-EuP v2语料库中的斯洛伐克语原生演讲,预测议员的性别。数据仅包含最初以斯洛伐克语发表的演讲,涉及政府和口语领域。数据集包含训练集(508个样本)和测试集(128个样本),每个样本包括文本和标签(0或1,代表不同性别)。标签分布显示训练集中标签1有435个样本,标签0有73个样本;测试集中标签1有109个样本,标签0有19个样本。文本长度平均约1017个字符(训练集)和1107个字符(测试集)。该数据集基于源数据集unimelb-nlp/MultiEup-v2派生而来,适用于评估嵌入模型在斯洛伐克语性别分类任务上的性能。
MultiEupSlovakGenderClassification is a dataset for text classification, part of the MTEB (Massive Text Embedding Benchmark). It focuses on a binary classification task to predict the gender of Members of the European Parliament from native Slovak speeches in the Multi-EuP v2 corpus. The data includes only speeches originally delivered in Slovak, covering government and spoken domains. The dataset consists of a training set (508 examples) and a test set (128 examples), each with text and label (0 or 1, representing different genders). Label distribution shows 435 samples with label 1 and 73 with label 0 in the training set, and 109 with label 1 and 19 with label 0 in the test set. Average text length is approximately 1017 characters for training and 1107 for test. Derived from the source dataset unimelb-nlp/MultiEup-v2, it is designed for evaluating embedding models on Slovak gender classification tasks.




