官方服务:
资源简介:
:unav
应用场景:
相关数据集
sumyeongahn/sst5
--- dataset_info: features: - name: idx dtype: int64 - name: label dtype: int64 - name: sentence dtype: string - name: paraphrase dtype: string splits: - name: train
Hugging Face2024-06-15 更新400
kdercksen/pragtag_sentence_classification
该数据集是作为Revise and Resubmit工作的一部分提供的,具体来说,是从项目Github页面上的prag.csv文件处理而来,以便于使用HuggingFace模型进行句子分类。数据集中的句子是根据Github仓库中的splits.csv文件进行分割的,分为训练集、验证集和测试集。
Hugging Face2024-02-28 更新140
toramaru-u/cc100-ja-nsp-32
该数据集包含用于训练模型的数据,主要特征包括索引(idx)、下一句标签(next_sentence_label)、句子A(sentence_a)和句子B(sentence_b)。数据集分为训练集,包含127,086,714个示例,总大小为31,150,074,680字节。下载大小为19,813,849,727字节。
Hugging Face2024-06-27 更新120
portuguese-benchmark-datasets/xpaws_pt
--- configs: - config_name: default data_files: - split: test path: data/test-* - split: validation path: data/validation-* dataset_info: features: - name: id dtype: int64 - na
Hugging Face2023-12-26 更新140
WendyHoang/news-shuffled-zh-JIEBA-nsp-test
该数据集包含两个句子(sentence1和sentence2)和一个标签(label),适用于句子对分类任务。训练集包含12736332个样本,测试集包含1415148个样本。
Hugging Face2025-04-13 更新120



