相关数据集
Pushshift Reddit
该数据集是从选定的一些提供正式问答讨论形式的Reddit子版块中精心筛选出来的,用于收集批评意见。它包含了多种筛选机制,以确保批评意见的质量和相关性。规模上,我们选择了16个最适合的Reddit子版块。该数据集的任务是生成批评意见以及回应的细化。
arXiv450
vdaita/gsm8k_mini_gemma_answer_2_train
--- dataset_info: features: - name: question dtype: string - name: answer dtype: string - name: id dtype: int64 - name: initial_answer dtype: string splits: - name: train
Hugging Face2024-05-27 更新50
316usman/thematic1b
--- license: bsd dataset_info: features: - name: text dtype: string - name: thematic dtype: string - name: sub-thematic dtype: string - name: country dtype: string - name:
Hugging Face2024-01-02 更新30
abbassix/ComNumPlus
--- dataset_info: features: - name: original dtype: string - name: label dtype: int64 - name: char dtype: string - name: sci_10E dtype: string - name: sci_10E_char dtyp
Hugging Face2024-01-04 更新140
CATIE-AQ/frenchNLI
这是一个由多个法语自然语言推理子数据集组成的集合,包含前提文本、假设文本和对应的标签信息,分为训练集、验证集和测试集,共569,895条记录。适用于文本分类任务,尤其是自然语言推理领域。
Hugging Face2025-07-29 更新40



