相关数据集
google_wellformed_query
GoogleWellformedQuery专注于识别自然语言问题的规范性,其核心在于判断搜索查询是否为符合语法的、明确的且无拼写错误的提问。数据集包含从Paralex语料库中抽取的25100条查询,每条查询都由五位众包工作者进行标注,最终得到一个0到1之间的评分,代表查询的规范程度。该数据集支持文本分类和文本评分任务,并采用CC BY-SA 4.0授权许可。其中训练集包含17500条数据,验证集包
Opencsg2024-07-19 更新90
Datasets used in this work and their main characteristics.
Columns LTr, LVa, U contain the numbers of tweets in the training set, held-out validation set, and test set, respectively. Column “Shift” contains the values of distribution shift between L ≡ LTr ⋃ L
NIAID Data Ecosystem50
konwoo/dclm-train-1M-shuffled
--- dataset_info: features: - name: bff_contained_ngram_count_before_dedupe dtype: int64 - name: language_id_whole_page_fasttext struct: - name: en dtype: float64 - name: met
Hugging Face2025-12-13 更新30
TTTXXX01/DPO_Orz-30K_filtered
该数据集包含消息内容、角色、以及数据集名称等信息。消息内容包括文本内容和发送者的角色。此外,数据集还提供了每个示例的真相或正确答案。训练集包含3000个示例,总大小为1,458,969字节。
Hugging Face2025-10-13 更新60
lygonzal/dataset_sentimiento_manual
--- dataset_info: features: - name: texto dtype: string - name: label dtype: string splits: - name: train num_bytes: 2004 num_examples: 50 download_size: 2143 dataset_siz
Hugging Face2026-03-23 更新40



