官方服务:
资源简介:
roberta large mega set
应用场景:
创建时间:
2021-07-14
相关数据集
AIME2025-long-hints
该数据集包含了问题、答案、完整句子及其正确性、领域、注释、有提示和无提示情况下的准确率等字段。它被设计用于训练模型,其中的训练集包含了30个示例,数据集总大小为2,152,843字节。
Hugging Face2025-04-27 更新170
mlfoundations-dev/instruction_filtering_fasttext_seed_data_code
该数据集包含了对话指令种子(instruction_seed)、来源(source)、以及对应的70B蒸馏模型响应(r1_distill_70b_response)等字段。此外,每个样本还有一个原始行索引(__original_row_idx)和会话信息(conversations),其中会话信息包含发送者(from)和消息内容(value)。数据集被拆分为训练集(train),包含2500个示例
Hugging Face2025-02-14 更新50
The Finnish Sub-corpus of the JRC-Acquis Multilingual Parallel Corpus, Downloadable Version
This is the legal subcorpus of the [Helsinki Korp Version of the Finnish TreeBank 3](http://urn.fi/urn:nbn:fi:lb-2016042602). The corpus is available for online browsing through the concordancer Korp
SSH Open MarketPlace2025-07-04 更新20
RylanSchaeffer/collapse_gemma-2-9b_hs2_accumulate_iter2_sftsd0_temp1_max_seq_len512
该数据集包含两个主要特征:prompt(提示)和response(响应),均为字符串类型。数据集包含一个训练集(train),其中包含12531个样本,总大小为15358934字节。下载大小为904870字节,数据集总大小为15358934字节。数据集的配置文件指定了默认配置,数据文件路径为data/train-*。
Hugging Face2024-11-24 更新60
AnanthZeke/tamil_sentences_sample
该数据集名为tamil_combined_sentences,包含泰米尔语的句子,主要用于句子相似性和零样本分类任务。数据集来源于OSCAR和Wikipedia,大小在1M到10M之间。数据集的特征包括句子,分割为训练集,包含2,391,475个例子,总大小为1,164,550,978字节。
Hugging Face2023-10-13 更新30



