遇见数据集

JingweiNi/train_prm800k_gpt-oss-120b_annotated_qwen3_1.7b_thinking_5000_shards_4_7_32

收藏
Hugging Face2025-12-17 更新2025-12-20 收录
官方服务:

资源简介:

该数据集包含多个特征,如question(问题)、answer(答案)、input_ids(输入标识符)、reply(回复)、original_index(原始索引)、claims(声明)和verified(已验证)。claims特征进一步细分为aligned_token_ids(对齐的令牌标识符)、claim_text(声明文本)和sentence(句子)。数据集包含一个名为train的训练集,共有2,500个样本,总大小为55,706,369字节,下载大小为18,086,051字节。配置文件指定了train分割的数据文件路径。

The dataset includes features such as question, answer, input_ids, reply, original_index, claims, and verified. The claims feature is subdivided into aligned_token_ids, claim_text, and sentence. The dataset has a single train split with 2,500 examples, totaling 55,706,369 bytes in size, and a download size of 18,086,051 bytes. The configuration specifies the data files for the train split.

提供机构:
JingweiNi
二维码
社区交流群
二维码
科研交流群
商业服务