RyanYr/pg_sais-dapo_shuffled-offline-pg-dapo-qwen3-4B-Base-mbs128-n4_matheval
收藏资源简介:
该数据集包含用于训练和评估AI模型的数据,特别侧重于问题解答和对话生成任务。数据特征包括:数据源(data_source)、问题(problem)、解决方案(solution)、答案(answer)、提示(prompt,含角色和内容字段)、奖励模型信息(reward_model,含真实标签和风格字段)以及响应列表(responses)。数据集分为mixed和hard两种类型,每种类型有多个版本(如240、220等),对应不同的数据子集,其中mixed子集示例数较多(1447个),hard子集示例数较少(100个)。数据可能用于强化学习、自然语言处理模型训练或评估,尤其是在多轮对话和奖励建模场景中。
This dataset contains data for training and evaluating AI models, with a focus on question answering and dialogue generation tasks. Features include: data source (data_source), problem (problem), solution (solution), answer (answer), prompt (with role and content fields), reward model information (reward_model, including ground truth and style fields), and a list of responses (responses). The dataset is divided into mixed and hard types, each with multiple versions (e.g., 240, 220), corresponding to different data subsets. The mixed subsets have a larger number of examples (1447), while the hard subsets have fewer (100). The data is likely used for reinforcement learning, natural language processing model training, or evaluation, particularly in multi-turn dialogue and reward modeling scenarios.




