ganglii/8B_if_n10
收藏资源简介:
该数据集是一个用于自然语言处理任务的数据集,包含2000个训练示例,总大小约49.5 MB。特征包括prompt(提示)、question(问题)、answer(答案),以及多个生成相关字段(如generations、ave_token_logprob、seq_len、seq_logprob)及其对应ground truth版本(如generations_gt、ave_token_logprob_gt),用于评估或比较文本生成模型的输出。数据以训练集形式组织,适用于文本生成、问答或模型评估任务。
This dataset is designed for natural language processing tasks, containing 2000 training examples with a total size of approximately 49.5 MB. Features include prompt, question, answer, and multiple generation-related fields (e.g., generations, ave_token_logprob, seq_len, seq_logprob) along with their ground truth counterparts (e.g., generations_gt, ave_token_logprob_gt), intended for evaluating or comparing text generation model outputs. The data is organized as a training split and is suitable for text generation, question answering, or model evaluation tasks.




