遇见数据集

jihuny/llama3.2-1b_nectar_10000_triplet_l2_armo

收藏
Hugging Face2026-03-27 更新2026-03-29 收录
官方服务:

资源简介:

--- dataset_info: features: - name: prompt dtype: string - name: prompt_idx dtype: int32 - name: response_a dtype: string - name: response_b dtype: string - name: response_a_idx dtype: int32 - name: response_b_idx dtype: int32 - name: logprob_a dtype: float32 - name: logprob_b dtype: float32 - name: embedding_difference list: float32 length: 2048 - name: score_a dtype: float64 - name: score_b dtype: float64 splits: - name: train num_bytes: 5127172776 num_examples: 450000 download_size: 1461589400 dataset_size: 5127172776 configs: - config_name: default data_files: - split: train path: data/train-* ---

数据集信息: 特征字段: - 提示词(prompt):数据类型为字符串 - 提示词索引(prompt_idx):数据类型为32位整型 - 回复A(response_a):数据类型为字符串 - 回复B(response_b):数据类型为字符串 - 回复A索引(response_a_idx):数据类型为32位整型 - 回复B索引(response_b_idx):数据类型为32位整型 - 回复A对数概率(logprob_a):数据类型为32位浮点型 - 回复B对数概率(logprob_b):数据类型为32位浮点型 - 嵌入向量差异(embedding_difference):为长度2048的32位浮点型列表 - 回复A评分(score_a):数据类型为64位浮点型 - 回复B评分(score_b):数据类型为64位浮点型 数据集划分: - 训练集(train):字节大小为5127172776,样本数量为450000 下载大小:1461589400字节,数据集总大小:5127172776字节 配置项: - 默认配置(default):对应训练集划分的数据文件,数据存储路径为data/train-*

提供机构:
jihuny
二维码
社区交流群
二维码
科研交流群
商业服务