kylelovesllms/hi_hf_frames_d4_100_heldoutdepth_4
收藏资源简介:
该数据集包含多个字段,用于处理文本或序列数据,具体包括:frame_id(帧标识符)、depth(深度值)、hi(初始文本)、hf(最终文本)、hi_bracketed(带括号的初始文本)、hf_bracketed(带括号的最终文本)和n_tokens(令牌数量)。数据集分为训练集(8964个示例)、验证集(4800个示例)和测试集(4800个示例),总大小约为14.6MB,适用于机器学习模型(如自然语言处理模型)的训练和评估任务。
This dataset includes multiple fields for processing text or sequence data, specifically: frame_id (frame identifier), depth (depth value), hi (initial text), hf (final text), hi_bracketed (initial text with brackets), hf_bracketed (final text with brackets), and n_tokens (number of tokens). The dataset is divided into training set (8,964 examples), validation set (4,800 examples), and test set (4,800 examples), with a total size of approximately 14.6MB, suitable for machine learning model training and evaluation tasks, such as natural language processing models.



