filtered_LWLLM_combined_data
收藏资源简介:
该数据集包含三个主要特征:id(字符串类型)、conversations(列表类型,包含role和content,均为字符串类型)和text(字符串类型)。数据集分为三个部分:训练集(包含327754个样本,大小为1008157466.9685072字节)、验证集(包含100个样本,大小为112606字节)和测试集(包含100个样本,大小为112606字节)。数据集的总下载大小为164804219字节,总大小为1008382678.9685072字节。数据集配置为default,数据文件路径分别为data/train-*、data/valid-*和data/test-*。
This dataset comprises three core attributes: id (string type), conversations (a list containing two string-type fields, role and content), and text (string type). The dataset is split into three subsets: the training set with 327,754 samples and a size of 1008157466.9685072 bytes, the validation set with 100 samples and a size of 112606 bytes, and the test set with 100 samples and a size of 112606 bytes. The total download size of the entire dataset is 164804219 bytes, and the total cumulative size is 1008382678.9685072 bytes. The dataset is configured under the default setting, and its data file paths are data/train-*, data/valid-*, and data/test-* respectively.




