lowry02/prova
收藏资源简介:
该数据集是一个包含隐藏状态和标注信息的NLP数据集,用于训练、验证和测试模型。数据集中每个样本具有多个特征,包括数据集样本ID、数据集划分(如训练集、验证集、测试集)、层索引、token位置、token ID、token标签和逻辑标签,以及一个表示隐藏状态的浮点数列表。数据集分为训练集(351,242个样本)、验证集(39,890个样本)和测试集(39,892个样本),总大小约为24.6 MB。
This dataset is an NLP dataset containing hidden states and annotation information, designed for training, validation, and testing models. Each sample in the dataset includes multiple features such as dataset sample ID, dataset split (e.g., train, validation, test), layer index, token position, token ID, token label, logic label, and a list of floating-point numbers representing hidden states. The dataset is divided into a training set (351,242 samples), a validation set (39,890 samples), and a test set (39,892 samples), with a total size of approximately 24.6 MB.




