遇见数据集

Leonardo6/gpt-2-crosscoder-diff

收藏
Hugging Face2024-12-02 更新2024-12-14 收录
官方服务:

资源简介:

该数据集包含100,000个训练样本,每个样本由一系列整数(int32类型)表示,存储在名为input_ids的特征中。数据集的总大小为410,000,000字节,下载大小为100,361,137字节。数据集的训练分割存储在路径为data/train-*的文件中。

This dataset contains 100,000 training examples, each represented by a sequence of integers (int32 type) stored in a feature named input_ids. The total size of the dataset is 410,000,000 bytes, with a download size of 100,361,137 bytes. The training split of the dataset is stored in files with the path data/train-*.

提供机构:
Leonardo6
二维码
社区交流群
二维码
科研交流群
商业服务