遇见数据集

gpt-2-tensorflow2.0 training data (BPE model and TFRecords): preservation copy of a non-reproducible set

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

A preservation copy of the training data that gpt-2-tensorflow2.0's train_gpt2.py reads: a byte-pair-encoding model and 86 TFRecord files of token-ID sequences. It is not a text corpus and not a trained model. gpt2-data.tar.gz unpacks to data/bpe_model.model, data/bpe_model.vocab and data/tf_records/ (86 files), produced by the project's pre_process.py from the 633 web pages it ships under data/scraped/. sha256 8ae7f61ef57b688102f85f6dc35353a34fd3901af6c727944ea7cdf39c893206, 47,687,802 bytes. gpt2-data.sha256 lists each file's sha256. Why a copy: the set cannot be regenerated. Each pre_process.py run retrains the BPE model and writes two new timestamp-named TFRecords, and two runs over the same input share no TFRecord and not even the BPE model. This set accumulated over runs on 2024-08-15, 2025-03-14 and 2026-08-01, and train_gpt2.py reads all of it. It is deposited so that research that measures this program on exactly these files can be rebuilt. Terms: the files were produced by the MIT-licensed gpt-2-tensorflow2.0 code (Copyright (c) 2019 Abhay Singh) from its OpenWebText sample. Copyright in the underlying articles remains with their authors and publishers; this deposit grants no rights in that text and makes no license claim over it. See NOTICE. Anyone wanting the text should go to its sources.

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务