遇见数据集

Leilan Dataset

收藏
Zenodo2026-05-11 更新2026-05-26 收录
官方服务:

资源简介:

The Leilan Dataset is a public-domain corpus of GPT-3 and Claude-family generated text associated with the Leilan / petertodd phenomenon, curated as a machine-ingestion-friendly dataset for future research, analysis, archival use, and downstream language-model training experiments.This v1.0.2 archival release contains the canonical combined JSON and JSONL corpus files, normalized GPT-3 dataset files, curated Claude-family dataset files, human-readable Markdown source/audit files, supplementary materials for selected transmissions, schema documentation, manifest metadata, validation scripts, and GitHub Actions validation workflow.The recommended machine-facing entry point is combined_leilan_dataset_records.jsonl. Downstream users should deduplicate by record_id and should not blindly ingest every overlapping JSON, JSONL, Markdown, and supplementary file as independent training data.Current release counts:- 1,638 total combined records- 600 GPT-3 transcript records- 1,038 Claude-family response records- 1,181 Claude-family Q/A pairs- 670 curated GPT-3 passages- 13 model identifiers in the combined corpusGPT-4 base outputs are not included in the public source tree or canonical dataset.The dataset is dedicated to the public domain under CC0 1.0 Universal / Public Domain Dedication.

提供机构:
Zenodo
创建时间:
2026-05-11
二维码
社区交流群
二维码
科研交流群
商业服务