遇见数据集

bergson-wikitext-512-chunks

收藏
魔搭社区2026-08-03 更新2026-08-09 收录
官方服务:

资源简介:

# bergson-wikitext-512-chunks Wikitext-2 (`Salesforce/wikitext`, `wikitext-2-raw-v1`) pre-chunked into 512-GPT-2-token rows for training-data-attribution experiments with [bergson](https://github.com/EleutherAI/bergson), replicating the data setup of the MAGIC paper (Ilyas & Engstrom 2025, arXiv:2504.16430): each row is one attribution unit / query. - **train**: first 4,608 chunks of the concatenated, GPT-2-tokenized wikitext-2 train split (empty rows dropped before concatenation). - **test**: first 256 chunks of the wikitext-2 test split, same procedure. Each row's `text` field decodes/re-tokenizes to exactly 512 GPT-2 tokens. Built by `scripts/build_chunked_wikitext.py` in the bergson repo.

提供机构:
maas
创建时间:
2026-07-09
二维码
社区交流群
二维码
科研交流群
商业服务