xavierdurawa/proof-pile-2-streaming
收藏资源简介:
--- task_categories: - text-generation language: - en tags: - math size_categories: - 10B<n<100B --- <img src="proofpile_logo.jpg" width="500"> [ArXiv](http://arxiv.org/abs/2310.10631) | [Models](https://huggingface.co/EleutherAI/llemma_34b) | [Data](https://huggingface.co/datasets/EleutherAI/proof-pile-2) | [Code](https://github.com/EleutherAI/math-lm) | [Blog](https://blog.eleuther.ai/llemma/) | [Sample Explorer](https://llemma-demo.github.io/) [Zhangir Azerbayev](https://zhangir-azerbayev.github.io/), [Hailey Schoelkopf](https://github.com/haileyschoelkopf), [Keiran Paster](https://keirp.com), [Marco Dos Santos](https://github.com/dsantosmarco), [Stephen McAleer](https://www.andrew.cmu.edu/user/smcaleer/), [Albert Q. Jiang](https://albertqjiang.github.io/), [Jia Deng](https://www.cs.princeton.edu/~jiadeng/), [Stella Biderman](https://www.stellabiderman.com/), [Sean Welleck](https://wellecks.com/) The **Proof-Pile-2** is a 55 billion token dataset of mathematical and scientific documents. This dataset was created in order to train the [Llemma 7B](https://huggingface.co/EleutherAI/llemma_7b) and [Llemma 34B](https://huggingface.co/EleutherAI/llemma_34b) models. It consists of three subsets: - `arxiv` (29B tokens): the ArXiv subset of [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T) - `open-web-math` (15B tokens): The [OpenWebMath](https://huggingface.co/datasets/open-web-math/open-web-math) dataset, which contains much of the high-quality mathematical text from the internet. - `algebraic-stack` (11B tokens): A new dataset of mathematical code, including numerical computing, computer algebra, and formal mathematics. You can download the dataset as follows ```python from datasets import load_dataset ds = load_dataset("EleutherAI/proof-pile-2") # To load only a specific subset, pass it as an argument, e.g ds_arxiv = load_dataset("EleutherAI/proof-pile-2", "arxiv") ``` ### Schema Each dataset row has the following structure ```python { "text": ..., # document text "meta": ..., # JSON string of metadata, schema specific to data source } ``` ### Dataset Contents For detailed documentation of the ArXiv and web subsets, refer to [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T) and [OpenWebMath](https://huggingface.co/datasets/open-web-math/open-web-math). The following table enumerates the contents of the AlgebraicStack by programming language. The AlgebraicStack is filtered to only include documents that contain mathematics, as judged by hand-crafted, language-specific heuristics. | Language | AlgebraicStack tokens | |-----------|-----------------------| | Agda | 35.2 M | | C | 25.1 M | | C++ | 954.1 M | | Coq | 281.9 M | | Fortran | 724.9 M | | GAP | 3.6 M | | Haskell | 9.1 M | | Idris | 10.9 M | | Isabelle | 1,089.7 M | | Julia | 531.0 M | | Jupyter | 199.1 M | | Lean | 285.6 M | | Maple | 2.0 M | | Matlab | 65.8 M | | Python | 6,098.8 M | | R | 71.3 M | | Tex | 567.7 M | | **Total** | **10,955.7 M** | ### License We do not alter the license of any of the underlying data. ### Version History **v1.1.0**: Contains an updated version of OpenWebMath, precisely the one available at [open-web-math/open-web-math](https://huggingface.co/datasets/open-web-math/open-web-math). This version of OpenWebMath has slightly improved filtering, for example, removal of very short documents. **v1.0.0**: The data used to train the [Llemma 7B](https://huggingface.co/EleutherAI/llemma_7b) and [Llemma 34B](https://huggingface.co/EleutherAI/llemma_34b). Uses a development version of OpenWebMath. ### Citation For the entire Proof-Pile-2, cite ``` @misc{azerbayev2023llemma, title={Llemma: An Open Language Model For Mathematics}, author={Zhangir Azerbayev and Hailey Schoelkopf and Keiran Paster and Marco Dos Santos and Stephen McAleer and Albert Q. Jiang and Jia Deng and Stella Biderman and Sean Welleck}, year={2023}, eprint={2310.10631}, archivePrefix={arXiv}, primaryClass={cs.CL} } ``` For the ArXiv subset, cite ``` @software{together2023redpajama, author = {Together Computer}, title = {RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset}, month = April, year = 2023, url = {https://github.com/togethercomputer/RedPajama-Data} } ``` For OpenWebMath, cite ``` @misc{paster2023openwebmath, title={OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text}, author={Keiran Paster and Marco Dos Santos and Zhangir Azerbayev and Jimmy Ba}, year={2023}, eprint={2310.06786}, archivePrefix={arXiv}, primaryClass={cs.AI} } ```
任务类别:文本生成 语言:英语 标签:数学 规模类别:100亿 < Token数 < 1000亿 <img src="proofpile_logo.jpg" width="500"> [ArXiv](http://arxiv.org/abs/2310.10631) | [Models](https://huggingface.co/EleutherAI/llemma_34b) | [Data](https://huggingface.co/datasets/EleutherAI/proof-pile-2) | [Code](https://github.com/EleutherAI/math-lm) | [Blog](https://blog.eleuther.ai/llemma/) | [Sample Explorer](https://llemma-demo.github.io/) [Zhangir Azerbayev](https://zhangir-azerbayev.github.io/), [Hailey Schoelkopf](https://github.com/haileyschoelkopf), [Keiran Paster](https://keirp.com), [Marco Dos Santos](https://github.com/dsantosmarco), [Stephen McAleer](https://www.andrew.cmu.edu/user/smcaleer/), [Albert Q. Jiang](https://albertqjiang.github.io/), [Jia Deng](https://www.cs.princeton.edu/~jiadeng/), [Stella Biderman](https://www.stellabiderman.com/), [Sean Welleck](https://wellecks.com/) **Proof-Pile-2**是一个包含550亿Token的数学与科学文档数据集。本数据集旨在训练[Llemma 7B](https://huggingface.co/EleutherAI/llemma_7b)与[Llemma 34B](https://huggingface.co/EleutherAI/llemma_34b)模型。数据集包含三个子集: - `arxiv`(290亿Token):[RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T)的ArXiv子集 - `open-web-math`(150亿Token):[OpenWebMath](https://huggingface.co/datasets/open-web-math/open-web-math)数据集,该数据集收录了互联网上大量高质量数学文本 - `algebraic-stack`(110亿Token):全新的数学代码数据集,涵盖数值计算、计算机代数与形式化数学内容。 可通过以下方式下载该数据集: python from datasets import load_dataset ds = load_dataset("EleutherAI/proof-pile-2") # 若仅需加载特定子集,可传入对应参数,例如: ds_arxiv = load_dataset("EleutherAI/proof-pile-2", "arxiv") ### 数据模式 每条数据集样本遵循以下结构: python { "text": ..., # 文档文本 "meta": ..., # 元数据JSON字符串,格式由数据源决定 } ### 数据集内容 关于ArXiv与网页子集的详细文档,请参考[RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T)与[OpenWebMath](https://huggingface.co/datasets/open-web-math/open-web-math)。下表按编程语言枚举了AlgebraicStack的内容。AlgebraicStack经过过滤,仅保留经人工编写的语言专属启发式规则判定为包含数学内容的文档。 | 编程语言 | AlgebraicStack Token数 | |----------|-----------------------| | Agda | 3520万 | | C | 2510万 | | C++ | 9.541亿 | | Coq | 2.819亿 | | Fortran | 7.249亿 | | GAP | 360万 | | Haskell | 910万 | | Idris | 1090万 | | Isabelle | 10.897亿 | | Julia | 5.310亿 | | Jupyter | 1.991亿 | | Lean | 2.856亿 | | Maple | 200万 | | Matlab | 6580万 | | Python | 60.988亿 | | R | 7130万 | | Tex | 5.677亿 | | **总计** | **109.557亿** | ### 授权协议 我们未修改任何原始数据的授权协议。 ### 版本历史 **v1.1.0**:包含更新版OpenWebMath,即[open-web-math/open-web-math](https://huggingface.co/datasets/open-web-math/open-web-math)当前可用版本。该版本的OpenWebMath优化了过滤逻辑,例如移除了过短的文档。 **v1.0.0**:用于训练[Llemma 7B](https://huggingface.co/EleutherAI/llemma_7b)与[Llemma 34B](https://huggingface.co/EleutherAI/llemma_34b)的数据集,使用了开发版OpenWebMath。 ### 引用 若引用完整的Proof-Pile-2数据集,请使用以下BibTeX条目: bibtex @misc{azerbayev2023llemma, title={Llemma: An Open Language Model For Mathematics}, author={Zhangir Azerbayev and Hailey Schoelkopf and Keiran Paster and Marco Dos Santos and Stephen McAleer and Albert Q. Jiang and Jia Deng and Stella Biderman and Sean Welleck}, year={2023}, eprint={2310.10631}, archivePrefix={arXiv}, primaryClass={cs.CL} } 若引用ArXiv子集,请使用: bibtex @software{together2023redpajama, author = {Together Computer}, title = {RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset}, month = April, year = 2023, url = {https://github.com/togethercomputer/RedPajama-Data} } 若引用OpenWebMath数据集,请使用: bibtex @misc{paster2023openwebmath, title={OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text}, author={Keiran Paster and Marco Dos Santos and Zhangir Azerbayev and Jimmy Ba}, year={2023}, eprint={2310.06786}, archivePrefix={arXiv}, primaryClass={cs.AI} }
数据集概述
Proof-Pile-2 是一个包含 550 亿个标记的数学和科学文档数据集。该数据集旨在训练 Llemma 7B 和 Llemma 34B 模型。它由三个子集组成:
arxiv(290 亿个标记): ArXiv 子集,来自 RedPajamaopen-web-math(150 亿个标记): OpenWebMath 数据集,包含大量高质量的互联网数学文本。algebraic-stack(110 亿个标记): 一个新的数学代码数据集,包括数值计算、计算机代数和形式数学。
数据集加载
可以使用以下代码下载数据集: python from datasets import load_dataset ds = load_dataset("EleutherAI/proof-pile-2")
仅加载特定子集,例如 arxiv
ds_arxiv = load_dataset("EleutherAI/proof-pile-2", "arxiv")
数据集结构
每个数据集行具有以下结构: python { "text": ..., # 文档文本 "meta": ..., # 元数据的 JSON 字符串,模式特定于数据源 }
数据集内容
详细文档请参考 RedPajama 和 OpenWebMath。以下表格列举了 AlgebraicStack 按编程语言的内容:
| 语言 | AlgebraicStack 标记数 |
|---|---|
| Agda | 35.2 M |
| C | 25.1 M |
| C++ | 954.1 M |
| Coq | 281.9 M |
| Fortran | 724.9 M |
| GAP | 3.6 M |
| Haskell | 9.1 M |
| Idris | 10.9 M |
| Isabelle | 1,089.7 M |
| Julia | 531.0 M |
| Jupyter | 199.1 M |
| Lean | 285.6 M |
| Maple | 2.0 M |
| Matlab | 65.8 M |
| Python | 6,098.8 M |
| R | 71.3 M |
| Tex | 567.7 M |
| 总计 | 10,955.7 M |
许可证
我们不更改任何基础数据的许可证。
版本历史
- v1.1.0: 包含 OpenWebMath 的更新版本,改进了过滤,例如移除非常短的文档。
- v1.0.0: 用于训练 Llemma 7B 和 Llemma 34B 的数据。
引用
对于整个 Proof-Pile-2,引用:
@misc{azerbayev2023llemma, title={Llemma: An Open Language Model For Mathematics}, author={Zhangir Azerbayev and Hailey Schoelkopf and Keiran Paster and Marco Dos Santos and Stephen McAleer and Albert Q. Jiang and Jia Deng and Stella Biderman and Sean Welleck}, year={2023}, eprint={2310.10631}, archivePrefix={arXiv}, primaryClass={cs.CL} }
对于 ArXiv 子集,引用:
@software{together2023redpajama, author = {Together Computer}, title = {RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset}, month = April, year = 2023, url = {https://github.com/togethercomputer/RedPajama-Data} }
对于 OpenWebMath,引用:
@misc{paster2023openwebmath, title={OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text}, author={Keiran Paster and Marco Dos Santos and Zhangir Azerbayev and Jimmy Ba}, year={2023}, eprint={2310.06786}, archivePrefix={arXiv}, primaryClass={cs.AI} }




