遇见数据集

finepdfs-10M

收藏
魔搭社区2026-04-28 更新2026-07-15 收录
官方服务:

资源简介:

## Sampling Methodology This dataset was created using **reservoir sampling**, a statistically unbiased random sampling algorithm that guarantees each sample from the source dataset has an equal probability of being included. This ensures the 10M token sample is representative of the full dataset's characteristics. **Source Dataset**: [HuggingFaceFW/finepdfs](https://huggingface.co/datasets/HuggingFaceFW/finepdfs) **Sample Size**: 10M tokens **Content**: High-quality textbook-style pdfs Reservoir sampling enables rapid experimentation and ablation studies without processing the entire source dataset, while maintaining statistical validity of results. For details on how this dataset was used in optimal pre-training data composition research, see the [blog post](https://huggingface.co/blog/codelion/optimal-dataset-mixing/). ## Citation If you use this model/dataset, please cite: ```bibtex @article{sharma2025billion, title={The 1 Billion Token Challenge: Finding the Perfect Pre-training Mix}, author={Sharma, Asankhaya}, year={2025}, url={https://huggingface.co/blog/codelion/optimal-dataset-mixing/} } ``` For more details, see the [blog post](https://huggingface.co/blog/codelion/optimal-dataset-mixing/).

提供机构:
maas
创建时间:
2025-10-22
二维码
社区交流群
二维码
科研交流群
商业服务