遇见数据集

token-counts

收藏
魔搭社区2026-07-19 更新2026-07-19 收录
官方服务:

资源简介:

# Marin Token Counts Token counts for all datasets used in [Marin](https://github.com/marin-community/marin) pretraining runs. ## Schema | Column | Type | Description | |--------|------|-------------| | `dataset` | string | Dataset identifier | | `marin_tokens` | int | Number of tokens after tokenization | | `category` | string | Content domain (web, code, math, academic, books, etc.) | | `synthetic` | bool | Whether the data is LLM-generated or LLM-translated | ## Categories - **web** — Quality-classified Common Crawl text (Nemotron-CC) - **code** — Source code and code-related documents - **math** — Math-focused extractions and competition problems - **academic** — Peer-reviewed papers and abstracts - **reasoning** — Cross-domain reasoning and formal logic - **books** — Digitized public domain and open access books - **legal** — Court decisions, regulations, patents - **government** — Parliamentary proceedings and publications - **education** — Open educational resources and textbooks - **encyclopedic** — Wiki-style reference content - **forum** — Q&A sites and chat logs - **documents** — PDF-extracted document text - **translation** — Parallel translation corpora - **news** — CC-licensed news articles - **media** — Transcribed audio/video - **supervised** — Curated task datasets - **reference** — Niche reference sites - **general** — General-domain content ## Updates This dataset is updated by running `experiments/count_tokens.py` from the Marin repo, which reads tokenized dataset stats from GCS and pushes the results here.

提供机构:
maas
创建时间:
2026-03-31
二维码
社区交流群
二维码
科研交流群
商业服务