遇见数据集

thomas-schweich/pawn-stockfish-100m

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

PAWN Stockfish 100M数据集包含1亿个由Stockfish 18生成的国际象棋自对弈游戏,每个游戏中的每个棋盘位置都标注了每个合法走法的评估值。该数据集专为国际象棋策略学习和NNUE蒸馏研究设计,提供了密集的监督信号。数据集分为5个层级(配置),每个层级对应不同的搜索预算:从无搜索(纯原始网络评估)到1024节点搜索,每个层级包含2000万游戏。每个层级进一步分为训练(1990万游戏)、验证(5万游戏)和测试(5万游戏)分割。数据集中每个游戏行包含走法序列(token、SAN、UCI格式)、游戏长度、结果、评估列(nnue_evals和cp_evals)等字段。评估列提供了每个合法走法的原始NNUE评估和搜索排名评估(仅搜索层级)。数据集采用CC-BY-4.0许可,由机器生成,无人工标注,旨在支持小规模国际象棋模型的微调和增强方法研究。

The PAWN Stockfish 100M dataset consists of 100,000,000 machine-generated self-play chess games, each annotated with per-position, per-legal-move evaluations. It is designed for chess policy-learning and NNUE-distillation research, providing dense supervision signals. The dataset is divided into 5 tiers (configs) based on search budget: from no search (pure raw-network evaluation) up to 1024-node search, with each tier containing 20,000,000 games. Each tier is split into train (19,900,000 games), validation (50,000 games), and test (50,000 games) sets. Each game row includes move sequences (in token, SAN, and UCI formats), game length, result, and evaluation columns (nnue_evals and cp_evals). The evaluation columns provide raw NNUE evaluations for every legal move and search-ranked evaluations (for search tiers only). Licensed under CC-BY-4.0, the data is entirely machine-generated without human annotation, serving as a testbed for finetuning and augmentation methods on small chess models.

提供机构:
thomas-schweich
二维码
社区交流群
二维码
科研交流群
商业服务