dancinlab/hexa-forge-corpus-stack-v2-sample-v0.1.3
收藏资源简介:
这是一个确定性5%样本,来自bigcode/the-stack-v2的permissive子集,覆盖hexa-forge核心编程语言:Python、Rust、TypeScript、Go、C和Zig。通过tool/stack_v2_sample.py生成,选择基于repo/path的稳定BLAKE2b哈希(种子42),因此样本可从相同源快照中复现。它是papers/datasets-source-manifest.md中STRUCT行pretrain-bias的伴生数据集,用于通过直接测量采样子集并外推至100%来验证约600B令牌的估计。
Deterministic 5% sample of the bigcode/the-stack-v2 permissive subset across the hexa-forge core languages: Python, Rust, TypeScript, Go, C, Zig. Produced by tool/stack_v2_sample.py ; selection is by stable BLAKE2b hash of repo/path (seed 42), so the sample is reproducible from the same source snapshot. Companion to papers/datasets-source-manifest.md STRUCT row pretrain-bias. Used to validate the ~600B-token estimate by direct measurement on the sampled subset and extrapolation to 100%.



