AxiomResearch/synthetic-dataset-quality90plus
收藏资源简介:
Axiopedia v3 — Quality 90+ 是一个包含50,593篇长篇幅教育文章的数据集,经过严格质量过滤,质量评分≥90/100。该数据集专为持续预训练设计,而非简单的合成数据堆砌。文章平均长度约1,180词,中位数1,213词,涵盖算法、系统、法律、经济学、物理、数学、历史、生物学等50个多样领域。数据集通过生成、精确去重、MinHash LSH去重、质量评分和过滤流程构建,确保无重复内容。生成模型为openai/gpt-oss-20b,采用随机种子领域采样方法,最小生成长度为1,000个令牌。质量评分基于长度、词汇多样性、重复性、结构和清洁度等启发式标准。
Axiopedia v3 — Quality 90+ is a dataset of 50,593 long-form educational articles, filtered for high quality with scores ≥90/100. It is designed to be actually usable for continued pretraining, not just another synthetic dump. Articles have a median length of 1,200 words and cover 50 diverse domains including algorithms, systems, law, economics, physics, mathematics, history, and biology. The dataset is built through a pipeline of generation, exact deduplication, MinHash LSH deduplication, quality scoring, and filtering, ensuring 0% duplicates. It was generated using the openai/gpt-oss-20b model with random-seed domain sampling and a minimum of 1,000 tokens. Quality is assessed via a heuristic rubric based on length, lexical diversity, repetition, structure, and cleanliness.




