遇见数据集

AxiomResearch/synthetic-dataset-v3

收藏
Hugging Face2026-05-08 更新2026-05-31 收录
官方服务:

资源简介:

Axiopedia V1 是一个合成的教育数据集,使用 openai/gpt-oss-20b 模型生成。它包含 185,300 行数据,总大小为 711.2MB,内容混合比例为 70% 的教育主题(如物理等)和 30% 的编码示例(如 Python 等)。数据集提供了详细的列描述,包括唯一标识符、文本内容、领域、类型、字数统计、字符数统计、令牌计数、生成时间戳、使用模型和来源等信息,旨在支持机器学习和教育相关应用。

Axiopedia V1 is a synthetic educational dataset generated using the openai/gpt-oss-20b model. It contains 185,300 rows with a total size of 711.2MB, consisting of a mix of 70% educational content (e.g., physics) and 30% coding examples (e.g., Python). The dataset includes detailed column descriptions such as unique identifier, text content, domain, type, word count, character count, token count, generation timestamp, model used, and source, designed to support machine learning and educational applications.

提供机构:
AxiomResearch
二维码
社区交流群
二维码
科研交流群
商业服务