遇见数据集

ChavyvAkvar/MegaMath-1M-Sample-288

收藏
Hugging Face2025-09-26 更新2025-10-25 收录
官方服务:

资源简介:

该数据集包含多个字段,如文本内容(text)、唯一标识符(id)、文件路径(cc-path)、领域(domain)、数学分数(finemath_score)、语言(lang)、语言分数(lang_score)、时间戳(timestamp)和网址(url)等。数据集被划分为训练集,共有100万条示例,大小为4,769,954,445字节。数据集的具体应用场景和内容未在README中描述。

The dataset includes multiple fields such as text content (text), unique identifier (id), file path (cc-path), domain (domain), math score (finemath_score), language (lang), language score (lang_score), timestamp (timestamp), and URL (url). The dataset is split into a training set with a total of 1,000,000 examples, totaling 4,769,954,445 bytes in size. The specific application scenario and content of the dataset are not described in the README.

提供机构:
ChavyvAkvar
二维码
社区交流群
二维码
科研交流群
商业服务