遇见数据集

ChavyvAkvar/MegaMath-1M-Sample-231

收藏
Hugging Face2025-09-26 更新2025-10-25 收录
官方服务:

资源简介:

这是一个包含文本数据的数据集,每个样本都包括文本内容、唯一标识符、文件路径、域名、数学和语言相关分数、使用的语言、时间戳和URL等信息。数据集分为训练集,共有100万条样本,总数据大小为4.9GB。

This is a dataset containing text data, with each sample including text content, unique identifier, file path, domain, math and language scores, language used, timestamp, and URL information. The dataset is split into a training set with a total of 1 million samples, with a total dataset size of 4.9GB.

提供机构:
ChavyvAkvar
二维码
社区交流群
二维码
科研交流群
商业服务