遇见数据集

openeurollm/contaminated-documents

收藏
Hugging Face2025-11-25 更新2026-01-03 收录
官方服务:

资源简介:

--- dataset_info: config_name: nemotron_sample features: - name: warc_record_id dtype: string - name: file_part dtype: string - name: benchmark dtype: string - name: matched_ngram dtype: string - name: benchmark_text dtype: string - name: train dtype: string splits: - name: single_collision num_bytes: 189307298 num_examples: 7428 - name: all_collisions num_bytes: 1044885141 num_examples: 32841 download_size: 212808419 dataset_size: 1234192439 configs: - config_name: nemotron_sample data_files: - split: single_collision path: nemotron_sample/single_collision-* - split: all_collisions path: nemotron_sample/all_collisions-* --- This repository will include the contaminated documents from Nemotron and HPLT, extracted using nemo-curator. The benchmarks are obtained from [here](https://docs.google.com/spreadsheets/d/1uBji9fJFLdaaOnzYI71RGW2IiuHJ2vrmuHsXPIsx_rk/edit?gid=1345345034#gid=1345345034), and use the split defined for benchmarking by lm-evaluation-harness

数据集信息: 配置名称:nemotron_sample 特征: - 字段名:warc_record_id,数据类型:字符串 - 字段名:file_part,数据类型:字符串 - 字段名:benchmark,数据类型:字符串 - 字段名:matched_ngram,数据类型:字符串 - 字段名:benchmark_text,数据类型:字符串 - 字段名:train,数据类型:字符串 数据拆分: - 拆分名称:single_collision,字节数:189307298,样本数:7428 - 拆分名称:all_collisions,字节数:1044885141,样本数:32841 下载大小:212808419 数据集总大小:1234192439 配置项: - 配置名称:nemotron_sample,数据文件: - 拆分:single_collision,路径:nemotron_sample/single_collision-* - 拆分:all_collisions,路径:nemotron_sample/all_collisions-* 本仓库收录由 nemo-curator 从 Nemotron 与 HPLT 中提取的污染文档。 本数据集所用基准集获取自[此处](https://docs.google.com/spreadsheets/d/1uBji9fJFLdaaOnzYI71RGW2IiuHJ2vrmuHsXPIsx_rk/edit?gid=1345345034#gid=1345345034),并采用 lm-evaluation-harness 为基准测试定义的数据拆分规则。

提供机构:
openeurollm
二维码
社区交流群
二维码
科研交流群
商业服务