遇见数据集

anonymous-eandd-2026/HainaWeb-Sci-sample

收藏
Hugging Face2026-05-04 更新2026-05-31 收录
官方服务:

资源简介:

HainaWeb-Sci样本是一个科学网络语料库的代表性样本,用于NeurIPS 2026的双盲同行评审。数据集覆盖14个STEM学科,包括航空航天工程、农业、天文学等。数据来源于Common Crawl快照(2013年至2026年3月)和DCLM-Pool,经过数据中心的处理流程,包括启发式过滤、两层级内容驱动的学科路由、基于评分标准的质量评分以及去重。数据以gzip压缩的JSONL格式存储,每个JSON对象包含文本字段和元数据,如原始URI、检测语言、学科分布和质量分数。个人身份信息已使用标准方法进行掩码处理。样本仅用于评审目的,完整语料库将在非匿名位置发布。许可证为Open Data Commons Attribution License (ODC-By) v1.0,并遵循Common Crawl使用条款。

HainaWeb-Sci Sample is a representative sample of a scientific web corpus, released anonymously for double-blind peer review at NeurIPS 2026. It covers 14 STEM disciplines, including AerospaceEngineering, Agriculture, Astronomy, etc. Data is sourced from Common Crawl snapshots (2013–March 2026) and DCLM-Pool, processed through a data-centric curation pipeline involving heuristic filtering, two-tier discipline routing, rubric-based quality scoring, and deduplication. Documents are stored as gzip-compressed JSONL shards, with each JSON object containing text and metadata fields such as original URI, detected language, discipline distribution, and quality scores. Personally identifiable information has been masked. The sample is intended solely for reviewer evaluation, with the full corpus to be released at a non-anonymous location. It is licensed under ODC-By v1.0 and subject to Common Crawl terms of use.

提供机构:
anonymous-eandd-2026
二维码
社区交流群
二维码
科研交流群
商业服务