遇见数据集

chaannwooff/CC-korean

收藏
Hugging Face2026-05-08 更新2026-05-31 收录
官方服务:

资源简介:

CC-korean是一个韩语网络文本语料库,从CommonCrawl的WET文件(去除HTML的文本格式)中提取。数据收集覆盖了15个快照(CC-MAIN-2025-05至CC-MAIN-2026-12,对应2025年至2026年),使用fasttext的lid.176.bin模型进行韩语检测,仅收集置信度高于0.5的文档。初步质量过滤包括启发式行级噪声去除和KenLM困惑度过滤。收集后,应用了额外的过滤步骤:行级内联去除(如URL、表情符号、括号标签、记者署名、来源标注等)和行级去除(如非韩语行、UI/导航残留、日期/元标签、地址、商业信息、版权文本、广告/赌博/成人关键词、论坛标题、分页等),总计去除约4300万行。文档级过滤从3,695,378个输入文档中筛选出3,570,675个通过(通过率96.6%)。数据集规模为约36GB,包含3,570,675个文档。列包括:url(原始文档URL)、date(WET文件收集日期)、text(过滤后的正文文本)、char_count(字符数)、line_count(行数)、lang_conf(语言检测置信度)和license(许可证信息)。数据集基于Common Crawl使用条款,适用于研究目的,商业使用需获得原始版权所有者的许可。

CC-korean is a Korean web text corpus extracted from the WET files (HTML-stripped text format) of Common Crawl. The data collection covers 15 snapshots (CC-MAIN-2025-05 to CC-MAIN-2026-12, corresponding to 2025 to 2026). The fasttext lid.176.bin model was employed for Korean language detection, and only documents with a confidence score higher than 0.5 were collected. Preliminary quality filtering includes heuristic line-level noise removal and KenLM perplexity filtering. After collection, additional filtering steps were applied: inline line-level removal (such as URLs, emojis, bracket tags, journalist bylines, source annotations, etc.) and line-level removal (such as non-Korean lines, UI/navigation residues, date/meta tags, addresses, commercial information, copyright text, advertising/gambling/adult keywords, forum titles, pagination, etc.), with a total of approximately 43 million lines removed. For document-level filtering, 3,570,675 documents passed the screening out of 3,695,378 input documents, resulting in a pass rate of 96.6%. The dataset has a total size of approximately 36 GB and contains 3,570,675 documents. Its columns are as follows: url (original document URL), date (WET file collection date), text (filtered main body text), char_count (number of characters), line_count (number of lines), lang_conf (language detection confidence score), and license (license information). The dataset is governed by the Common Crawl Terms of Use, and is permitted for research purposes only; commercial use requires prior permission from the original copyright holders.

提供机构:
chaannwooff
二维码
社区交流群
二维码
科研交流群
商业服务