polyglot-tagger/nlp-noise-snippets
收藏官方服务:
资源简介:
--- license: apache-2.0 task_categories: - text-classification size_categories: - 100K<n<1M --- # Synthetic Noise Pool For Text Classification purposes, as many models may consider code snippets, html artifacts, and math as "English". Around 50K are latex snippets from `im2latex-100k`
许可证:Apache 2.0 任务类别: - 文本分类 规模类别: - 10万<样本数<100万 # 合成噪声池(Synthetic Noise Pool) 本数据集旨在服务于文本分类任务,因诸多模型可能将代码片段、HTML 残留标记以及数学表达式误判为“英语”。 其中约5万条样本为来自`im2latex-100k`的LaTeX片段。
提供机构:
polyglot-tagger


