Cleaned Lang-8 Dataset
收藏资源简介:
Cleaned Lang-8 Dataset是由VIT Bhopal University的研究团队创建的一个用于语法错误检测的高质量数据集。该数据集包含200,000条经过严格清理的句子对,其中一列包含语法错误的句子,另一列包含相应的修正版本。数据集的创建过程包括多个清理步骤,如去除相似句子、文本归一化、去除多余空格、大小写转换、处理缩写、去除标点符号以及基于Levenshtein距离的过滤等。该数据集主要用于训练和评估语法错误检测模型,旨在提高自然语言处理系统对语法错误的识别和纠正能力,特别适用于第二语言学习者的写作辅助工具。
Cleaned Lang-8 Dataset is a high-quality dataset for grammatical error detection, developed by a research team from VIT Bhopal University. This dataset contains 200,000 strictly cleaned sentence pairs, where one column holds grammatically incorrect sentences and the other contains their corresponding corrected versions. The dataset creation process includes multiple cleaning steps, such as removing similar sentences, text normalization, eliminating redundant spaces, case conversion, handling abbreviations, removing punctuation, and filtering based on Levenshtein distance. This dataset is mainly used for training and evaluating grammatical error detection models, aiming to enhance the grammatical error recognition and correction capabilities of natural language processing systems, and is particularly suitable for writing assistance tools for second language learners.




