OpenAlex Human–LLM rewritten abstract dataset
收藏资源简介:
The OpenAlex Human–LLM rewritten abstract dataset is a balanced text classification corpus designed for research on AI-generated text detection in scholarly writing. It contains 143,976 abstract-level text samples, including 71,988 human-written abstracts collected from OpenAlex and 71,988 LLM-rewritten counterparts derived from those abstracts. Each record includes three fields: new_id (a unique identifier), text (the abstract text), and label (binary class label, where 0 denotes the original human-written abstract and 1 denotes the LLM-rewritten version). The dataset is provided in JSONL format and is intended for training, validating, and benchmarking models that detect LLM-mediated rewriting in scientific abstracts. It should be interpreted as a benchmark for distinguishing original scholarly abstracts from their LLM-rewritten versions, rather than as evidence that the original source documents themselves were written with LLM assistance.



