ALHD
收藏资源简介:
ALHD(阿拉伯语大型语言模型和人类数据集)是一个大规模、多语种和多方言的语料库,旨在区分人类和大型语言模型生成的文本。该数据集跨越三个语种(新闻、社交媒体、评论),涵盖了阿拉伯语和阿拉伯方言,包含超过40万个平衡样本,由三个领先的大型语言模型生成,并来自多个人类来源。ALHD数据集为研究阿拉伯语大型语言模型生成文本检测的可迁移性提供了基础。
ALHD (Arabic Large Language Model and Human Dataset) is a large-scale, multilingual and multi-dialectal corpus designed to distinguish between texts generated by humans and large language models (LLMs). The dataset covers three text domains, namely news, social media and reviews, encompasses both Modern Standard Arabic and various Arabic dialects, and contains over 400,000 balanced samples. These samples are generated by three leading large language models and collected from diverse human sources. The ALHD dataset serves as a foundational resource for research on the transferability of text detection methods for Arabic large language model-generated texts.




