遇见数据集

Fediverse HWT & AIGT Corpus

收藏
Zenodo2026-02-09 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is a large-scale Fediverse (Mastodon) social media corpus designed for benchmarking AI-Generated Text (AIGT) detection and related community-aware analyses. It contains a balanced composition of: Human-Written Text (HWT) collected from the pre-LLM era, and AI-Generated Text (AIGT) produced with state-of-the-art Large Language Models (LLMs). A key feature of this release is the explicit preservation of community context: posts are drawn from 263 distinct Mastodon instances, where each instance is treated as a “community” with its own norms, topics, and moderation policies. This structure supports community-aware modeling, cross-instance generalization studies, and analyses of linguistic variation in decentralized social networks.

提供机构:
Zenodo
创建时间:
2026-02-09
二维码
社区交流群
二维码
科研交流群
商业服务