StyloBench
收藏资源简介:
StyloBench是一个用于评估个性化机器生成文本(MGT)检测器鲁棒性的基准数据集。该数据集由文学和博客文本及其LLM生成的模仿文本组成,旨在研究现有MGT检测方法在个性化场景下的性能。数据集包含两个子集,分别代表文学作品和博客的个性化场景。StyloBench涵盖了两种子场景:Stylo-Literary和Stylo-Blog。Stylo-Literary子集包含来自七位著名作家的文学作品,而Stylo-Blog子集则包含来自Blog-1K数据集的博客文章。该数据集用于评估现有MGT检测方法在个性化场景下的性能,并揭示了现有检测器在个性化场景下的性能下降和反转现象。
StyloBench is a benchmark dataset for evaluating the robustness of personalized machine-generated text (MGT) detectors. It consists of original literary and blog texts paired with their LLM-generated imitative counterparts, with the goal of investigating the performance of existing MGT detection methods in personalized scenarios. The dataset includes two subsets corresponding to the personalized scenarios of literary works and blog posts, namely Styl-Literary and Styl-Blog. The Styl-Literary subset contains literary works from seven renowned writers, while the Styl-Blog subset comprises blog articles sourced from the Blog-1K dataset. This benchmark is used to evaluate the performance of existing MGT detection methods in personalized scenarios, and it reveals the performance degradation and reversal phenomena of current detectors under such scenarios.

- 1通过MBZUAI, ByteDance, National University of Singapore, Wuhan University · 2025年



