showerthoughts-dataset
收藏资源简介:
该数据集源自r/Showerthoughts Reddit社区,用于研究大型语言模型在特定领域写作风格适应中的机智、创造力和可检测性。数据集通过Pushshift API服务提供,但由于Reddit在2023年更改了其条款,限制了对数据的访问,因此本仓库不提供原始Reddit数据的下载链接,仅提供AI生成的Showerthoughts数据。
This dataset originates from the r/Showerthoughts Reddit community and is utilized to investigate the wit, creativity, and detectability of large language models in adapting to specific domain writing styles. The dataset is provided via the Pushshift API service. However, due to Reddit's updated terms in 2023, which restrict access to the data, this repository does not offer download links to the original Reddit data but only provides AI-generated Showerthoughts data.
数据集概述
数据集名称
- 名称: showerthoughts-dataset
数据集来源
- 来源: 基于Reddit社区r/Showerthoughts的数据,通过Pushshift API服务获取。
数据集内容
- 内容: 包含AI生成的Showerthoughts数据。
数据获取
- 获取方式: 由于Reddit在2023年更改了数据访问条款,本仓库不提供原始Reddit数据的下载链接。如需获取完整数据集,请参考仓库中的单独README文件。
引用信息
-
引用格式:
@inproceedings{ buz2024investigating, title={Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddits Showerthoughts}, author={Tolga Buz and Benjamin Frost and Nikola Genchev and Moritz Schneider and Lucie-Aimée Kaffee and Gerard de Melo}, booktitle={The 13th Joint Conference on Lexical and Computational Semantics}, year={2024}, url={https://openreview.net/forum?id=VAYdzStvFj} }




