遇见数据集

Anonymous1789/grokset

收藏
Hugging Face2026-05-08 更新2026-05-31 收录
官方服务:

资源简介:

@GrokSet是首个从公共社交媒体收集的大规模多参与者人类-LLM互动数据集。与现有语料库(如WildChat、LMSYS-Chat-1M)不同,这些语料库捕捉的是私人的、二元(一对一)的用户-助手互动,而@GrokSet捕捉的是大型语言模型Grok在X(前身为Twitter)上作为多用户线程中的公共参与者的行为。数据集时间跨度为2025年3月至2025年10月,覆盖了超过1百万条推文和18.2万+个对话线程。它旨在研究LLM在对抗性、社会嵌入和“公共广场”环境中的行为。数据集以脱水格式(推文ID + 注释 + 结构元数据)发布,以符合平台服务条款,并提供了专门的水合工具包来重建数据集的文本和元数据。关键特点包括:多参与者动态(捕捉复杂的互动图,而不仅仅是线性查询)、真实世界上下文(包括互动指标如点赞、转发、回复以衡量社会验证)和丰富注释(包括预计算的毒性(Detoxify)、主题(BERTopic)、网络指标(中心性、传递性)等标签)。

@GrokSet is the first large-scale dataset of multi-party human–LLM interactions collected from public social media. Unlike existing corpora (e.g., WildChat, LMSYS-Chat-1M) that capture private, dyadic (one-on-one) user-assistant interactions, @GrokSet captures the Grok Large Language Model acting as a public participant in multi-user threads on X (formerly Twitter). The dataset spans from March to October 2025, covering over 1 million tweets across 182,000+ conversation threads. It is designed to study the behavior of LLMs in adversarial, socially embedded, and public square environments. This dataset is released in a dehydrated format (Tweet IDs + annotations + structural metadata) to comply with platform ToS. A specialized rehydration toolkit is provided to reconstruct the datasets text and metadata. Key Features: Multi-Party Dynamics (captures complex interaction graphs, not just linear queries), Real-World Context (includes engagement metrics like likes, reposts, replies to measure social validation), and Rich Annotations (includes pre-computed labels for Toxicity (Detoxify), Topics (BERTopic), and Network Metrics (Centrality, Transitivity)).

提供机构:
Anonymous1789
二维码
社区交流群
二维码
科研交流群
商业服务