x_dataset_70
收藏资源简介:
该数据集是Bittensor Subnet 13去中心化网络的一部分,包含从X(原Twitter)预处理的数据。数据由网络矿工持续更新,提供实时的推文流,适用于各种分析和机器学习任务。数据集主要为英文,但也可能包含多语言内容。每个实例代表一条推文,包含推文内容、标签、使用的标签、发布日期、编码的用户名和编码的URL等字段。数据集没有固定的分割,用户应根据需求和数据的时间戳创建自己的分割。数据集的创建遵循X平台的条款和API使用指南,用户名和URL均被编码以保护用户隐私。用户应注意数据中可能存在的社会影响和偏见,以及数据质量和时效性等方面的限制。数据集根据MIT许可证发布,使用时需遵守X的使用条款。
This dataset is a component of the Bittensor Subnet 13 decentralized network, comprising preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, delivering real-time tweet streams suitable for a wide range of analytical and machine learning tasks. The dataset is primarily in English, though it may also contain multilingual content. Each entry corresponds to a single tweet, with fields including tweet content, tags, used tags, publish date, encoded usernames, encoded URLs, and additional relevant information. There are no predefined data splits; users should create custom splits based on their specific requirements and the data's timestamps. The dataset was constructed in compliance with X's Terms of Service and API usage guidelines, with both usernames and URLs encoded to safeguard user privacy. Users should be mindful of potential social impacts and biases present in the dataset, as well as limitations pertaining to data quality and timeliness. The dataset is released under the MIT License, and users must adhere to X's Terms of Service when utilizing the dataset.
Bittensor Subnet 13 X (Twitter) Dataset
数据集描述
- 仓库地址: LadyMia/x_dataset_70
- 子网: Bittensor Subnet 13
- 矿工热键: 5GBQoeC7P5VdbsHiSQUGTCah82kwMyZvsj9uxFDQJa7MANAS
数据集概述
该数据集是Bittensor Subnet 13去中心化网络的一部分,包含来自X(原Twitter)的预处理数据。数据由网络矿工持续更新,提供实时推文流,适用于各种分析和机器学习任务。
支持的任务
该数据集的多功能性允许研究人员和数据科学家探索社交媒体动态的各个方面,并开发创新应用。用户可以利用这些数据进行以下任务:
- 情感分析
- 趋势检测
- 内容分析
- 用户行为建模
语言
主要语言:数据集主要是英语,但由于去中心化的创建方式,可能包含多语言内容。
数据集结构
数据实例
每个实例代表一条推文,包含以下字段:
数据字段
text(字符串): 推文的主要内容。label(字符串): 推文的情感或主题类别。tweet_hashtags(列表): 推文中使用的标签列表。如果没有标签,则为空。datetime(字符串): 推文的发布日期。username_encoded(字符串): 用户名的编码版本,以保护用户隐私。url_encoded(字符串): 推文中包含的URL的编码版本。如果没有URL,则为空。
数据分割
该数据集持续更新,没有固定的分割。用户应根据其需求和数据的时间戳创建自己的分割。
数据集创建
源数据
数据收集自X(Twitter)上的公开推文,遵守平台的条款服务和API使用指南。
个人信息和敏感信息
所有用户名和URL均经过编码以保护用户隐私。数据集不包含故意包含的个人或敏感信息。
使用数据的注意事项
社会影响和偏见
用户应注意X(Twitter)数据中固有的潜在偏见,包括人口统计和内容偏见。该数据集反映了X上表达的内容和意见,不应被视为一般人口的代表性样本。
局限性
- 由于数据收集和预处理的去中心化性质,数据质量可能有所不同。
- 数据集可能包含社交平台常见的噪声、垃圾邮件或无关内容。
- 由于实时收集方法,可能存在时间偏差。
- 数据集仅限于公开推文,不包括私人账户或直接消息。
- 并非所有推文都包含标签或URL。
附加信息
许可信息
该数据集在MIT许可下发布。使用该数据集还需遵守X的使用条款。
引用信息
如果您在研究中使用此数据集,请按如下方式引用:
@misc{LadyMia2024datauniversex_dataset_70, title={The Data Universe Datasets: The finest collection of social media data the web has to offer}, author={LadyMia}, year={2024}, url={https://huggingface.co/datasets/LadyMia/x_dataset_70}, }
贡献
如需报告问题或贡献数据集,请联系矿工或使用Bittensor Subnet 13的治理机制。
数据集统计
- 总实例数: 61999518
- 日期范围: 2024-12-05T00:00:00Z 至 2024-12-12T00:00:00Z
- 最后更新时间: 2024-12-12T08:02:20Z
数据分布
- 带标签的推文: 42.68%
- 不带标签的推文: 57.32%
前10个标签
| 排名 | 主题 | 总数 | 百分比 |
|---|---|---|---|
| 1 | NULL | 34582222 | 56.65% |
| 2 | #tiktok | 227473 | 0.37% |
| 3 | #騎士aリプ返24時間 | 166082 | 0.27% |
| 4 | #riyadh | 158974 | 0.26% |
| 5 | #ad | 148641 | 0.24% |
| 6 | #bbkingvivian | 121098 | 0.20% |
| 7 | #apma2024 | 113335 | 0.19% |
| 8 | #冬もピッコマでポイ活 | 103341 | 0.17% |
| 9 | #مجلس_الصياهد | 78705 | 0.13% |
| 10 | #pr | 76366 | 0.13% |
更新历史
| 日期 | 新增实例数 | 总实例数 |
|---|---|---|
| 2024-12-05T05:59:36Z | 954263 | 954263 |
| 2024-12-05T06:00:00Z | 1313009 | 2267272 |
| 2024-12-08T19:44:55Z | 30692032 | 32959304 |
| 2024-12-12T08:02:20Z | 29040214 | 61999518 |




