reddit_dataset_212
收藏资源简介:
Bittensor Subnet 13 Reddit数据集是一个包含预处理后的Reddit帖子和评论的数据集,数据持续由网络矿工更新,提供实时的Reddit内容流,适用于各种分析和机器学习任务。数据集以英文为主,但也支持多语言。数据集结构包括文本内容、标签、数据类型、社区名称、日期、用户名编码和URL编码等字段。数据集不包含固定的分割,用户应根据需求和时间戳自行创建分割。数据来源于Reddit的公共帖子和评论,并遵循平台的服务条款和API使用指南。所有用户名和URL都经过编码处理以保护隐私。使用该数据集时,应注意潜在的偏见和限制。
The Bittensor Subnet 13 Reddit Dataset is a curated collection of preprocessed Reddit posts and comments. It is continuously updated by network miners to provide a real-time Reddit content stream, suitable for various analytics and machine learning tasks. The dataset is primarily in English while supporting multilingual content. Its structure includes fields such as text content, labels, data types, community names, dates, encoded usernames, and encoded URLs. The dataset does not include pre-defined splits, and users should create custom splits based on their specific requirements and timestamps. The data is sourced from public Reddit posts and comments, and complies with the platform's Terms of Service and API usage guidelines. All usernames and URLs are encoded to protect user privacy. Users should be aware of potential biases and limitations when utilizing this dataset.




