遇见数据集

SalatielJordao/orkut-communities

收藏
Hugging Face2026-05-22 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含从存档的Orkut社区页面提取的文本内容,源自互联网档案馆/Archive Team的Orkut WARC和CDX文件,并转换为分层Parquet格式供研究使用。数据集专注于人类撰写的文本字段,包括社区名称和描述、论坛主题元数据、用户名、回复日期、回复标题和回复正文,但不包括图像和其他二进制资源。每一行代表一个Orkut社区论坛主题或线程,回复作为嵌套列表存储在其父线程行中。数据集结构包括顶层线程字段(如社区ID、名称、描述、URL、时间戳等)和嵌套回复字段(如回复ID、作者ID、名称、日期、正文等)。数据规模较大,包含约1.206亿个线程行和约8.973亿个回复。数据集适用于历史网络研究、计算社会科学、语言建模研究、网络存档分析和数据集保存研究等用途,但需注意其中可能包含用户生成的敏感内容,并遵守相关法律和伦理规范。

This dataset contains text extracted from archived Orkut community pages. It is derived from Internet Archive / Archive Team Orkut WARC and CDX files and converted into a hierarchical Parquet dataset for research use. The dataset focuses on human-authored text fields found in community and forum pages, including community names and descriptions, forum topic metadata, usernames, reply dates, reply titles, and reply bodies. Images and other binary assets are not included. Each row represents one forum topic/thread. Replies are stored as a nested list inside their parent thread row. The dataset structure includes top-level thread fields (e.g., community_id, community_name, community_description, community_url, community_timestamp) and nested reply fields (e.g., reply_id, author_id, author_name, date, body). It contains approximately 120.64 million thread rows and 897.28 million replies. The dataset is suitable for uses such as historical web research, computational social science, language modeling research, web-archive analysis, and dataset preservation studies, but users should be aware that it may contain user-generated sensitive content and must comply with legal and ethical standards.

提供机构:
SalatielJordao
二维码
社区交流群
二维码
科研交流群
商业服务