遇见数据集

trentmkelly/looksmaxxing-forum

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含从Looksmaxxing论坛收集的公开论坛内容,用于研究和分析,旨在保留论坛、线程和帖子之间的讨论结构。数据集分为四个部分:forums(20行)包含论坛和子论坛记录,如论坛标题、URL、类别元数据及可见的线程/消息计数;threads(172,917行)包含线程级元数据,如线程标题、源/规范URL、页面数、论坛标识符、首次观察到的论坛页面和任何前缀标签;posts(2,594,973行)包含个别论坛帖子,如线程/页面位置、作者元数据、时间戳、永久链接、反应、附件/提及/引用元数据、提取的文本、HTML和源页面URL;thread_classifications(2,007行)包含部分线程的可选线程级分类输出,如模型/状态元数据、标签、解释、响应JSON和提示使用字段。数据集以zstd压缩的Parquet文件格式分发。

This dataset contains public forum content collected from Looksmaxxing Forum for research and analysis. It is intended to preserve discussion structure across forums, threads, and posts. The dataset includes four splits: forums (20 rows) with forum and subforum records such as forum title, URL, category metadata, and visible thread/message counts; threads (172,917 rows) with thread-level metadata like thread title, source/canonical URLs, page count, forum identifiers, first observed forum page, and prefix labels; posts (2,594,973 rows) with individual forum posts including thread/page position, author metadata, timestamps, permalink, reactions, attachment/mention/quote metadata, extracted text, HTML, and source page URL; thread_classifications (2,007 rows) with optional thread-level classification outputs for a subset of threads, including model/status metadata, labels, explanations, response JSON, and prompt usage fields. The dataset is distributed as zstd-compressed Parquet files.

提供机构:
trentmkelly
二维码
社区交流群
二维码
科研交流群
商业服务