GPTS-Crawler-Dataset
收藏资源简介:
该数据集包含10000多个GPTs的详细元数据,通过从互联网多个来源(如Google搜索、Github)爬取获得。数据集以gizmos.jsonl文件形式提供,每行包含一个完整GPTs的元数据,如ID、作者信息、显示名称、描述、类别、标签等。仓库支持爬取、去重和持续更新数据集,旨在构建公开的GPTS数据集资源。
This dataset contains detailed metadata for over 10,000 GPTs, collected via web crawling from multiple online sources including Google Search and GitHub. The dataset is provided in the format of a gizmos.jsonl file, where each line stores the complete metadata of a single GPT, such as ID, author information, display name, description, category, tags and other relevant attributes. The associated repository supports dataset crawling, deduplication and continuous updates, aiming to build a publicly available GPTs dataset resource.
数据集概述
数据集名称:GPTS-Crawler-Dataset
数据规模:包含超过10000个GPTs的元数据。
数据格式:文件 gizmos.jsonl,每一行是一个完整的GPTs元数据,采用JSON格式。
关键字段说明:
id:GPTs的唯一标识符。organization_id:所属组织ID。short_url:短链接。author:作者信息,包括用户ID、显示名称、链接、验证状态等。display:显示信息,包括名称、描述、欢迎消息、提示启动词、头像URL、分类等。share_recipient:分享接收方(如marketplace)。updated_at:更新时间。tags:标签(如public,reportable)。
数据来源:
- 通过互联网多来源爬取。
- 支持通过Google搜索和GitHub获取GPTs数据。
更新方式:
- 新获取的GPTs元数据会追加到
gizmos.jsonl文件中。 - 支持断点续传和异常重试。
辅助功能:
- 提供去重脚本(
deduplicate-urls和deduplicate-gpts)。 - 支持从Google和GitHub批量获取GPTs URL。
相关资源:
贡献方式:
- 在GitHub Issue中提交GPTs URL。
- 直接更新
gpts-url-list或gizmos.jsonl文件。 - 添加Google搜索关键词。
致谢数据源:
- Hybird AI search: https://github.com/memfreeme/memfree
- gpts-works: https://github.com/all-in-aigc/gpts-works
- gptshunter issue: https://github.com/airyland/gptshunter.com/issues/1
- GPTHub: https://github.com/lencx/GPTHub/blob/main/gpthub.json





