librarian-bots/dataset_cards_with_metadata
收藏资源简介:
该数据集包含了Hugging Face Hub上托管模型的社区创建的dataset cards,这些卡片提供了关于托管在Hugging Face Hub上的数据集的信息。数据集每天更新,包括Hugging Face Hub上公开可用的数据集。数据集的主要用途包括文本挖掘、分析数据集卡片的格式/内容、主题建模以及在数据集卡片上训练语言模型。数据集的结构包括一个单一的分割,数据来源于Hugging Face Hub上的README.md文件,通过CRON作业每天下载。数据集卡片由社区创建,可能包含个人或敏感信息,且可能存在偏见。
This dataset comprises dataset cards created by the community for datasets hosted on the Hugging Face Hub, which provide comprehensive information about those hosted datasets. The dataset is updated daily and includes all publicly available datasets on the Hugging Face Hub. Its primary use cases include text mining, analysis of the format and content of dataset cards, topic modeling, and training language models on the dataset cards. The dataset structure features a single data split. The data is sourced from README.md files on the Hugging Face Hub and downloaded daily via CRON jobs. Since dataset cards are created by the community, they may contain personal or sensitive information, as well as potential biases.
数据集概述
数据集描述
- 大小类别: 10K<n<100K
- 任务类别: 文本检索
数据集信息
- 特征:
datasetId: 字符串author: 字符串last_modified: 时间戳[微秒, 时区=UTC]downloads: 64位整数likes: 64位整数tags: 字符串序列task_categories: 字符串序列createdAt: 时间戳[微秒, 时区=UTC]card: 字符串
- 分割:
train: 字节数: 659378245, 样本数: 116715
- 下载大小: 153642105
- 数据集大小: 659378245
配置
- 默认配置:
- 数据文件:
- 分割:
train - 路径:
data/train-*
- 分割:
- 数据文件:
标签
ethicsdocumentation
数据集创建
- 数据来源: Hugging Face Hub上托管的数据集的
README.md文件。 - 数据收集和处理: 使用CRON作业每日下载数据。
- 数据生产者: 数据集卡片的创建者,包括社区中的各种人员。
注释
- 注释过程: 无
- 注释者: 无
个人和敏感信息
- 未进行匿名化处理。
偏差、风险和限制
- 数据集卡片由社区创建,内容不受控制。
- 可能包含偏差和敏感信息。
推荐
- 用户应了解数据集的风险、偏差和技术限制。
引用
- 无需正式引用,但使用时请包含数据集页面链接。
数据集卡片作者和联系人




