models-for-tokenizers-metadata
收藏资源简介:
该数据集包含关于机器学习模型的详细元数据,旨在支持模型发现、评估和管理任务。数据集结构包含一个训练集,共有417,492个样本,总大小约为2.09 GB。每个样本包含丰富的字段,如模型ID、作者、下载量、创建和修改时间戳、库名称、点赞数、趋势分数等。技术细节包括模型架构、任务类型、输入输出模态、许可证信息以及相关数据集和语言。此外,还包含模型的安全参数(safetensors_params和gguf_params)和性能指标。此数据集适用于模型推荐系统、性能分析、模型分类和趋势预测等应用场景。
This dataset contains detailed metadata for machine learning models, designed to support tasks including model discovery, evaluation, and management. The dataset structure includes a training set with 417,492 samples and a total size of approximately 2.09 GB. Each sample includes comprehensive fields such as model ID, author, download count, creation and modification timestamps, library name, like count, trend score, and more. Technical details include model architecture, task type, input and output modalities, license information, as well as related datasets and languages. Additionally, it includes model security parameters (safetensors_params and gguf_params) and performance metrics. This dataset is applicable to application scenarios such as model recommendation systems, performance analysis, model classification, and trend prediction.
数据集概述
数据集基本信息
- 数据集名称: models-for-tokenizers-metadata
- 发布者: christopher
- 数据来源: Hugging Face Hub
- 数据集地址: https://huggingface.co/datasets/christopher/models-for-tokenizers-metadata
数据集结构与内容
- 数据配置: 默认配置 (
default) - 数据文件: 训练集 (
train),路径模式为data/train-* - 数据量: 训练集包含 417,492 个样本
- 数据格式: 结构化数据,包含多个特征字段
数据特征字段
数据集包含以下主要特征字段:
模型标识与元数据
_id: 内部标识符 (字符串)id: 模型标识符 (字符串)author: 作者 (字符串)model_index: 模型索引 (字符串)sha: SHA校验值 (字符串)
模型关系与组成
base_models: 基础模型信息 (结构体)models: 模型列表_id: 内部标识符 (字符串)id: 模型标识符 (字符串)
relation: 关系类型 (字符串)
siblings: 兄弟文件列表 (字符串列表)architectures: 架构列表 (字符串列表)
时间信息
created_at: 创建时间 (UTC时间戳)last_modified: 最后修改时间 (UTC时间戳)
统计与交互数据
downloads: 下载次数 (整型)downloads_all_time: 历史总下载次数 (整型)likes: 点赞数 (整型)trending_score: 趋势分数 (浮点型)
模型技术属性
library_name: 库名称 (字符串)pipeline_tag: 流水线标签 (字符串)safetensors: Safetensors格式信息 (字符串)gguf: GGUF格式信息 (字符串)config: 配置信息 (字符串)transformers_info: Transformers库信息 (结构体)auto_model: 自动模型类型 (字符串)custom_class: 自定义类 (字符串)pipeline_tag: 流水线标签 (字符串)processor: 处理器类型 (字符串)
模型参数规模
safetensors_params: Safetensors参数数量 (浮点型)gguf_params: GGUF参数数量 (浮点型)
内容与标签
tags: 标签列表 (字符串列表)licenses: 许可证列表 (字符串列表)datasets: 相关数据集列表 (字符串列表)languages: 语言列表 (字符串列表)metrics: 评估指标列表 (字符串列表)tasks: 任务列表 (字符串列表)modalities: 模态列表 (字符串列表)input_modalities: 输入模态列表 (字符串列表)output_modalities: 输出模态列表 (字符串列表)
访问控制与卡片数据
gated: 访问控制状态 (字符串)card_data: 卡片数据 (字符串)card: 卡片信息 (字符串)spaces: 空间信息 (空值)
数据集规模
- 下载大小: 727,686,332 字节 (约 694 MB)
- 数据集大小: 2,094,602,980 字节 (约 1.95 GB)
- 训练集大小: 2,094,602,980 字节 (约 1.95 GB)
- 样本数量: 417,492 个




