jorjiiie/cpp_dump
收藏Hugging Face2024-04-14 更新2024-06-12 收录
下载链接:
https://hf-mirror.com/datasets/jorjiiie/cpp_dump
下载链接
链接失效反馈官方服务:
资源简介:
---
dataset_info:
features:
- name: hexsha
dtype: string
- name: size
dtype: int64
- name: ext
dtype: string
- name: lang
dtype: string
- name: max_stars_repo_path
dtype: string
- name: max_stars_repo_name
dtype: string
- name: max_stars_repo_head_hexsha
dtype: string
- name: max_stars_repo_licenses
sequence: string
- name: max_stars_count
dtype: float64
- name: max_stars_repo_stars_event_min_datetime
dtype: string
- name: max_stars_repo_stars_event_max_datetime
dtype: string
- name: max_issues_repo_path
dtype: string
- name: max_issues_repo_name
dtype: string
- name: max_issues_repo_head_hexsha
dtype: string
- name: max_issues_repo_licenses
sequence: string
- name: max_issues_count
dtype: float64
- name: max_issues_repo_issues_event_min_datetime
dtype: string
- name: max_issues_repo_issues_event_max_datetime
dtype: string
- name: max_forks_repo_path
dtype: string
- name: max_forks_repo_name
dtype: string
- name: max_forks_repo_head_hexsha
dtype: string
- name: max_forks_repo_licenses
sequence: string
- name: max_forks_count
dtype: float64
- name: max_forks_repo_forks_event_min_datetime
dtype: string
- name: max_forks_repo_forks_event_max_datetime
dtype: string
- name: content
dtype: string
- name: avg_line_length
dtype: float64
- name: max_line_length
dtype: int64
- name: alphanum_fraction
dtype: float64
splits:
- name: train
num_bytes: 24968990
num_examples: 10000
download_size: 12453480
dataset_size: 24968990
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
---
提供机构:
jorjiiie
原始信息汇总
数据集概述
数据集特征
- hexsha:字符串类型
- size:整数类型(int64)
- ext:字符串类型
- lang:字符串类型
- max_stars_repo_path:字符串类型
- max_stars_repo_name:字符串类型
- max_stars_repo_head_hexsha:字符串类型
- max_stars_repo_licenses:字符串序列类型
- max_stars_count:浮点数类型(float64)
- max_stars_repo_stars_event_min_datetime:字符串类型
- max_stars_repo_stars_event_max_datetime:字符串类型
- max_issues_repo_path:字符串类型
- max_issues_repo_name:字符串类型
- max_issues_repo_head_hexsha:字符串类型
- max_issues_repo_licenses:字符串序列类型
- max_issues_count:浮点数类型(float64)
- max_issues_repo_issues_event_min_datetime:字符串类型
- max_issues_repo_issues_event_max_datetime:字符串类型
- max_forks_repo_path:字符串类型
- max_forks_repo_name:字符串类型
- max_forks_repo_head_hexsha:字符串类型
- max_forks_repo_licenses:字符串序列类型
- max_forks_count:浮点数类型(float64)
- max_forks_repo_forks_event_min_datetime:字符串类型
- max_forks_repo_forks_event_max_datetime:字符串类型
- content:字符串类型
- avg_line_length:浮点数类型(float64)
- max_line_length:整数类型(int64)
- alphanum_fraction:浮点数类型(float64)
数据集分割
- 训练集(train):
- 字节数:24968990
- 示例数:10000
数据集大小
- 下载大小:12453480字节
- 数据集大小:24968990字节
配置
- 默认配置(default):
- 数据文件路径:
data/train-*
- 数据文件路径:



