遇见数据集

DDSC/da-wikipedia-queries-gemma-processed

收藏
Hugging Face2024-11-19 更新2025-04-12 收录
官方服务:

资源简介:

--- dataset_info: features: - name: anchor dtype: string - name: positive dtype: string - name: negative dtype: string - name: negative_index_pos dtype: int64 splits: - name: train num_bytes: 25553993 num_examples: 30280 download_size: 17534937 dataset_size: 25553993 configs: - config_name: default data_files: - split: train path: data/train-* language: - da --- This is a processed version of https://huggingface.co/datasets/DDSC/da-wikipedia-queries-gemma The dataset was created using this script: https://github.com/Dansk-Data-Science-Community/embedding_model/blob/main/create_processed_data.py

数据集信息: 特征项: - 锚点(anchor):字符串类型 - 正样本(positive):字符串类型 - 负样本(negative):字符串类型 - 负样本索引位置(negative_index_pos):64位整型 数据集划分: - 训练集(train):占用字节数25553993,样本总数30280 下载大小:17534937字节 数据集总占用大小:25553993字节 配置项: - 默认配置(default):数据文件: - 训练集划分:路径为data/train-* 语言:丹麦语(da) 本数据集为https://huggingface.co/datasets/DDSC/da-wikipedia-queries-gemma 的预处理版本。 本数据集通过以下脚本生成: https://github.com/Dansk-Data-Science-Community/embedding_model/blob/main/create_processed_data.py

提供机构:
DDSC
二维码
社区交流群
二维码
科研交流群
商业服务