遇见数据集

Madjakul/halvest-contrastive

收藏
Hugging Face2025-11-26 更新2025-12-20 收录
官方服务:

资源简介:

--- dataset_info: - config_name: base-10 features: - name: query_halid dtype: string - name: query dtype: string - name: query_year dtype: string - name: query_domain sequence: string - name: query_affiliations sequence: string - name: query_authorids sequence: string - name: pos_halid dtype: string - name: positive dtype: string - name: pos_year dtype: string - name: pos_domain sequence: string - name: pos_affiliations sequence: string - name: pos_authorids sequence: string - name: neg_halids dtype: string - name: negative dtype: string - name: neg_year dtype: string - name: neg_domain sequence: string - name: neg_affiliations sequence: string - name: neg_authorids sequence: string splits: - name: train num_bytes: 3834152941.0486026 num_examples: 730394 - name: test num_bytes: 39129259.033585 num_examples: 7454 - name: valid num_bytes: 39124009.60253676 num_examples: 7453 download_size: 2227392557 dataset_size: 3912406209.6847243 - config_name: base-2 features: - name: query_halid dtype: string - name: query dtype: string - name: query_year dtype: string - name: query_domain sequence: string - name: query_affiliations sequence: string - name: query_authorids sequence: string - name: pos_halid dtype: string - name: positive dtype: string - name: pos_year dtype: string - name: pos_domain sequence: string - name: pos_affiliations sequence: string - name: pos_authorids sequence: string - name: neg_halids dtype: string - name: negative dtype: string - name: neg_year dtype: string - name: neg_domain sequence: string - name: neg_affiliations sequence: string - name: neg_authorids sequence: string splits: - name: train num_bytes: 2439002041.429833 num_examples: 1875668 - name: test num_bytes: 24888465.908128202 num_examples: 19140 - name: valid num_bytes: 24887165.57030646 num_examples: 19139 download_size: 1472798499 dataset_size: 2488777672.9082675 - config_name: base-4 features: - name: query_halid dtype: string - name: query dtype: string - name: query_year dtype: string - name: query_domain sequence: string - name: query_affiliations sequence: string - name: query_authorids sequence: string - name: pos_halid dtype: string - name: positive dtype: string - name: pos_year dtype: string - name: pos_domain sequence: string - name: pos_affiliations sequence: string - name: pos_authorids sequence: string - name: neg_halids dtype: string - name: negative dtype: string - name: neg_year dtype: string - name: neg_domain sequence: string - name: neg_affiliations sequence: string - name: neg_authorids sequence: string splits: - name: train num_bytes: 3207787504.4813294 num_examples: 1401495 - name: test num_bytes: 32732595.62223732 num_examples: 14301 - name: valid num_bytes: 32732595.62223732 num_examples: 14301 download_size: 1922379318 dataset_size: 3273252695.725804 - config_name: base-6 features: - name: query_halid dtype: string - name: query dtype: string - name: query_year dtype: string - name: query_domain sequence: string - name: query_affiliations sequence: string - name: query_authorids sequence: string - name: pos_halid dtype: string - name: positive dtype: string - name: pos_year dtype: string - name: pos_domain sequence: string - name: pos_affiliations sequence: string - name: pos_authorids sequence: string - name: neg_halids dtype: string - name: negative dtype: string - name: neg_year dtype: string - name: neg_domain sequence: string - name: neg_affiliations sequence: string - name: neg_authorids sequence: string splits: - name: train num_bytes: 3642651240.5619364 num_examples: 1111598 - name: test num_bytes: 37170445.63024946 num_examples: 11343 - name: valid num_bytes: 37170445.63024946 num_examples: 11343 download_size: 2152793615 dataset_size: 3716992131.8224354 - config_name: base-8 features: - name: query_halid dtype: string - name: query dtype: string - name: query_year dtype: string - name: query_domain sequence: string - name: query_affiliations sequence: string - name: query_authorids sequence: string - name: pos_halid dtype: string - name: positive dtype: string - name: pos_year dtype: string - name: pos_domain sequence: string - name: pos_affiliations sequence: string - name: pos_authorids sequence: string - name: neg_halids dtype: string - name: negative dtype: string - name: neg_year dtype: string - name: neg_domain sequence: string - name: neg_affiliations sequence: string - name: neg_authorids sequence: string splits: - name: train num_bytes: 3802480108.820647 num_examples: 891803 - name: test num_bytes: 38804950.72384451 num_examples: 9101 - name: valid num_bytes: 38800686.91209593 num_examples: 9100 download_size: 2230176353 dataset_size: 3880085746.4565873 - config_name: ict-1 features: - name: halid dtype: string - name: year dtype: string - name: affiliations sequence: string - name: domains sequence: string - name: authors sequence: string - name: query dtype: string - name: positive dtype: string - name: negative dtype: string splits: - name: train num_bytes: 270852390.17447275 num_examples: 295645 - name: test num_bytes: 2763996.215584178 num_examples: 3017 - name: valid num_bytes: 2763996.215584178 num_examples: 3017 download_size: 181600181 dataset_size: 276380382.60564107 - config_name: ict-2 features: - name: halid dtype: string - name: year dtype: string - name: affiliations sequence: string - name: domains sequence: string - name: authors sequence: string - name: query dtype: string - name: positive dtype: string - name: negative dtype: string splits: - name: train num_bytes: 347528995.34114176 num_examples: 203771 - name: test num_bytes: 3547415.0409507477 num_examples: 2080 - name: valid num_bytes: 3545709.552950291 num_examples: 2079 download_size: 221935538 dataset_size: 354622119.9350428 - config_name: ict-3 features: - name: halid dtype: string - name: year dtype: string - name: affiliations sequence: string - name: domains sequence: string - name: authors sequence: string - name: query dtype: string - name: positive dtype: string - name: negative dtype: string splits: - name: train num_bytes: 433973609.79633397 num_examples: 174150 - name: test num_bytes: 4430692.381383185 num_examples: 1778 - name: valid num_bytes: 4428200.428412779 num_examples: 1777 download_size: 271221960 dataset_size: 442832502.60612994 - config_name: ict-4 features: - name: halid dtype: string - name: year dtype: string - name: affiliations sequence: string - name: domains sequence: string - name: authors sequence: string - name: query dtype: string - name: positive dtype: string - name: negative dtype: string splits: - name: train num_bytes: 491254281.5417554 num_examples: 149850 - name: test num_bytes: 5015809.481207112 num_examples: 1530 - name: valid num_bytes: 5012531.17435665 num_examples: 1529 download_size: 302317093 dataset_size: 501282622.19731915 configs: - config_name: base-10 data_files: - split: train path: base-10/train-* - split: test path: base-10/test-* - split: valid path: base-10/valid-* - config_name: base-2 data_files: - split: train path: base-2/train-* - split: test path: base-2/test-* - split: valid path: base-2/valid-* - config_name: base-4 data_files: - split: train path: base-4/train-* - split: test path: base-4/test-* - split: valid path: base-4/valid-* - config_name: base-6 data_files: - split: train path: base-6/train-* - split: test path: base-6/test-* - split: valid path: base-6/valid-* - config_name: base-8 data_files: - split: train path: base-8/train-* - split: test path: base-8/test-* - split: valid path: base-8/valid-* - config_name: ict-1 data_files: - split: train path: ict-1/train-* - split: test path: ict-1/test-* - split: valid path: ict-1/valid-* - config_name: ict-2 data_files: - split: train path: ict-2/train-* - split: test path: ict-2/test-* - split: valid path: ict-2/valid-* - config_name: ict-3 data_files: - split: train path: ict-3/train-* - split: test path: ict-3/test-* - split: valid path: ict-3/valid-* - config_name: ict-4 data_files: - split: train path: ict-4/train-* - split: test path: ict-4/test-* - split: valid path: ict-4/valid-* task_categories: - text-classification - feature-extraction language: - en pretty_name: HALvest-Contrastive size_categories: - 1M<n<10M --- <div align="center"> <h1> HALvest-Contrastive </h1> <h3> Contrastive triplets Harvested from HAL </h3> </div> --- ## Citation ```bib @misc{kulumba2024harvestingtextualstructureddata, title={Harvesting Textual and Structured Data from the HAL Publication Repository}, author={Francis Kulumba and Wissam Antoun and Guillaume Vimont and Laurent Romary}, year={2024}, eprint={2407.20595}, archivePrefix={arXiv}, primaryClass={cs.DL}, url={https://arxiv.org/abs/2407.20595}, } ``` ## Dataset Copyright The licence terms for HALvest strictly follows the one from HAL. Please refer to the below license when using this dataset. - [HAL license](https://doc.archives-ouvertes.fr/en/legal-aspects/)

# 数据集信息 该数据集包含以下8个配置项: ## 配置项:base-10 ### 特征字段: - 查询HAL标识符(query_halid):字符串类型 - 查询文本(query):字符串类型 - 查询发表年份(query_year):字符串类型 - 查询所属领域(query_domain):字符串序列类型 - 查询作者机构(query_affiliations):字符串序列类型 - 查询作者ID(query_authorids):字符串序列类型 - 正样本HAL标识符(pos_halid):字符串类型 - 正样本文本(positive):字符串类型 - 正样本发表年份(pos_year):字符串类型 - 正样本所属领域(pos_domain):字符串序列类型 - 正样本作者机构(pos_affiliations):字符串序列类型 - 正样本作者ID(pos_authorids):字符串序列类型 - 负样本HAL标识符列表(neg_halids):字符串类型 - 负样本文本(negative):字符串类型 - 负样本发表年份(neg_year):字符串类型 - 负样本所属领域(neg_domain):字符串序列类型 - 负样本作者机构(neg_affiliations):字符串序列类型 - 负样本作者ID(neg_authorids):字符串序列类型 ### 数据集划分: - 训练集(train):数据量3834152941.0486026字节,样本数730394 - 测试集(test):数据量39129259.033585字节,样本数7454 - 验证集(valid):数据量39124009.60253676字节,样本数7453 ### 整体统计:下载大小2227392557字节,数据集总大小3912406209.6847243字节 ## 配置项:base-2 ### 特征字段:与base-10配置一致 ### 数据集划分: - 训练集(train):数据量2439002041.429833字节,样本数1875668 - 测试集(test):数据量24888465.908128202字节,样本数19140 - 验证集(valid):数据量24887165.57030646字节,样本数19139 ### 整体统计:下载大小1472798499字节,数据集总大小2488777672.9082675字节 ## 配置项:base-4 ### 特征字段:与base-10配置一致 ### 数据集划分: - 训练集(train):数据量3207787504.4813294字节,样本数1401495 - 测试集(test):数据量32732595.62223732字节,样本数14301 - 验证集(valid):数据量32732595.62223732字节,样本数14301 ### 整体统计:下载大小1922379318字节,数据集总大小3273252695.725804字节 ## 配置项:base-6 ### 特征字段:与base-10配置一致 ### 数据集划分: - 训练集(train):数据量3642651240.5619364字节,样本数1111598 - 测试集(test):数据量37170445.63024946字节,样本数11343 - 验证集(valid):数据量37170445.63024946字节,样本数11343 ### 整体统计:下载大小2152793615字节,数据集总大小3716992131.8224354字节 ## 配置项:base-8 ### 特征字段:与base-10配置一致 ### 数据集划分: - 训练集(train):数据量3802480108.820647字节,样本数891803 - 测试集(test):数据量38804950.72384451字节,样本数9101 - 验证集(valid):数据量38800686.91209593字节,样本数9100 ### 整体统计:下载大小2230176353字节,数据集总大小3880085746.4565873字节 ## 配置项:ict-1 ### 特征字段: - HAL标识符(halid):字符串类型 - 发表年份(year):字符串类型 - 作者机构(affiliations):字符串序列类型 - 所属领域(domains):字符串序列类型 - 作者ID(authors):字符串序列类型 - 查询文本(query):字符串类型 - 正样本文本(positive):字符串类型 - 负样本文本(negative):字符串类型 ### 数据集划分: - 训练集(train):数据量270852390.17447275字节,样本数295645 - 测试集(test):数据量2763996.215584178字节,样本数3017 - 验证集(valid):数据量2763996.215584178字节,样本数3017 ### 整体统计:下载大小181600181字节,数据集总大小276380382.60564107字节 ## 配置项:ict-2 ### 特征字段:与ict-1配置一致 ### 数据集划分: - 训练集(train):数据量347528995.34114176字节,样本数203771 - 测试集(test):数据量3547415.0409507477字节,样本数2080 - 验证集(valid):数据量3545709.552950291字节,样本数2079 ### 整体统计:下载大小221935538字节,数据集总大小354622119.9350428字节 ## 配置项:ict-3 ### 特征字段:与ict-1配置一致 ### 数据集划分: - 训练集(train):数据量433973609.79633397字节,样本数174150 - 测试集(test):数据量4430692.381383185字节,样本数1778 - 验证集(valid):数据量4428200.428412779字节,样本数1777 ### 整体统计:下载大小271221960字节,数据集总大小442832502.60612994字节 ## 配置项:ict-4 ### 特征字段:与ict-1配置一致 ### 数据集划分: - 训练集(train):数据量491254281.5417554字节,样本数149850 - 测试集(test):数据量5015809.481207112字节,样本数1530 - 验证集(valid):数据量5012531.17435665字节,样本数1529 ### 整体统计:下载大小302317093字节,数据集总大小501282622.19731915字节 # 配置项数据文件映射 各配置项对应的数据文件路径如下: - base-10:训练集对应`base-10/train-*`,测试集对应`base-10/test-*`,验证集对应`base-10/valid-*` - base-2:训练集对应`base-2/train-*`,测试集对应`base-2/test-*`,验证集对应`base-2/valid-*` - base-4:训练集对应`base-4/train-*`,测试集对应`base-4/test-*`,验证集对应`base-4/valid-*` - base-6:训练集对应`base-6/train-*`,测试集对应`base-6/test-*`,验证集对应`base-6/valid-*` - base-8:训练集对应`base-8/train-*`,测试集对应`base-8/test-*`,验证集对应`base-8/valid-*` - ict-1:训练集对应`ict-1/train-*`,测试集对应`ict-1/test-*`,验证集对应`ict-1/valid-*` - ict-2:训练集对应`ict-2/train-*`,测试集对应`ict-2/test-*`,验证集对应`ict-2/valid-*` - ict-3:训练集对应`ict-3/train-*`,测试集对应`ict-3/test-*`,验证集对应`ict-3/valid-*` - ict-4:训练集对应`ict-4/train-*`,测试集对应`ict-4/test-*`,验证集对应`ict-4/valid-*` # 基本属性 - 任务类别:文本分类(text-classification)、特征提取(feature-extraction) - 语言:英语(en) - 数据集名称:HALvest-Contrastive(HALvest-Contrastive) - 样本规模:100万<样本数<1000万 --- <div align="center"> <h1> HALvest-Contrastive(HALvest-Contrastive)</h1> <h3> 从HAL学术出版库中采集的对比三元组数据集</h3> </div> --- ## 引用信息 bib @misc{kulumba2024harvestingtextualstructureddata, title={Harvesting Textual and Structured Data from the HAL Publication Repository}, author={Francis Kulumba and Wissam Antoun and Guillaume Vimont and Laurent Romary}, year={2024}, eprint={2407.20595}, archivePrefix={arXiv}, primaryClass={cs.DL}, url={https://arxiv.org/abs/2407.20595}, } ## 数据集版权 本数据集的许可条款严格遵循HAL平台的许可协议,使用本数据集时请参照下述许可协议: - [HAL许可协议](https://doc.archives-ouvertes.fr/en/legal-aspects/)

提供机构:
Madjakul
二维码
社区交流群
二维码
科研交流群
商业服务