官方服务:
资源简介:
Word Embeddings to our paper and conll converted data of the shared task
应用场景:
相关数据集
tyzhu/fw_num_bi_train_10_eval_10
--- configs: - config_name: default data_files: - split: train path: data/train-* - split: train_doc2id path: data/train_doc2id-* - split: train_id2doc path: data/train_id2doc-*
Hugging Face2023-08-22 更新140
Efe2898/turkish-reasoning-merged-tokenized-gemma3
--- dataset_info: features: - name: input_ids list: int32 - name: attention_mask list: int8 - name: labels list: int64 - name: length_before_padding dtype: int64 splits:
Hugging Face2026-03-25 更新80
AraSeg-2026-Shared-Task-NoPnx-PA
该数据集是一个文本分类或序列标注数据集,包含约13,453个样本,划分为训练采样集(train_sampled,3,514个样本)、开发集(dev,5,066个样本)和测试集(test,4,873个样本)。每个样本包含四个字段:文档ID(doc_id,字符串类型)、段落ID(paragraph_id,整型)、文本内容(text,字符串列表形式)以及对应的标签(labels,整型列表形式)。数据以结
Hugging Face2026-05-18 更新50
AraSeg-2026-Shared-Task-Pnx-NP
该数据集包含三个预定义的数据分割:测试集(test,262个样本)、开发集(dev,222个样本)和采样训练集(train_sampled,174个样本)。每个数据样本由三个字段构成:一个字符串类型的文档ID(doc_id)、一个字符串列表类型的文本内容(text),以及一个整型列表类型的标签(labels)。数据集总大小约为8.6 MB。
Hugging Face2026-05-18 更新110
CoNLL-based Extended Czech Named Entity Corpus 2.0
This is a Czech Named Entity Corpus 2.0 transformed into the CoNLL format. The original corpus can be downloaded from: http://hdl.handle.net/11858/00-097C-0000-0023-1B22-8. The...
B2FIND80



