相关数据集
GENIAC-team-haijima/yutohub-002-sft-data
--- dataset_info: features: - name: output dtype: string - name: input dtype: string - name: instruction dtype: string splits: - name: train num_bytes: 8959599 num_exam
Hugging Face2024-05-26 更新170
AbderrahmanSkiredj1/ahadith_translation_34k
该数据集包含三个特征:文本(text)、标签(label)和提示(prompt),数据类型均为字符串。数据集包含一个训练集(train),大小为56339232字节,包含34088个样本。下载大小为19500537字节,数据集总大小为56339232字节。数据集的任务类别为翻译(translation),语言为阿拉伯语(ar)。
Hugging Face2024-07-15 更新60
stephantulkens/triviaqa-query-gte-modernbert-pooled
这是一个嵌入阿里巴巴-NLP/gte-modernbert-base模型的数据集,用于大规模蒸馏、检索和相似度搜索等任务。每个数据集示例包括一个文本字段用于嵌入和一个维度为768的嵌入字段。数据集被分为一个包含87622个示例的训练集。嵌入是通过Hugging Face Hub中的阿里巴巴-NLP/gte-modernbert-base模型生成的,数据集的生成得到了Mixedbread AI的支持
Hugging Face2025-10-29 更新50
mmdjiji/bert-chinese-idioms
--- license: gpl-3.0 --- For the detail, see [github:mmdjiji/bert-chinese-idioms](https://github.com/mmdjiji/bert-chinese-idioms). [preprocess.js](preprocess.js) is a Node.JS script to generate the
Hugging Face2022-06-28 更新100



