官方服务:
资源简介:
Freeling-based text tokenizer.
应用场景:
相关数据集
GermanSTSBenchmark
该数据集是一个针对德语文本相似度评估的基准测试集,常用于评价句子对之间的相似度。它特别适用于评估双语模型在德语任务上的性能表现,其任务类型为语义文本相似度。
arXiv140
yaswanth-iitkgp/Refined_Prompts
--- dataset_info: features: - name: System_Prompt dtype: string - name: Raw_Prompts dtype: string - name: Total_Chars dtype: int64 - name: Total_Texts dtype: int64 - name:
Hugging Face2024-06-08 更新130
vinilazzari/apis_test
--- dataset_info: features: - name: instruction dtype: string - name: input dtype: string - name: text dtype: string - name: output dtype: string splits: - name: train
Hugging Face2024-04-23 更新60
multi-train/nq-train-multikilt_1107
--- configs: - config_name: default data_files: - split: train path: data/train-* dataset_info: features: - name: query dtype: string - name: pos sequence: string - name: neg
Hugging Face2023-11-10 更新140
tzwilliam0/instruction_following_dpo_filtered_add
该数据集包含三个字段:提示(prompt)、选中(chosen)和拒绝(rejected),均为字符串类型。它被划分为训练集,共有18828个示例,数据集总大小为23622317字节。
Hugging Face2025-10-21 更新80



