官方服务:
资源简介:
Regression test to compare NLTK vs Tatarus Porter stemmer
应用场景:
创建时间:
2017-08-19
相关数据集
doushabao4766/ontonotes_zh_ner
--- dataset_info: features: - name: id dtype: int64 - name: tokens sequence: string - name: ner_tags sequence: int64 splits: - name: train num_bytes: 7629700 num_exampl
Hugging Face2023-05-26 更新310
achintasandia/cnn_2018_ecoli
该数据集包含多个特征字段,如日期、年份、月份、日、作者、标题、文章、URL、部分、出版物以及由Meta-Llama-3-70B-Instruct模型生成和清理的声明等。数据集包含一个训练集分割,共有50个样本,文件大小为907927字节。
Hugging Face2024-06-26 更新100
Dataset for the Article: Automatic Subject Descriptions of Polish Library and Information Science Articles: A Comparison of DESCRIPTOR E-Service and GPT-4o
This dataset contains a set of materials and metrics used to compare the performance of two automatic indexing tools: DESCRIPTOR e-service and GPT-4o. Comprehensive details regarding the contents of t
DataCite Commons2025-07-08 更新80
takara-ai/micropajama_embedded_qwen_embed_8B
这是一个包含文本数据和元数据的机器学习数据集,适用于训练自然语言处理模型。数据集由大约25万个训练样本组成,每个样本包括文本内容和一个包含redpajama_set_name字段的元数据结构,以及一个浮点数列表qwen。
Hugging Face2025-09-10 更新130
techiaith/cofnodycynulliad_en-cy
该数据集由英语-威尔士语句对组成,这些语句对是通过解析威尔士议会网站提供的数据获得的。数据集支持翻译、文本分类和句子相似性等任务,语言包括英语和威尔士语。数据集的结构包括源语言和目标语言字段,数据分割为训练集。数据集的创建过程使用了DVC和Python的内部管道。源数据收集和标准化过程中,如果句子包含过多拼写错误或句子长度差异过大,则会被丢弃。源语言数据来自Senedd全体会议的记录及其翻译。数据
Hugging Face2025-04-07 更新100



