官方服务:
资源简介:
:unav
应用场景:
相关数据集
Comentários tóxicos PT-BR
Dataset com comentários tóxicos coletados de outros datasets
kaggle2023-01-11 更新80
escavador/finepdfs_edu_pt_filtered
--- configs: - config_name: default data_files: - split: train path: data/train-* dataset_info: features: - name: text dtype: string - name: id dtype: string - name: dump d
Hugging Face2026-04-03 更新50
Affect in Tweets PT
This is a data set of Portuguese tweets labelled with the emotion conveyed in the tweet. Each tweet is labelled with an emotion (i.e., anger, fear, joy, sadness). The corpus is available from PORTULAN
SSH Open MarketPlace2025-07-04 更新50
TucanoBR/GigaVerbo-Text-Filter
GigaVerbo Text-Filter是一个包含110,000个随机选择的样本的数据集,这些样本来自9个子集的GigaVerbo(即那些不是合成的子集)。这个数据集用于训练在论文《Tucano: Advancing Neural Text Generation for Portuguese》中描述的文本质量过滤器。为了创建文本嵌入,我们使用了sentence-transformers/LaBS
Hugging Face2025-07-24 更新60
SentiLex-PT 02
SentiLex-PT is a sentiment lexicon for Portuguese, made up of 7,014 lemmas, and 82,347 inflected forms. In detail, the lexicon describes: 4,779 (16,863) adjectives, 1,081...
B2FIND50



