akcit-ijf/jigsaw-toxic-comment-train-processed-seqlen128_translated_padronizado
收藏资源简介:
该数据集主要用于文本分类任务,特别是识别和分类文本中的有毒内容。数据集包含多个特征,如唯一标识符(id)、有毒内容标记(toxic)、严重有毒内容标记(severe_toxic)、淫秽内容标记(obscene)、威胁内容标记(threat)、侮辱内容标记(insult)、身份仇恨内容标记(identity_hate)、输入单词ID(input_word_ids)、输入掩码(input_mask)、所有段ID(all_segment_id)、文本内容(text)和是否为有毒内容标记(is_toxic)。数据集包含一个训练集,大小为424,890,269字节,包含223,549个示例。下载大小为118,843,841字节。
This dataset is primarily used for text classification tasks, particularly for identifying and classifying toxic content in text. The dataset includes multiple features such as unique identifier (id), toxic content label (toxic), severe toxic content label (severe_toxic), obscene content label (obscene), threat content label (threat), insult content label (insult), identity hate content label (identity_hate), input word IDs (input_word_ids), input mask (input_mask), all segment IDs (all_segment_id), text content (text), and is toxic label (is_toxic). The dataset contains a training set with a size of 424,890,269 bytes, comprising 223,549 examples. The download size is 118,843,841 bytes.



