fpadovani/goldfish-dan-latn-100mb-tokenized
收藏资源简介:
该数据集是一个用于自然语言处理任务的结构化数据集,包含两个数据划分:训练集(train)和验证集(validation)。训练集有383,104个样本,占用约113.55 MB空间;验证集有37,763个样本,占用约11.33 MB空间。数据集总大小约为124.89 MB,下载大小约为198.01 MB。特征包括input_ids(int32列表)和attention_mask(int8列表),这些通常用于表示文本序列的编码和注意力掩码,适用于预训练或微调模型,如基于Transformer的NLP模型。数据集没有提供具体的描述性内容(如主题、语言或应用领域),因此描述基于技术规格推断。
This dataset is a structured dataset for natural language processing tasks, comprising two splits: train and validation. The train split contains 383,104 examples and occupies approximately 113.55 MB, while the validation split contains 37,763 examples and occupies approximately 11.33 MB. The total dataset size is about 124.89 MB, with a download size of approximately 198.01 MB. Features include input_ids (list of int32) and attention_mask (list of int8), which are commonly used to represent encoded text sequences and attention masks, suitable for pre-training or fine-tuning models such as Transformer-based NLP models. No descriptive content (e.g., topic, language, or application domain) is provided in the README, so the description is inferred from technical specifications.




