数据链接:
官方服务:
资源简介:
Polifonia_Corpus_Wikipedia_Annotations_IT
应用场景:
创建时间:
2022-07-15
相关数据集
Polifonia Corpus - Books Module Metadata - Italian Language (Full)
We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpu
Mendeley Data2024-05-10 更新160
TNSA/PT-1000B
--- task_categories: - text-generation language: - en pretty_name: Red Pajama 1T --- ### Getting Started The dataset consists of 2084 jsonl files. You can download the dataset using HuggingFace: ```p
Hugging Face2026-03-31 更新60
koki0702/zero-llm-data
Zero LLM数据集提供了从多个公共数据集中提取的经过清理和标准化的文本语料库,专门为书籍《从零开始深度学习❻——LLM篇》准备。每个语料库都组织在自己的目录中(如codebot/、storybot/、webbot/),以便文本数据、BPE编码数据和合并规则分组在一起。这使得数据集易于理解、导航和复现。数据集包括Python代码、TinyStories V2和OpenWebText样本,提供原始
Hugging Face2026-01-23 更新150
Språkbanken Text
Språkbanken was established in 1975 as a national center located in the Faculty of Arts, University of Gothenburg. Allén's groundbreaking corpus linguistic research resulted in the creation of one of
re3data.org100



