gplsi/alia_gva_grammar
收藏资源简介:
ALIA_GVA_Grammar数据集是一个专为文本生成设计的资源。数据集由Markdown格式的文本文档组成,以结构化的JSONL条目形式提供。每个条目包含文本的语言、格式、文本内容、来源和元数据等信息。具体来说,语言为瓦伦西亚语(Valencian),格式为JSON Lines(.jsonl),规模在1,000到10,000个条目之间。数据集可能包含Markdown格式化结构(如标题、列表、强调等),并遵循CC BY 4.0许可证。该数据集由西班牙数字化与公共职能部资助,欧盟NextGenerationEU共同出资,作为ALIA模型开发项目的一部分。
The ALIA_GVA_Grammar dataset is a resource specifically designed for text generation. It consists of Markdown-formatted text documents, distributed as structured JSON Lines (JSONL) entries. Each entry contains information such as the text's language, format, content, source, and metadata. Specifically, the language of the text is Valencian, the dataset uses JSON Lines (.jsonl) as its format, and it includes between 1,000 and 10,000 entries. The dataset may contain Markdown formatting structures such as headings, lists, emphasis, and other similar elements, and is licensed under CC BY 4.0. This dataset was funded by the Spanish Ministry of Digital Transformation and Public Administration, co-financed by the European Union's NextGenerationEU, and is part of the ALIA model development project.



