kasuboski/gleam-code-corpus
收藏资源简介:
Gleam代码语料库是一个结构化、带属性的语料库,包含从1,528个GitHub仓库收集的22,581个Gleam源代码文件,用于持续预训练和代码生成研究。每个文件都包含完整的属性信息,如仓库所有者、名称、URL、许可证、星标数量以及可用的Hex.pm包元数据。该数据集适用于代码聚焦语言模型的持续预训练、Gleam代码生成的微调、小型编程语言生态系统研究以及Gleam习语和模式的参考语料库。它不是指令/响应对(参见配套的gleam-sft数据集)、文档(语言导览、hexdocs——单独的数据集)或原始未过滤数据(已移除生成文件、未许可代码和重复项)。
A structured, attributed corpus of 22,581 Gleam source files collected from 1,528 GitHub repositories for continued pre-training and code generation research. Every file includes full attribution — repo owner, name, URL, license, star count, and Hex.pm package metadata where available. What this is for: Continued pre-training (CPT) of code-focused language models, Fine-tuning models for Gleam code generation, Research on small programming language ecosystems, Reference corpus for Gleam idioms and patterns. What this is NOT: Instruction/response pairs (see the companion `gleam-sft` dataset, coming later), Documentation (language tour, hexdocs — separate dataset, coming later), Raw/unfiltered data (generated files, unlicensed code, and duplicates have been removed).




