Apexgridapps/asotele-corpus-methodology
收藏资源简介:
该数据集是Asotele训练语料库方法论(第三阶段)的文档,并非实际训练数据。它详细描述了为尼日利亚和新兴市场金融推理而微调的开源经济智能语言模型asotele-econ背后的训练语料库设计方法。内容包括完整的阶段计划(SFT和DPO语料库)、数据源(三个现有和三个计划来源)、审核流程(验证器→批评者→标注器)、去重和分层规则,以及推迟某些部分(如GRPO到v2)的明确理由。该文档旨在提高模型训练数据的透明度和可评估性,供银行、资助项目官员、学术合作者和下游微调者在模型发布前验证训练数据过程。文档以Markdown和PDF格式提供,遵循CC BY 4.0许可证。
This dataset is the documentation for the Asotele Training-Corpus Methodology (Phase 3), not the actual training corpus. It details the methodology behind the training corpus for the open economic-intelligence language model asotele-econ, fine-tuned for Nigerian and emerging-market financial reasoning. It includes the full phased plan covering SFT and DPO corpora, three live and three planned sources, the audit pipeline (Validator → Critic → Labeler), deduplication and stratification rules, and explicit reasoning for deferred components (e.g., GRPO to v2). The document aims to enhance transparency and evaluability of model training data, allowing reviewers such as banks, grant officers, academic collaborators, and downstream fine-tuners to verify the training-data process before model release. It is provided in Markdown and PDF formats under the CC BY 4.0 license.




