kgrabko/JiRackBooksDataset
收藏资源简介:
JiRack Boooks数据集是为1.5B模型设计的,专门针对JiRack分词器格式化。该数据集用于训练一个紧凑的1.5B参数模型,基于广泛的110亿token语料库,通过接近7:1的token与参数比例实现高知识密度和推理能力。它优化用于安全、低延迟的银行应用,支持欺诈预防、垃圾邮件过滤、风险评估和反洗钱检测等任务。建议初始化模型时使用4K上下文窗口以确保稳定性,然后扩展到8K上下文以处理长程依赖。
The JiRack Boooks dataset is formatted for the JiRack tokenizer and designed for the 1.5B model. It is used to train a compact 1.5B parameter model on an extensive 11 billion token corpus, achieving exceptional knowledge density and reasoning capabilities with a token-to-parameter ratio of nearly 7:1. Optimized for secure, low-latency banking applications, it supports tasks such as fraud prevention, spam filtering, risk assessment, and Anti-Money Laundering detection. It is recommended to initialize the model with a 4K context window for stability before scaling to 8K context for long-range dependency handling.



