遇见数据集

专利大模型和生物医药大模型应用

收藏
江苏数据交易所2026-01-30 收录
官方服务:

资源简介:

智慧芽已成功训练专利大模型和生物医药大模型,并积极更多垂直领域,正在训练面向材料、通信等领域的大模型,上述大模型合称“智慧芽垂直领域大模型”。其中,专利大模型通过中国专利代理师资格考试的水平,生物医药大模型达到了通过中国执业药师职业资格考试、美国注册药剂师考试(NAPLEX)的水平。在MMLU、C-Eval,Patent-Bench等综合测评结果显示,智慧芽垂直领域大模型在问答、总结、写作、翻译、分类等方面能力整体优于商业通用大模型。在训练数据方面,得益于智慧芽十余年积累的海量高质量科技创新数据,智慧芽垂直领域大模型的预训练数据达到了千亿级token的规模。另外,在智慧芽垂直领域独特的数据配方构成上,还加入了7000余本专业书籍、丰富的行业常识等内容。在AI算法方面,智慧芽围绕数据、算法训练、测试、强化学习构筑了“四位一体”的训练平台。算法上,采用增强式预训练的策略,基于专利和医药领域超40位专家反馈及其2万多条对比数据的强化学习,配合RAG技术,加强大模型理解能力,减少幻觉,对齐人类意图,将大模型精度提升至80%。在应用场景方面,智慧芽面向知识产权、研发创新、生物医药和科创金融等领域的数据产品和服务拥有百万级的专业用户,与其业务流程深度整合。

Zhihuiya has successfully trained patent-specific and biopharmaceutical large language models (LLMs), and is actively expanding into more vertical domains, with ongoing LLM development for fields such as materials and telecommunications. The aforementioned LLMs are collectively referred to as "Zhihuiya Vertical Domain LLMs". Specifically, the patent LLM matches the proficiency level required to pass the Chinese Patent Attorney Qualification Examination, while the biopharmaceutical LLM meets the standards for passing both the Chinese Licensed Pharmacist Professional Qualification Examination and the North American Pharmacist Licensure Examination (NAPLEX). Comprehensive evaluations including MMLU, C-Eval, and Patent-Bench have shown that Zhihuiya Vertical Domain LLMs outperform commercial general-purpose LLMs overall in capabilities such as question answering, summarization, writing, translation, and classification. Regarding training data, benefiting from the massive high-quality scientific and technological innovation data accumulated by Zhihuiya over the past decade, the pre-training corpus of Zhihuiya Vertical Domain LLMs reaches a scale of hundreds of billions of tokens. Additionally, the unique data formulation for these vertical domain LLMs incorporates over 7,000 professional books, extensive industry common knowledge and other relevant content. In terms of AI algorithms, Zhihuiya has built a "four-in-one" training platform centered on data, algorithm training, testing and reinforcement learning. Specifically, an enhanced pre-training strategy is adopted: based on reinforcement learning with feedback from more than 40 experts in the patent and pharmaceutical fields and over 20,000 sets of comparative data, combined with Retrieval-Augmented Generation (RAG) technology, the platform enhances the LLMs' comprehension ability, reduces hallucinations, aligns with human intentions, and boosts the model accuracy to 80%. In terms of application scenarios, Zhihuiya's data products and services targeting intellectual property, R&D innovation, biopharmaceuticals and tech-enabled financial services have millions of professional users, and are deeply integrated with the business workflows of these users.

搜集汇总
背景与挑战
背景概述
该数据集聚焦于智慧芽垂直领域大模型,特别是专利和生物医药大模型的应用。这些模型在专业资格考试和综合测评中表现卓越,整体能力优于商业通用大模型,得益于千亿级token的预训练数据和独特的增强式算法训练,精度提升至80%。它们深度整合于知识产权、研发创新等领域的业务流程,服务于百万级专业用户。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务