mAIEnergy
收藏资源简介:
mAIEnergy数据集是由UBITECH能源数字化小组构建的多模态开放语料库,旨在为能源领域的大语言模型应用提供基础支持。该数据集整合了约5万份文本文档、2万张图像、2500万条数值时间序列记录以及200万条地理空间与关系数据,涵盖政策法规、科学文献、新闻文章、卫星影像、电力系统测量、天气观测及能源基础设施地理表征等多源异构数据。其创建过程通过系统化的工作流实现,包括数据识别、检索与准备三个阶段,并对原始数据进行清洗、模式统一与跨模态关联,最终形成符合FAIR原则的结构化知识库。该数据集主要应用于能源领域的AI驱动研究、建模与决策支持,能够有效支撑大语言模型的持续预训练与检索增强生成系统开发,以解决能源转型中复杂多源信息融合与智能分析的核心挑战。
The mAIEnergy Dataset is a multimodal open corpus developed by the UBITECH Energy Digitalization Team, aiming to provide foundational support for large language model (LLM) applications in the energy sector. This dataset integrates approximately 50,000 text documents, 20,000 images, 25 million numerical time-series records, and 2 million geospatial and relational data entries, covering multi-source heterogeneous data including policies and regulations, scientific literature, news articles, satellite imagery, power system measurements, weather observations, and geospatial representations of energy infrastructure. Its construction follows a systematic workflow comprising three stages: data identification, retrieval, and preparation. Raw data undergoes cleaning, schema unification, and cross-modal correlation, ultimately forming a structured knowledge base compliant with the FAIR Principles. This dataset is primarily applied in AI-driven research, modeling, and decision support within the energy field, effectively supporting continuous pre-training of LLMs and the development of retrieval-augmented generation (RAG) systems to address core challenges of complex multi-source information fusion and intelligent analysis in energy transition.

- 1A Multimodal Dataset for Large Language Model Applications in the Energy DomainUBITECH·能源数字化小组 · 2026年




