遇见数据集

ELMAS dataset

收藏
Figshare2023-09-03 更新2026-04-08 收录
官方服务:

资源简介:

This dataset provides a set of 18 load profiles with an hourly temporal resolution that represent main industrial and tertiary sectors in France for the year 2018.The ELMAS dataset is derived from a total of 55,730 consumption time series initially split into 424 business sectors and three levels of subscribed capacity. The customer’s field of activity follows the Statistical Classification of Economic Activities in the European Community (NACE), which is a four-digit industry standard classification used in the European Union composed of 21 sections, 88 divisions, 272 groups, and 615 classes. For anonymity concerns, the initial times series are averaged according to their NACE coding and level of subscribed capacity.<br><br>Discrepancies between the temporal patterns of customers that belong to the same NACE section highlight the need to resort to another clustering approach. Thus, a K-means algorithm is used to gather the business groups sharing similar temporal patterns into 18 clusters. The resulting clustering shows that numerous NACE sections are scattered over various clusters, which increases the global heterogeneity of the clustering while spoiling the interpretation. The proportion of these dispersed NACE classes in terms of annual energy consumption remains low, which suggests that a manual reorganisation has little impact on the global consistency of the clusters. This manual reclassification is conducted in such a way that scattered NACE classes are gathered in the cluster that possesses the highest share of the considered NACE section. The energy consumption time series dataset represents a limited panel composed of 55,730 customers, which may bias the output load profiles in comparison with the whole French panel of industrial and tertiary customers. To fill this gap, Enedis provides the annual energy consumption of a wider range of customers for the year 2019. This annual energy consumption dataset is used to generate weights implemented in the clustering approach and to derive weighted average time series for the clusters.

本数据集包含18条时间分辨率为小时级的负荷曲线,代表法国2018年主要工业与第三产业部门的用电情况。ELMAS数据集源自共计55730条用电时间序列,最初被划分为424个业务部门与三级报装容量等级。客户的活动领域遵循《欧洲共同体经济活动统计分类》(NACE),该分类体系为欧盟通用的四位行业标准分类,包含21个大类、88个中类、272个小类以及615个细类。出于匿名性考量,原始时间序列将根据其NACE编码与报装容量等级进行平均化处理。 同一NACE大类下的客户,其时间用电模式存在差异,这凸显了采用其他聚类方法的必要性。为此,研究人员使用K均值(K-means)聚类算法,将具有相似时间用电模式的业务群组划分为18个簇。最终聚类结果显示,大量NACE细类分散在不同簇中,这既提升了聚类的整体异质性,也降低了结果的可解释性。这类分散的NACE细类在年度能耗中的占比仍然较低,意味着手动重组对簇的整体一致性影响有限。本次手动重分类遵循以下原则:将分散的NACE细类归入其所归属NACE大类占比最高的簇中。本能耗时间序列数据集仅包含55730名客户的有限样本,与法国全部工业及第三产业客户群体相比,可能会对输出的负荷曲线产生偏倚。为弥补这一缺陷,Enedis提供了2019年更广范围客户的年度能耗数据。该年度能耗数据集被用于生成聚类方法所需的权重,并推导得到各簇的加权平均时间序列。

创建时间:
2023-09-03
二维码
社区交流群
二维码
科研交流群
商业服务