遇见数据集

Research on ML algorithms for segmenting shopping data

收藏
data.europa2024-08-08 更新2025-04-19 收录
官方服务:

资源简介:

<p>The dataset contains research results of available ML algorithms for analyzing shopping data from retail outlets (data by product (SKU)) as part of the project "Research and development work on the creation of a platform with a built-in AI/ML engine, addressed to participants of the distribution chain of the FMCG and Consumer Health markets".</p> <p>The study consisted of checking the effectiveness of the algorithms depending on the segmentation parameters used, e.g. the number of SKUs in the purchasing data, the assumed number of purchasing patterns, data processing time and the number of segments in the resulting data.</p> <p>Synthetic shopping data in many configurations (input parameters) was used for the study.</p> <p>Included resources include:</p> <p>- input data to the segmentation process</p> <p>- segmentation result data along with a comparison of the effectiveness of the algorithms</p> <p>The file structure is described in additional documents.</p> <p>The published research results are for one data set for each combination: number of SKUs/number of concepts.</p> <p>Two factorization methods were tested in TRL3, the remaining one was found to be more effective.</p> <p>Details of research on individual TRLs:</p> <p>- TRL3: results using LDA and NMF factoring + study of algorithms: hc algorithm (ALG1) and kmeans algorithm (ALG2), clique algorithm (ALG3), DBScan algorithm (ALG4) and APC - Affinity Propagation algorithm (ALG5)</p> <p>- TRL4: results using NMF factoring + study of algorithms: hc algorithm (ALG1) and kmeans algorithm (ALG2)</p> <p>- TRL5: results using NMF factoring + study of algorithms: hc algorithm (ALG1) and kmeans algorithm (ALG2)</p> <p style="margin-left:0cm; margin-right:0cm"><span style="color:#000000"><span style="color:black">- TRL6: results using real data and NMF factoring + study of algorithms: hc algorithm (ALG1), kmeans algorithm (ALG2) and clique algorithm (ALG3)</span></span></p> <p style="margin-left:0cm; margin-right:0cm"> </p> <p><span style="color:#000000"><span style="color:black">Summary:</span></span></p> <p><span style="color:#000000"><span style="color:black">1. Synthetic data sets (number of concepts: 2÷10) were subjected to factorization processes using the NMF algorithm and segmentation using the KMeans algorithm and hierarchical Agglomerative Clustering, as well as the proprietary click algorithm.<br /> 2. The AI/ML Platform's efficiency for segmentation was estimated in the context of processing synthetic data with different numbers of SKUs (Stock Keeping Units), i.e. 15, 90, 540 and 1080.<br /> 3. The processing process of each synthetic data set was precisely measured, i.e. information about the processed data set, information about process parameters and values ​​of measures evaluating the segmentation stage were collected. For this purpose, a module monitoring the processing time was configured and launched.<br /> 4. The analysis of the data processing time by the AI/ML Platform showed a clear dependence on the number of SKUs, i.e. on the size of the product portfolio. The broader the product portfolio, the longer the time required for data processing. This conclusion highlights that organizations with a broader product range can expect longer data processing times.</span></span></p> <p><span style="color:#000000"><span style="color:black">5. Factoring and segmentation performed on real data required the development of a segmentation solution operating in an unsupervised learning regime.</span></span></p> <p><span style="color:#000000"><span style="color:black">6. The test results showed that the Agglomerative Clustering algorithm for data segmentation had a longer processing time compared to the KMeans algorithm.</span></span></p> <p><span style="color:#000000"><span style="color:black">7. The KMeans algorithm and the Agglomerative Clustering algorithm in combination with NMF factorization achieved similar results in segmentation quality measures on synthetic data.</span></span></p> <p><span style="color:#000000"><span style="color:black">8. In the tests with synthetic data, the author's clique algorithm was also included, which prioritizes high segment homogeneity. However, this algorithm obtained a lower value of the corrected Rand Index than the KMeans and Agglomerative Clustering algorithms, but this applies to synthetic data. Therefore, the decision was made to conduct additional tests in TRL VI using the clique algorithm on real data. As a result of this activity, it was confirmed that the click algorithm worked well on real data and proved to be more effective for this type of data (PSD purchasing data), where the primary business goal was to achieve high homogeneity of segments.</span></span></p>

创建时间:
2024-05-16
二维码
社区交流群
二维码
科研交流群
商业服务