遇见数据集

Exploring Statistical Change Point Detection Techniques for Performance Anomaly Detection at Mozilla

收藏
Zenodo2026-05-25 更新2026-05-26 收录
官方服务:

资源简介:

Dataset descriptionThis repository contains datasets collected as part of a study investigating performance regression detection and change point evaluation in software performance monitoring systems. The data originates from controlled annotation tasks and follow-up surveys conducted with Mozilla engineers.The dataset is composed of 2 main components: Annotation Experiment Data (tasks_anonymized.csv and annotations_anonymized.csv) Timeseries data & attributes (timeseries.zip and timeseries_attributes.json) Annotation Experiment DataThis dataset captures the results of a manual annotation experiment where participants analyzed performance datasets and identified potential change points. It consists of two CSV files: Tasks Dataset (tasks_anonymized.csv): The tasks dataset records the annotation tasks assigned to participants and metadata related to their completion. Its columns are: Column Description TaskID Unique identifier of the annotation task DatasetName Identifier of the performance dataset analyzed in the task Difficulty Self-reported difficulty level of the task TimeSpent Time spent completing the task (in seconds) Problem Optional field for reporting issues encountered during the task UserID Identifier of the participant who completed the task Annotations Dataset (annotations_anonymized.csv): The annotations dataset contains the actual change point annotations produced by participants while analyzing the datasets. This dataset is used to construct the ground truth annotations and analyze agreement between annotators. Its columns are: Column Description DatasetName Identifier of the dataset being annotated UserID Identifier of the participant who made the annotation AnnotationIndex Index position of the annotated change point within the time series AnnotationType Type of change detected (mean, variance, and mean_variance) It is worth noting that the variance annotations were dropped upon doing the evaluation of the change point detection methods mentioned in the paper as there was low agreement on this type of annotation points. Timeseries data & attributes (timeseries.zip and timeseries_attributes.json): They corespond to the time series data used in the annotations process (a total of 174 time series), and their attreibutes which some of them were displayed to annotators upon performing the annotation tasks. Refer to this publication and its replication package here for more context on the time series data.

数据集说明 本仓库收录了为一项研究收集的数据集,该研究旨在探究软件性能监控系统中的性能退化检测与变点(change point)评估工作。数据来源于针对Mozilla工程师开展的受控标注任务与后续调研。本数据集包含两大核心组成部分: ### 标注实验数据(tasks_anonymized.csv 与 annotations_anonymized.csv) ### 时间序列(time series)数据与属性集(timeseries.zip 与 timeseries_attributes.json) #### 标注实验数据 本数据集收录了一项人工标注实验的结果,实验中参与者需分析性能数据集并识别潜在变点。其包含两个CSV文件: ##### 任务数据集(tasks_anonymized.csv) 该数据集记录了分配给参与者的标注任务及其完成相关的元数据,各字段说明如下: | 字段名 | 说明 | | ---- | ---- | | TaskID | 标注任务的唯一标识符 | | DatasetName | 任务中分析的性能数据集的标识符 | | Difficulty | 参与者自我报告的任务难度等级 | | TimeSpent | 完成任务所花费的时长(单位:秒) | | Problem | 用于报告任务过程中遇到的问题的可选字段 | | UserID | 完成任务的参与者的标识符 | ##### 标注数据集(annotations_anonymized.csv) 本数据集包含参与者在分析数据集时生成的实际变点标注结果,可用于构建基准真值(ground truth)标注集并分析标注者间的一致性。各字段说明如下: | 字段名 | 说明 | | ---- | ---- | | DatasetName | 待标注数据集的标识符 | | UserID | 生成该标注的参与者的标识符 | | AnnotationIndex | 时间序列中被标注变点的索引位置 | | AnnotationType | 检测到的变化类型(均值、方差与均值-方差) | 值得注意的是,在评估论文中提及的变点检测方法时,方差类型的标注已被剔除,原因是该类标注点的标注一致性较低。 #### 时间序列数据与属性集(timeseries.zip 与 timeseries_attributes.json) 二者分别对应标注流程中使用的时间序列数据(共计174条时间序列)及其属性集,其中部分属性会在参与者执行标注任务时展示给标注者。如需了解该时间序列数据的更多背景信息,请参阅本论文及其复现包。

提供机构:
Zenodo
创建时间:
2026-05-25
二维码
社区交流群
二维码
科研交流群
商业服务