遇见数据集

A Dataset of Annotated Semantic Descriptions of UI Components for Desktop Environments

收藏
Zenodo2025-07-02 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is part of the paper "Enriching Process-Related UI Logs via Screenshot-Based Activity Labeling Using Vision-Language Models". It contains two main files: screenshots_&_soms.zip and semantic_labels.csv. The screenshot_&_soms.zip file includes the Screenshots referenced in the different events, alongside their respective "Screen Object Models" (SOM), which includes the different UI Components in the Screenshots and their hierarchical relationship at different levels of depth, ranging from "Screen" or "Application" components, all the way down to "Icon" or "Text", passing through intermediate components such as "Container" or "Sidebar". The contents of dataset sums up to a total of 559 manually labelled UI Elements across 100 different screenshots. Each label represents a semantic description of an UI Element which shortly describes its purpose and meaning (e.g., refresh button). This dataset is used in the original paper to evaluate the capacity of LLMs to extract accurate semantic descriptions of UI Elements with varying techniques and models. More precisely, the semantic_labels.csv presents the data in the form a UI log with independent events containing the following properties: Screenshot: Each event has an associated screenshot which captures the UI at the moment of the user interacting with the user interface. EventType: Represents the user input. In this dataset, it is always "left_click". Coordinates: Click coordinates in the event. This information is used to locate the UI Element with which an user has interacted. Class: Indicates the class of the UI Element the user has interacted with (e.g., Button). This information is included for results analysis purposes. Depth: Indicates the depth of the UI Element the user has interacted with within the SOM of the image. This information is included for results analysis purposes. Density: The screenshots included within the dataset are divided into "Low Density", "High Density", and "Medium Density", depending on the number of UI Elements within the image. Through this property we represent in which of the three groups of images the element is located. This information is included for results analysis purposes. GroundTruth: Human generated activity label, following the same set of instructions provide to the VLMs during their evaluation.

提供机构:
Zenodo
创建时间:
2025-03-31
二维码
社区交流群
二维码
科研交流群
商业服务