遇见数据集

[Research Data] Page Relevance Classification for Software Development with LLMs

收藏
Zenodo2025-12-04 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the dataset associated with the study “Classification of Software-Development-Related Web Pages Using Large Language Models (LLMs)”. The research evaluates how effectively modern LLMs can classify and rank web pages related to software development tasks, specifically those involving software reuse. Contents Included in Data.zip file: The data.zip archive contains the raw and processed datasets used to validate the study. The data is organized into directories reflecting the two main stages of the research: the initial classification (by AI and Manual annotators) and the consensus analysis (Reviewer Agreements/Disagreements). The dataset covers four distinct software development tasks (Queries), represented as Q1 through Q4: Q1: Implementation of a Menu in JavaFX. Q2: Image Upload implementation using Java Spring. Q3: Implementation of CRUD operations using Java JPA. Q4: Searching for a product item on an e-commerce web application using Java Selenium. The files are categorized as follows: 1. Directory: Data/ClassificationAI and Manual: This folder contains the primary datasets where web pages were evaluated based on two specific criteria. The files are provided in CSV format (converted from Excel spreadsheets): Focus/Unfocused (FocadoDesfocado): Classifications determining whether the retrieved web page content is strictly relevant ("Focused") or contains unnecessary noise ("Unfocused") regarding the user's intent. Existence (Existência): Classifications determining the completeness of the solution provided in the web page (e.g., "Non-existent," "Partially Existent," or "Totally Existent"). 2. Directory: Data/Reviewer's Agreements e Disagreements:This folder contains the Comparative files. These datasets document the validation process, highlighting the consensus and discrepancies between human reviewers and/or the LLM outputs. These files were essential for establishing the Ground Truth for the study. It includes comparative analysis files for both Existence and Focus/Unfocus metrics across all four queries (Q1–Q4). These files detail the specific reviews, justifications for classifications, and the final consensus reached for each web page URL analyzed.

本仓库包含与《基于大语言模型(Large Language Models, LLMs)的软件开发相关网页分类研究》相关的数据集。本研究评估了现代大语言模型对软件开发任务相关网页(尤其是涉及软件复用的任务)进行分类与排序的有效程度。 Data.zip归档文件包含本研究验证所用的原始与处理后数据集,具体内容如下: 数据按照研究的两个核心阶段进行目录划分:初始分类阶段(由人工智能与人工标注者完成)与共识分析阶段(评审者一致与分歧结果)。 本数据集涵盖四项独立的软件开发任务(查询项),编号为Q1至Q4: Q1:JavaFX菜单实现 Q2:基于Java Spring的图片上传功能实现 Q3:基于Java JPA的CRUD操作实现 Q4:使用Java Selenium在电商Web应用中搜索商品条目 文件分类如下: 1. 目录:Data/ClassificationAI and Manual 该文件夹包含核心数据集,其中网页将基于两项特定标准进行评估。文件以CSV格式提供(由Excel电子表格转换而来): - 聚焦/非聚焦(Focus/Unfocused):用于判定检索到的网页内容是否严格贴合用户意图,分为“聚焦”(内容严格相关)与“非聚焦”(包含无关冗余信息)两类。 - 存在性(Existence):用于判定网页中提供的解决方案的完整程度,例如“不存在”“部分存在”或“完全存在”。 2. 目录:Data/Reviewer's Agreements e Disagreements 该文件夹包含对比分析文件。此类数据集记录了验证流程,凸显了人工评审者与/或大语言模型输出结果之间的共识与差异,是确立本研究基准真值(Ground Truth)的核心依据。该目录包含针对四项查询项(Q1–Q4)的存在性与聚焦/非聚焦两项指标的对比分析文件。 这些文件详细记录了每一条被分析的网页URL的具体评审意见、分类依据以及最终达成的共识结果。

提供机构:
Zenodo
创建时间:
2025-12-04
二维码
社区交流群
二维码
科研交流群
商业服务