遇见数据集

JS/TS-Smell - A Curated, Metric-Enriched Dataset for Code Smells and Anti-Patterns in JavaScript and TypeScript

收藏
Zenodo2026-03-20 更新2026-05-26 收录
官方服务:

资源简介:

JS/TS-Smell: A Curated, Metric-Enriched Dataset for Code Smells and Anti-Patterns in JavaScript and TypeScript This dataset provides a rigorously curated collection of annotated JavaScript and TypeScript code snippets, focusing on widely recognized code smells and anti-patterns. It was constructed from actively maintained, high-quality open-source projects retrieved from GitHub, selected using strict inclusion criteria to ensure representativeness, diversity, and sustainability. Each snippet is annotated for selected smells and anti-patterns, guided by established taxonomies (Fowler, 2019; Brown et al., 1998), and validated through expert-based consensus with substantial inter-annotator agreement. In addition, the dataset is enriched with a comprehensive set of software metrics at both class and method/function levels, including measures of size, complexity, cohesion, and coupling. These metrics provide objective, reproducible features that support empirical software engineering research and machine learning applications. The dataset is released in two structured CSV files: Class-level metrics and annotations Method/function-level metrics and annotations Potential applications include: Benchmarking and training machine learning and large language model–based smell detection tools. Empirical studies on software maintainability and quality assessment. Replication and extension of research on code smells and anti-patterns in modern JavaScript/TypeScript systems. This resource addresses a critical gap in the availability of high-quality, language-specific datasets beyond Java, and aims to support both academic research and industrial applications in software quality analysis.

JS/TS-Smell:面向JavaScript与TypeScript代码异味(Code Smells)与反模式(Anti-Patterns)的精选、指标增强型数据集 本数据集精心甄选了一批经过标注的JavaScript与TypeScript代码片段,聚焦于业界公认的代码异味与反模式。其构建素材取自GitHub上活跃维护的高质量开源项目,并通过严格的入选标准筛选,以确保数据集具备代表性、多样性与可持续性。 每一段代码片段均依据已确立的分类体系(Fowler,2019;Brown等人,1998)针对选定的异味与反模式进行标注,并经由专家共识验证,标注者间一致性较高。此外,本数据集还补充了覆盖类与方法/函数两个层级的全面软件指标,包括规模、复杂度、内聚性与耦合性相关度量。这些指标可提供客观可复现的特征,支撑经验软件工程研究与机器学习应用。 本数据集以两个结构化CSV文件形式发布: 类级指标与标注 方法/函数级指标与标注 潜在应用场景包括: 基准测试与训练基于机器学习及大语言模型(Large Language Model,LLM)的异味检测工具。 开展软件可维护性与质量评估相关的经验研究。 复现并拓展针对现代JavaScript/TypeScript系统中代码异味与反模式的相关研究。 本数据集填补了Java之外缺乏高质量语言专属数据集的关键空白,旨在为软件质量分析领域的学术研究与工业应用提供支撑。

提供机构:
Zenodo
创建时间:
2026-03-20
二维码
社区交流群
二维码
科研交流群
商业服务