Dataset and precomputed outputs from a Bifurcated Hybrid Aspect Term Extraction Approach
收藏资源简介:
This deposit provides the data artifacts and precomputed outputs used in the experiments for the manuscript: “Enhancing Aspect Discovery through a Bifurcated Hybrid Approach for Extracting Relevant Terms from Unstructured-Unlabeled Energy Data.” The Zenodo record is intended to support reproducibility and reviewer inspection of (i) the input datasets used by the proposed pipeline and (ii) the exported outputs used by the multi-metric evaluation. The corresponding notebooks and scripts are available in the companion GitHub repository (see “Related identifiers”). Contents The deposit includes three ZIP archives: data.zipCore datasets and auxiliary resources used throughout the pipeline and evaluation, including: CSV files for the main segmented dataset and branch-specific datasets (e.g., UNSUP/RB variants). Text resources for normalization and filtering, such as custom_stopwords, interjeksi_stopwords, and slangword. dataset_baselines.zipBaseline and intermediate datasets exported in CSV format, including baseline aspect-term outputs and derived datasets used for benchmarking and comparison. Examples include outputs from multiple baseline models and branch-specific exports aligned at the segment level. outputs_bifurcated_hybrid_ate.zipFinal precomputed outputs (CSV) of the bifurcated pipeline for evaluation and auditing: unsup_output.csv (UNSUP-ATE output) rb_output.csv (RB-ATE output) hybrid_output.csv (HYBRID-ATE output) Data organization and expected usage After downloading, reviewers can extract the ZIP archives and place the files into the expected directories referenced by the evaluation notebooks (e.g., data/ and precomputed_outputs/, depending on the share-pack layout). The evaluation notebooks operate on segment-level alignment, using a stable segment identifier (e.g., comment_id + seg_id) to ensure fair and consistent comparisons across methods and sources. Notes on privacy and ethics The shared data are prepared for research reproducibility and do not include personally identifying information. The study uses publicly available online text that has been anonymized prior to analysis. Related code All notebooks, scripts, and instructions to reproduce the evaluation workflow are provided in the companion GitHub repository (see “Related identifiers” in this record).



