Dataset for CodeLite: A Low-Cost Framework for Code Classification Tasks via Enhanced Code Representations and Lexical Fusion
收藏资源简介:
This dataset accompanies the paper: "CodeLite: A Low-Cost Framework for Code Classification Tasks via Enhanced Code Representations and Lexical Fusion" The dataset package is intended to support reproducibility and future research on lightweight approaches for source code classification tasks. The archive includes datasets for the following tasks: Programming Language Identification (PLI) train.csv test.csv unique_tags.csv Source Code Plagiarism Detection train.csv test_0.csv to test_9.csv Source Code Authorship Attribution Authorship.csv The corresponding implementation code, experimental notebooks, and replication package are available in the associated GitHub repository. This dataset is released for academic and research purposes under the CC BY 4.0 license.



