DATASET Log Parsing: How Far Can Llama-2 Go?
收藏资源简介:
We have employed datasets sourced from LogPAI (Log Parsing and Anomaly Detection), a research-oriented platform dedicated to addressing anomaly detection and log data analysis challenges. To our knowledge, this is the most extensive compilation of log datasets. LogPAI have made every effort to keep the logs in their original, unsanitized, anonymized, and unaltered form. These datasets are openly available for research purposes. These datasets are integral to our log parsing study, encompassing seven distinct datasets originating from various systems, including distributed systems (e.g., HDFS, Spark), supercomputers (Thunderbird), operating systems (e.g., Windows), mobile systems (e.g., Android), server applications (e.g., Apache), and standalone software (e.g., Proxifier). Notably, each of these datasets comprises precisely 2000 manually labeled log messages which server as a groundtruth in our study which will be utilized to calculate accuracy of the model.



