Experimental Dataset of the Paper "Evaluating the Reliability of Large Language Models for Software Verification"
收藏资源简介:
This archive contains the dataset used to perform the experiments reported in the paper entitled "Evaluating the Reliability of Large Language Models for Software Verification".It contains the following hierarchy of directories: - no-data-race - altered - helped - neutral - tricked - original - termination - altered - helped - neutral - tricked - original - unreach-call - altered - helped - neutral - tricked - original - valid-deref - altered - helped - neutral - tricked - original Each root directory is associated to the corresponding SV-COMP property handled in the paper.The subdirectories "original" contain the list of files as they were originally proposed in the SV-COMP contest.The subdirectories "altered" contain the list of files "altered" by the techniques proposed in the paper (obfuscation and/or invertion).Each "altered" program exists under three forms: "neutral", meaning that it does not contain any comment potentially altering the LLM's reasoning, "helped", meaning that it contains comments aiming at helping the LLM's reasoning, and "tricked", meaning that it contains comments aiming at "tricking" the LLM's reasoning. The original and neutral versions of the programs are accompanied by their corresponding ".yml" file, indicating the expected validity of the property to verify.



