FIND
收藏资源简介:
FIND数据集是由麻省理工学院计算机科学与人工智能实验室和东北大学创建的,旨在评估自动化解释性方法的构建块。该数据集包含2000个程序,这些程序模拟了训练过的神经网络的组件,并附有我们希望生成的描述类型。这些函数跨文本和数值领域程序化构建,涉及噪声、组合、近似和偏差等多种现实复杂性。数据集旨在帮助研究人员评估和比较开放式标签工具的效能,以及探索语言模型在解释性任务中的应用。FIND数据集特别关注黑盒函数描述范式,因为这种描述作为现有自动化解释方法的子程序或唯一操作实现。
The FIND Dataset was developed by the MIT Computer Science and Artificial Intelligence Laboratory (MIT CSAIL) and Northeastern University, aiming to evaluate the building blocks of automated interpretability methods. This dataset contains 2000 programs that simulate components of trained neural networks, paired with the target types of descriptions we aim to generate. These programs are programmatically constructed across both textual and numerical domains, incorporating various real-world complexities such as noise, compositionality, approximation, and bias. The dataset is designed to help researchers evaluate and compare the performance of open-ended labeling tools, as well as explore the applications of language models in interpretability tasks. The FIND Dataset specifically focuses on the black-box function description paradigm, as such descriptions serve as either subroutines or the sole operational implementation of existing automated interpretability methods.




