遇见数据集

DeepDataFlow

收藏
Zenodo2020-11-05 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This dataset contains 493k LLVM-IRs taken from a wide range of projects and source programming languages, and includes labels for several compiler data analyses. We also include the logs for the machine learning jobs which produced our published experimental results. The uncompressed dataset uses the following layout: <code>labels/</code> Directory containing machine learning features and labels for programs for compiler data flow analyses. <code>labels/&lt;analysis&gt;/&lt;source&gt;.&lt;id&gt;.&lt;lang&gt;.ProgramFeaturesList.pb</code> A ProgramFeaturesList protocol buffer containing a list of features resulting from running a data flow analysis on a program. <code>graphs/</code> Directory containing ProGraML representations of LLVM IRs. <code>graphs/&lt;source&gt;.&lt;id&gt;.&lt;lang&gt;.ProgramGraph.pb</code> A ProgramGraph protocol buffer of an LLVM IR in the ProGraML representation. <code>ll/</code> Directory containing LLVM-IR files. <code>ir/&lt;source&gt;.&lt;id&gt;.&lt;lang&gt;.ll</code> An LLVM IR in text format, as produced by <code>clang -emit-llvm -S</code> or equivalent. <code>test/</code> A directory containing symlinks to graphs in the <code>graphs/</code> directory, indicating which graphs should be used as part of the test set. <code>train/</code> A directory containing symlinks to graphs in the <code>graphs/</code> directory, indicating which graphs should be used as part of the training set. <code>val/</code> A directory containing symlinks to graphs in the <code>graphs/</code> directory, indicating which graphs should be used as part of the validation set. <code>vocab/</code> Directory containing vocabulary files. <code>vocab/&lt;type&gt;.csv</code> A vocabulary file, which lists unique node texts, their frequency in the dataset, and the cumulative proportion of total unique node texts that is covered. For further information please see our ProGraML repository.

提供机构:
Zenodo
创建时间:
2020-10-28
二维码
社区交流群
二维码
科研交流群
商业服务