遇见数据集

ManyTypes4Py: A benchmark Python Dataset for Machine Learning-Based Type Inference

收藏
Zenodo2021-03-01 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The dataset is gathered on Sep. 17th 2020. It has more than 5.4K Python repositories that are hosted on GitHub. Check out the file <strong>ManyTypes4PyDataset.spec </strong>for repositories URL and their commit SHA. The dataset is also de-duplicated using the CD4Py tool. The list of duplicate files is provided in <strong>duplicate_files.txt </strong>file. All of its Python projects are processed in JSON-formatted files. They contain a seq2seq representation of each file, type-related hints, and information for machine learning models. The structure of JSON-formatted files is described in <strong>JSONOutput.md</strong> file. The dataset is split into train, validation and test sets by source code files. The list of files and their corresponding set is provided in <strong>dataset_split.csv</strong> file. Notable changes to each version of the dataset are documented in <strong>CHANGELOG.md</strong>.

提供机构:
Zenodo
创建时间:
2021-03-01
二维码
社区交流群
二维码
科研交流群
商业服务