reveng-labs/public-opj-corpus
收藏资源简介:
Public OPJ/OPJU Corpus 是一个内容寻址的语料库,包含从公共科学数据存储库收集的OriginLab项目文件(扩展名为.opj和.opju)。该数据集以分片形式组织,存储在opj/和opju/目录下的tar文件中,每个分片最多包含100个文件,并通过metadata.jsonl文件提供元数据,包括文件路径、SHA256哈希值、大小、来源信息(如存储库类型、ID、查看URL和DOI)以及下载方式。数据集经过去重处理,确保文件唯一性。用户可以通过HuggingFace数据集库加载元数据,并使用huggingface_hub下载分片文件以提取原始项目文件。每个文件的上游许可证可能不同,因此没有统一的数据集许可证,用户需通过来源信息遵守相应许可条款。
The Public OPJ/OPJU Corpus is a content-addressable corpus containing OriginLab project files with extensions .opj and .opju, collected from public scientific data repositories. This dataset is organized into shards, stored as tar files under the opj/ and opju/ directories. Each shard contains up to 100 files, and metadata is provided via the metadata.jsonl file, including file paths, SHA256 hashes, file sizes, source information such as repository type, ID, view URLs, DOIs, and download methods. The dataset has undergone deduplication to ensure file uniqueness. Users can load the metadata using the Hugging Face Datasets library, and download the shard files via huggingface_hub to extract the original project files. The upstream licenses for each individual file may vary, so there is no unified dataset license. Users are required to comply with the corresponding license terms based on the provided source information.




