Datasets for "Streaming Right Multiplication over Grammar-Compressed Matrices: A Memory-Bounded GPU Engine for Genotype and Graph Data" (ALENEX 2027)
收藏资源简介:
Datasets used in the ALENEX 2027 paper 'Streaming Right Multiplication over Grammar-Compressed Matrices'. For each benchmark matrix the record ships the original input — a dense int32 matrix for the genotypes, a textual "row col" edge list for the graphs (.sparse for Wikidata, zstandard-compressed swh_full.zst for Software Heritage) — and the RePair grammar consumed by the GPU engine: 12 genotype matrices (1000 Genomes chr20/21/22 + 5 msprime synthetic panels + a billion-nonzero crossover_synth probe), 7 Wikidata relation edge lists with grammars, and the billion-edge forward DAG of the Software Heritage graph export 2021-03-23-popular-3k-python (the file swh_full: 45,691,499 nodes, 1,218,488,928 edges — not the full Software Heritage archive) with grammar. Every input matrix is bit-reproducible (msprime 1.4.2 + seed 42; deterministic 1000 Genomes slice; sorted Wikidata output; WebGraph dump). The grammar core (.vc.C/.vc.R/.iv/.val) reproduces every table without the dense matrix and is bit-for-bit deterministic; the REANS .ansf.1 encoding is size-deterministic only. Data companion to https://github.com/ftosoni/g-mm-repair (REPRODUCIBILITY.md). Per-file MD5 in MANIFEST.md5.



