ParallelQA-18: Multilingual Parallel QA Predictions and Human Validation Dataset
收藏资源简介:
Package v3.4.1 (bases-only): Packaging fix that ships a single Dataset archive. Scientific content is unchanged from v3.4.0 (non-identifying human-validation pair decoding, hashed QC exclusions, and error-analysis sampling/overlap provenance). It contains no copyrighted WOL article bodies, direct participant identifiers, paper aggregates, or source-of-truth result file. Companion Builder remains 10.5281/zenodo.21774412. See FLORES_ATTRIBUTION.md and COPYRIGHT_AND_SOURCE_TERMS.md. ParallelQA-18 is a reconstruction-based multilingual question-answering research resource developed to support the study of the Cross-Lingual Comprehension Gap (CLCG), a construct designed to measure systematic losses in semantic comprehension when the same content is presented to a language model in different languages. The release covers 18 languages and provides the frozen research artifacts used in the associated study, including LLM-generated questions and gold answers, indexes for source-authored human questions, model responses by language, automatic metric scores, complete automated-judge outputs, anonymized human ratings, experimental configurations, sampling definitions, manifests, schemas, and derived CLCG estimates. The resource is distributed as a reconstruction package rather than as a redistribution of the original corpus. Copyrighted source articles and source-authored human questions are therefore not included. Instead, the release provides the corresponding source URLs, publication and paragraph identifiers, extraction metadata, content hashes, and companion software required to retrieve and reconstruct the corpus locally from the original public source. This design preserves the reproducibility of the study while respecting the copyright status of the underlying material. It also enables researchers to verify the provenance and integrity of reconstructed documents through stable identifiers and cryptographic hashes. This record contains the frozen dataset snapshot used for the associated paper: internal study freeze v0.1 and public release v3.4.1, covering 150 articles, 5 evaluated language models, and 18 languages. Human-validation ratings from the completed collection (200 annotation batches) are included in anonymized form. Companion Software (required): This record is not intended to be used alone. It must be used together with the companion software ParallelQA-18 Builder: Reconstruction and Evaluation Toolkit (DOI: 10.5281/zenodo.21774412). Record page: https://zenodo.org/records/21774412. The Dataset provides the frozen research artifacts; the Software provides the reconstruction, validation, and evaluation toolkit. Related-identifier metadata also links the two records.



