遇见数据集

GL-MedQuAD: A Curated English–Galician Medical Question Answering Dataset

收藏
Zenodo2026-07-31 更新2026-08-01 收录
官方服务:

资源简介:

This data descriptor introduces GL-MedQuAD, the first domain-specific biomedical Question Answering (QA) benchmark for the Galician language. Comprising 2,100 parallel records derived from the MedQuAD corpus—specifically from GHR (N=1,467) and NIHSeniorHealth (N=633)—the dataset addresses the critical scarcity of specialized clinical NLP resources in low-resource linguistic scenarios. The dataset was constructed through a hybrid, multi-stage generation-and-curation pipeline: Answer Translation Module: Initial English-to-Galician translations of the medical answers were executed using SalamandraTA 7B Instruct. Question–Answer Alignment & Synthetic Generation Module: To optimize question-to-answer semantic alignment, Gemma 3 27B Instruct was used to directly generate and align the enhanced questions in both English and Galician. Expert Human Curation: To guarantee domain accuracy and linguistic integrity, human post-editing was conducted in compliance with ISO 18587:2017 standards and validated against official Galician medical references (Diccionario galego de termos médicos and Vocabulario de Medicina). The repository includes a comprehensive documentation file along with three cross-aligned CSV datasets linked by a shared record identifier (original_id): README.md: Contains exhaustive documentation regarding dataset usage, column descriptors and error taxonomy codes. Primary Bilingual Parallel Corpus (GL_MedQuAD_bilingual_translations.csv): Includes raw English QA pairs, machine-translated Galician versions of the answers, ISO-curated Galician translations, synthetically generated questions in both English and Galician with curated Galician variants, and instruction-tuned response formats optimized for LLM fine-tuning and conversational healthcare applications. Translation Quality Assessment Dataset (GL_MedQuAD_translation_evaluation.csv)Combines reference-less automated quality metrics (COMETKiwi, LaBSE) with expert human evaluation scores on a 0–5 scale, complemented by a fine-grained 12-category human post-editing error taxonomy. Source-Text Linguistic Complexity Dataset (GL_MedQuAD_linguistic_complexity.csv)Captures the source-text linguistic complexity of the original English answers across three structured, complementary tiers, spanning human-perceived difficulty ratings provided by domain experts, composite complexity indices derived through Multi-Criteria Decision-Making (MCDM) frameworks—such as Analytic Hierarchy Process (AHP), CRITIC weighting, and Hybrid strategies—and fine-grained individual linguistic metrics.

提供机构:
Zenodo
创建时间:
2026-07-31
二维码
社区交流群
二维码
科研交流群
商业服务