遇见数据集

HausaSRS: A Hausa-English Parallel Corpus of Software Requirements Specifications

收藏
Mendeley Data2026-05-21 收录
官方服务:

资源简介:

This study hypothesizes that software requirements written in English can be systematically transformed into a high-quality Hausa Requirements Engineering dataset through a controlled pipeline that combines document harvesting, domain filtering, glossary-guided translation, weak annotation, and expert validation. A related hypothesis is that, in a low-resource setting, combining automated corpus construction with Human-in-the-Loop review can produce data of sufficient quality for downstream NLP tasks such as translation, FR/NFR classification, and token-level entity extraction. The data consists of a parallel and annotated corpus derived from ~350 Software Requirements Specification (SRS) documents collected across the health, education, and finance domains. These source documents were processed from PDF & DOCX formats. Text was extracted using pdfplumber and python-docx, cleaned with rule-based preprocessing to remove non-semantic artifacts such as page numbers, repeated spacing, URLs, and boilerplate labels, and filtered to retain requirement-like content using requirement-engineering keywords and modal patterns. English segments were retained, technical terms were anchored through a custom SRS glossary and named-entity recognition, and the retained text was translated into Hausa. Hausa outputs were normalized and weakly annotated with BIO tags and FR/NFR labels. Synthetic IEEE-style Hausa requirement templates were also introduced to strengthen corpus structure. After cleaning, deduplication, and removal of malformed rows, the dataset was stored as a cleaned silver corpus and partitioned into train, validation, and test subsets. Results show it is feasible to construct a domain-specific Hausa RE resource from heterogeneous SRS documents using a semi-automated workflow. It also shows that a glossary-aware translation and annotation strategy can preserve important software engineering concepts such as actors, system entities, constraints, and quality attributes in Hausa. Notably, automated annotation alone is not sufficient for reliable low-resource RE data; expert correction by a Hausa-speaking NLP specialist was necessary to refine mistranslations, resolve ambiguous labels, and correct token boundaries. This confirms the importance of expert validation in producing a gold-standard corpus from an initially silver dataset. The dataset is a structured representation of software requirements knowledge in Hausa, aligned with common RE tasks. The parallel English–Hausa component supports machine translation and cross-lingual modeling. The FR/NFR labels support requirement classification, while the BIO tags support sequence labeling and information extraction. Researchers can use the data for translation benchmarking, low-resource RE classification, domain-adaptive pretraining, or Hausa-specific entity extraction. More broadly, the data demonstrates a reproducible pathway for creating Requirements Engineering datasets in under-resourced languages.

本研究提出如下假设:以英文撰写的软件需求可通过一套整合了文档采集、领域过滤、术语表引导式翻译、弱标注与专家验证的标准化流程,系统转化为高质量豪萨语需求工程(Requirements Engineering, RE)数据集。另一相关假设则指出,在低资源语言场景下,将自动化语料库构建与人在回路(Human-in-the-Loop)审核相结合,可生成质量足以支撑下游自然语言处理(Natural Language Processing, NLP)任务的数据集,此类任务包括翻译、功能需求/非功能需求(Functional Requirement/Non-Functional Requirement, FR/NFR)分类以及Token级实体抽取。 本数据集源自覆盖健康、教育与金融三大领域的约350份软件需求规格说明(Software Requirements Specification, SRS)文档,为带标注的平行语料库。原始文档均来自PDF与DOCX格式文件:先通过pdfplumber与python-docx工具提取文本,再基于规则进行预处理以剔除页码、重复空格、URL与样板式标签等非语义冗余信息;随后借助需求工程关键词与情态动词模式进行过滤,保留类需求内容。保留的英文片段中,技术术语通过定制化SRS术语表与命名实体识别(Named Entity Recognition, NER)进行锚定,之后将文本翻译为豪萨语。豪萨语译文经归一化处理后,添加BIO标注(BIO tagging)与FR/NFR标签完成弱标注;同时引入合成的IEEE风格豪萨语需求模板,以强化语料库的结构规范性。经清洗、去重与畸形行剔除后,数据集以净化后的银标准语料库形式存储,并划分为训练集、验证集与测试集子集。 实验结果表明,依托半自动化工作流从异构SRS文档中构建领域专属的豪萨语RE资源是可行的;同时也证实,基于术语表的翻译与标注策略可在豪萨语译文中保留软件工程的核心概念,包括参与者、系统实体、约束条件与质量属性。值得注意的是,仅依靠自动化标注不足以生成可靠的低资源RE数据集,需由通晓豪萨语的NLP专家进行人工校正,以修正误译、消解标注歧义并修正Token边界。这一结论验证了专家验证在从初始银标准语料库打造金标准语料库过程中的关键作用。 本数据集为豪萨语软件需求知识的结构化表征,适配主流RE任务场景:其中英-豪萨语平行语料部分可支撑机器翻译与跨语言建模任务;FR/NFR标签可用于需求分类任务;BIO标注则支持序列标注与信息抽取任务。研究人员可将该数据集用于翻译基准测试、低资源RE分类、领域自适应预训练或豪萨语专属实体抽取任务。从更广泛的视角来看,本数据集为低资源语言的需求工程数据集构建提供了一条可复现的技术路径。

创建时间:
2026-04-30
二维码
社区交流群
二维码
科研交流群
商业服务