Minimal replication dataset for "Textual Overlap Rather Than Domain Alignment: A Comparative Study of Fine-Tuning Strategies for Specialised Machine Translation with Large Language Models"
收藏资源简介:
This record provides the minimal dataset and code package necessary to replicate the reported findings of the study. The deposited materials support reproduction of the test-set summaries, item-level translation-quality metrics, BLEU pass-rate likelihood-ratio G^2 tests, paired t-tests, paired Cohen's d_z calculations, and figure/table source data reported in the manuscript. The full bilingual fine-tuning corpus is not redistributed in this record because parts of the underlying Chinese-English political materials derive from third-party officially published sources that may be subject to copyright restrictions. To support transparency and reproducibility within these constraints, this record provides corpus-source metadata, dataset documentation, test-set evaluation data, item-level metric data, statistical analysis files, training logs, and code sufficient to reproduce the reported findings.



