German-Greek Law Machine Translated Corpus with Sentence-Level Alignment and Human Translation Samples
收藏资源简介:
This dataset consists of a corpus of German-Greek legal texts (DE_GR) and specialized legal terminology, designed to support research in legal translation, terminology studies, and machine translation evaluation. The data is provided in an Excel (.xlsx) file and is organized at the sentence level. Each row corresponds to a sentence extracted from German legal texts containing specialized legal terminology. The Excel file includes the following columns: 1. Source References Contains the URL links to the original source of each legal text. 2. German Law Texts Containing Specialized Journalistic Terminology The original German legal text, segmented at the sentence level, with each sentence appearing in a separate cell. 3. Neural Machine Translation Output Generated by Google Translate (2021) Machine-translated output produced using Google Translate (2021 version). 4. Translations Generated by a Deep Learning–Based Neural Machine Translation System (2025) Machine-translated output generated by a neural machine translation system (2025). 5. Human-Validated Corrections of Automatically Rendered Terminology Human-reviewed and corrected translations, focusing on improving the accuracy of specialized legal terminology. Purpose This corpus enables:- Comparative analysis of machine translation quality across time and MT Systems (2021 Google Translate vs 2025 DeepLearning Translate)- Evaluation of terminology translation in legal contexts- Research on human-in-the-loop correction workflows- Development and benchmarking of NLP systems for legal language Format - File format: Excel (.xlsx)- Unit of analysis: Sentence-level alignment- Language pair: German → Greek- Encoding: UTF-8 Notes - The dataset focuses on domain-specific terminology found in legal contexts.



