Saudi Legal English Corpus (SaLEC)
收藏资源简介:
The Saudi Legal English Corpus (SaLEC) is a corpus of Saudi legal English: statutes and implementing regulations, financial-regulator rulebooks and circulars, state model contracts, law-firm commentary, judicial and quasi-judicial decisions, technical and regulatory instruments, and a matched reference stratum of US and UK statutes. Version 2.0 is the unified corpus under a single annotation ruleset. Size: 6,552 documents, 22,484,616 running words, across 8 strata (text revision b1.6). Annotation: an L1 part-of-speech and lemma layer, ruleset cite-sent-r6 (fingerprint 57560a7073960f7d5befe055322157ea6913309bedbb95792d2ac4b2310a1adc), in CoNLL-U, VRT, and AntConc formats. The accuracy of record is a pre-declared blind measurement over four pools; the population-weighted corpus estimate is 0.9561, 95% CI [0.9427, 0.9670]. Full disclosure, including the error decomposition and the non-comparability note for earlier figures, is in the datasheet (Section 5). Licensing: release rests on two independent bases, a written rights-holder clearance and the Saudi Copyright Law Article 4 official-documents exclusion (see the datasheet, Section 6). Deposit contents: README, datasheet, a document-level manifest (6,552 rows), the extracted texts (revision b1.6), the cite-sent-r6 annotation layer in CoNLL-U / AntConc / VRT, governance sidecars, and a byte inventory (bundle-inventory.csv, sha256 per file). Cite as the Saudi Legal English Corpus (SaLEC), version 2.0, DOI 10.5281/zenodo.21328459.



