遇见数据集

UN General Assembly Lexical and Phraseological List (UN-GAPL): Corpus Data and Frequency Lists

收藏
Zenodo2026-06-29 更新2026-08-02 收录
官方服务:

资源简介:

This dataset contains the complete outputs of a corpus-driven pipeline applied to 2,188 English-language verbatim records of United Nations General Assembly plenary sessions (2000–2017, approximately 12.8 million tokens). The resource comprises the UN General Assembly Lexical and Phraseological List (UN-GAPL): a tiered, frequency- and keyness-ranked list of 1,802 lemmas (Tiers 1–3) and 67,347 multi-word units (Tier 4), produced through frequency analysis, log-likelihood keyness analysis (G²), n-gram extraction, and collocation profiling. The data accompany the article: [your article title], submitted to Transletters: International Journal of Translation and Interpreting. An interactive exploration platform is available at: UN GAPL App.

提供机构:
Zenodo
创建时间:
2026-06-29
二维码
社区交流群
二维码
科研交流群
商业服务