Gender Bias in Portuguese Literary Texts: A Masked Language Model Approach
收藏官方服务:
资源简介:
This repository contains all curated data used in the Simpósio em Tecnologia da Informação e da Linguagem Humana (STIL 2025) paper: "Gender Bias in Portuguese Literary Texts: A Masked Language Model Approach". It contains the following data: [corpus.zip] Curated corpus of Portuguese literary texts sourced from four well-established corpora, selected to ensure a broad temporal span (1804-1998) and balanced representation of both Brazilian and European Portuguese. The curated corpus has 592 prose works, with 1.2 million sentences and 17.6 million tokens. [word_lists] The lists of words (i.e., adjectives, verbs and noun phrases) to construct the sentence templates. [constructed_templates] The constructed sentence templates.
提供机构:
Zenodo创建时间:
2025-08-05



