遇见数据集

Gender Bias in Portuguese Literary Texts: A Masked Language Model Approach

收藏
Zenodo2025-08-05 更新2026-05-29 收录
官方服务:

资源简介:

This repository contains all curated data used in the Simpósio em Tecnologia da Informação e da Linguagem Humana (STIL 2025) paper: "Gender Bias in Portuguese Literary Texts: A Masked Language Model Approach". It contains the following data: [corpus.zip] Curated corpus of Portuguese literary texts sourced from four well-established corpora, selected to ensure a broad temporal span (1804-1998) and balanced representation of both Brazilian and European Portuguese. The curated corpus has 592 prose works, with 1.2 million sentences and 17.6 million tokens. [word_lists] The lists of words (i.e., adjectives, verbs and noun phrases) to construct the sentence templates. [constructed_templates] The constructed sentence templates.

提供机构:
Zenodo
创建时间:
2025-08-05
二维码
社区交流群
二维码
科研交流群
商业服务