FinBERTimbau: Curated Brazilian Portuguese Financial Text Corpus and Sentiment Annotation Data
收藏资源简介:
This repository provides the datasets used in the development of FinBERTimbau, a domain-adapted transformer model for sentiment analysis in Brazilian Portuguese financial texts. The corpus is composed of regulatory filings, institutional communications, and financial news articles, including documents from the Brazilian Central Bank (BCB), the Brazilian Securities Commission (CVM), and major financial news outlets. The data were used for both domain-adaptive pretraining and supervised fine-tuning. Sentence-level filtering procedures were applied to improve textual quality, including verb-based filtering and heuristic rules. Sentiment labels were constructed using a hybrid human-in-the-loop annotation process that combines large language model-assisted labeling with human validation under a strict consensus criterion. Due to copyright and licensing restrictions associated with portions of the news content, the files in this repository are released with restricted access. Access may be granted upon reasonable request for academic and non-commercial research purposes.



