遇见数据集

Brazilian Legal Proceedings

收藏
Figshare2020-01-28 更新2026-04-28 收录
官方服务:

资源简介:

Our data is composed by two datasets: a dataset of 3*10^6 unlabelled motions (mov_treino.txt) and a dataset containing legal proceedings (mov.txt), each with an individual and variable number of motions, but which have been labeled by law experts (tags.txt). They were labeled in 3 classes: arquivado (archived), ativo (active), suspenso (suspended). These datasets are random samples from the first (São Paulo) ans third (Rio de Janeiro) biggest State Courts. State Courts handle the most variable types of cases throughout the Courts in Brazil, and are responsible for 80% of the total amount of lawsuits. Therefore, these datasets are representative of a very significative portion of the variable use of language and expressions in Courts vocabulary.

创建时间:
2020-01-28
二维码
社区交流群
二维码
科研交流群
商业服务