ulysses-camara/ulysses-ner-br
收藏资源简介:
--- annotations_creators: [] language_creators: [] language: - pt license: [] multilinguality: - monolingual pretty_name: UlyssesNER-br size_categories: - 10K<n<100K source_datasets: [] task_categories: - token-classification task_ids: - named-entity-recognition --- # Dataset Card for UlyssesNER-Br ## Table of Contents - [Table of Contents](#table-of-contents) - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards) - [Languages](#languages) - [Dataset Structure](#dataset-structure) - [Data Instances](#data-instances) - [Data Fields](#data-fields) - [Data Splits](#data-splits) - [Dataset Creation](#dataset-creation) - [Curation Rationale](#curation-rationale) - [Source Data](#source-data) - [Annotations](#annotations) - [Personal and Sensitive Information](#personal-and-sensitive-information) - [Considerations for Using the Data](#considerations-for-using-the-data) - [Social Impact of Dataset](#social-impact-of-dataset) - [Discussion of Biases](#discussion-of-biases) - [Other Known Limitations](#other-known-limitations) - [Additional Information](#additional-information) - [Dataset Curators](#dataset-curators) - [Licensing Information](#licensing-information) - [Citation Information](#citation-information) - [Contributions](#contributions) ## Dataset Description - **Homepage:** [Convenio-Camara-dos-Deputados/ulyssesner-br-propor](https://github.com/Convenio-Camara-dos-Deputados/ulyssesner-br-propor) - **Repository:** [Convenio-Camara-dos-Deputados/ulyssesner-br-propor](https://github.com/Convenio-Camara-dos-Deputados/ulyssesner-br-propor) - **Paper:** [UlyssesNER-Br: a Corpus of Brazilian Legislative Documents for Named Entity Recognition](https://link.springer.com/chapter/10.1007/978-3-030-98305-5_1) - **Leaderboard:** - **Point of Contact:** ### Dataset Summary [More Information Needed] ### Supported Tasks and Leaderboards [More Information Needed] ### Languages Portuguese (Brazil). ## Dataset Structure ### Data Instances [More Information Needed] ### Data Fields [More Information Needed] ### Data Splits [More Information Needed] ## Dataset Creation ### Curation Rationale [More Information Needed] ### Source Data #### Initial Data Collection and Normalization [More Information Needed] #### Who are the source language producers? [More Information Needed] ### Annotations #### Annotation process [More Information Needed] #### Who are the annotators? [More Information Needed] ### Personal and Sensitive Information [More Information Needed] ## Considerations for Using the Data ### Social Impact of Dataset [More Information Needed] ### Discussion of Biases [More Information Needed] ### Other Known Limitations [More Information Needed] ## Additional Information ### Dataset Curators [More Information Needed] ### Licensing Information [More Information Needed] ### Citation Information ``` @inproceedings{UlyssesNER-Br, title={UlyssesNER-Br: A Corpus of Brazilian Legislative Documents for Named Entity Recognition}, author={Albuquerque, Hidelberg O. and Costa, Rosimeire and Silvestre, Gabriel and Souza, Ellen and da Silva, Nádia F. F. and Vitório, Douglas and Moriyama, Gyovana and Martins, Lucas and Soezima, Luiza and Nunes, Augusto and Siqueira, Felipe and Tarrega, João P. and Beinotti, Joao V. and Dias, Marcio and Silva, Matheus and Gardini, Miguel and Silva, Vinicius and de Carvalho, André C. P. L. F. and Oliveira, Adriano L. I.}, booktitle={Computational Processing of the Portuguese Language}, year={2022}, publisher={Springer International Publishing}, isbn={978-3-030-98305-5}, doi={https://doi.org/10.1007/978-3-030-98305-5_1} } ``` ### Contributions Thanks to [@augusnunes](https://github.com/augusnunes) for adding this dataset.
注释创建者: [] 语言创建者: [] 语言: - 葡萄牙语(pt) 许可证: [] 多语言类型: - 单语 数据集展示名称: UlyssesNER-br 样本规模类别: - 10000 < 样本量 < 100000 源数据集: [] 任务类别: - 词元分类(token-classification) 任务子任务: - 命名实体识别(named-entity-recognition) --- # UlyssesNER-Br 数据集卡片 ## 目录 - [目录](#table-of-contents) - [数据集描述](#dataset-description) - [数据集概述](#dataset-summary) - [支持的任务与排行榜](#supported-tasks-and-leaderboards) - [语言](#languages) - [数据集结构](#dataset-structure) - [数据实例](#data-instances) - [数据字段](#data-fields) - [数据划分](#data-splits) - [数据集构建](#dataset-creation) - [构建初衷](#curation-rationale) - [源数据](#source-data) - [标注](#annotations) - [个人与敏感信息](#personal-and-sensitive-information) - [数据集使用注意事项](#considerations-for-using-the-data) - [数据集的社会影响](#social-impact-of-dataset) - [偏差讨论](#discussion-of-biases) - [其他已知局限性](#other-known-limitations) - [附加信息](#additional-information) - [数据集策展人](#dataset-curators) - [许可证信息](#licensing-information) - [引用信息](#citation-information) - [贡献](#contributions) ## 数据集描述 - **主页:** [Convenio-Camara-dos-Deputados/ulyssesner-br-propor](https://github.com/Convenio-Camara-dos-Deputados/ulyssesner-br-propor) - **代码仓库:** [Convenio-Camara-dos-Deputados/ulyssesner-br-propor](https://github.com/Convenio-Camara-dos-Deputados/ulyssesner-br-propor) - **论文:** [UlyssesNER-Br:面向命名实体识别的巴西立法文档语料库](https://link.springer.com/chapter/10.1007/978-3-030-98305-5_1) - **排行榜:** - **联系人:** ### 数据集概述 【需补充更多信息】 ### 支持的任务与排行榜 【需补充更多信息】 ### 语言 巴西葡萄牙语 ## 数据集结构 ### 数据实例 【需补充更多信息】 ### 数据字段 【需补充更多信息】 ### 数据划分 【需补充更多信息】 ## 数据集构建 ### 构建初衷 【需补充更多信息】 ### 源数据 #### 初始数据收集与归一化 【需补充更多信息】 #### 源语言内容的创作者是谁? 【需补充更多信息】 ### 标注 #### 标注流程 【需补充更多信息】 #### 标注人员是谁? 【需补充更多信息】 ### 个人与敏感信息 【需补充更多信息】 ## 数据集使用注意事项 ### 数据集的社会影响 【需补充更多信息】 ### 偏差讨论 【需补充更多信息】 ### 其他已知局限性 【需补充更多信息】 ## 附加信息 ### 数据集策展人 【需补充更多信息】 ### 许可证信息 【需补充更多信息】 ### 引用信息 @inproceedings{UlyssesNER-Br, title={UlyssesNER-Br: A Corpus of Brazilian Legislative Documents for Named Entity Recognition}, author={Albuquerque, Hidelberg O. and Costa, Rosimeire and Silvestre, Gabriel and Souza, Ellen and da Silva, Nádia F. F. and Vitório, Douglas and Moriyama, Gyovana and Martins, Lucas and Soezima, Luiza and Nunes, Augusto and Siqueira, Felipe and Tarrega, João P. and Beinotti, Joao V. and Dias, Marcio and Silva, Matheus and Gardini, Miguel and Silva, Vinicius and de Carvalho, André C. P. L. F. and Oliveira, Adriano L. I.}, booktitle={Computational Processing of the Portuguese Language}, year={2022}, publisher={Springer International Publishing}, isbn={978-3-030-98305-5}, doi={https://doi.org/10.1007/978-3-030-98305-5_1} } ### 贡献 感谢 [@augusnunes](https://github.com/augusnunes) 添加此数据集。
数据集概述
数据集名称
- 名称: UlyssesNER-br
语言
- 语言: 葡萄牙语 (巴西)
许可证
- 许可证: 未指定
多语言性
- 多语言性: 单语种
大小类别
- 大小类别: 10K<n<100K
任务类别
- 任务类别: 令牌分类
任务ID
- 任务ID: 命名实体识别
引用信息
@inproceedings{UlyssesNER-Br, title={UlyssesNER-Br: A Corpus of Brazilian Legislative Documents for Named Entity Recognition}, author={Albuquerque, Hidelberg O. and Costa, Rosimeire and Silvestre, Gabriel and Souza, Ellen and da Silva, Nádia F. F. and Vitório, Douglas and Moriyama, Gyovana and Martins, Lucas and Soezima, Luiza and Nunes, Augusto and Siqueira, Felipe and Tarrega, João P. and Beinotti, Joao V. and Dias, Marcio and Silva, Matheus and Gardini, Miguel and Silva, Vinicius and de Carvalho, André C. P. L. F. and Oliveira, Adriano L. I.}, booktitle={Computational Processing of the Portuguese Language}, year={2022}, publisher={Springer International Publishing}, isbn={978-3-030-98305-5}, doi={https://doi.org/10.1007/978-3-030-98305-5_1} }




