遇见数据集

TeCla: Text Classification Catalan dataset

收藏
Zenodo2022-11-18 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

If you use this resource in your work, please cite our latest paper: @inproceedings{armengol-estape-etal-2021-multilingual,<br> title = "Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? {A} Comprehensive Assessment for {C}atalan",<br> author = "Armengol-Estap{\'e}, Jordi and<br> Carrino, Casimiro Pio and<br> Rodriguez-Penagos, Carlos and<br> de Gibert Bonet, Ona and<br> Armentano-Oller, Carme and<br> Gonzalez-Agirre, Aitor and<br> Melero, Maite and<br> Villegas, Marta",<br> booktitle = "Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021",<br> month = aug,<br> year = "2021",<br> address = "Online",<br> publisher = "Association for Computational Linguistics",<br> url = "https://aclanthology.org/2021.findings-acl.437",<br> doi = "10.18653/v1/2021.findings-acl.437",<br> pages = "4933--4946",<br> } <em>Corpus de notícies en català per a classificació textual, extret del web de l'Agència Catalana de Notícies sota llicència CC-BY-NC-ND</em> TeCla is a Catalan News corpus for thematic Text Classification tasks. It contains 153.265 articles classified under 30 different categories. The source data is crawled from the ACN (Catalan News Agency) site: http://www.acn.cat, and used under CC-BY-NC-ND 4.0 licence. The dataset is released under the same licence, and is intended exclusively for training Machine Learning models. This dataset was developed by BSC TeMU as part of the AINA project, and intended as part of CLUB (Catalan Language Understanding Benchmark).

提供机构:
Zenodo
创建时间:
2021-05-14
二维码
社区交流群
二维码
科研交流群
商业服务