The data set contains about 120,000 Polish words and sentences and their translations into Kashubian. It was created using two types of sources. The first one is the online dictionaries: kaszebe.org
Parallel texts downloaded from the websites of the Swedish Work Environment Authority. The txt files that are available are the result of running the pdf files through the pdftotext command from an ub
--- language: - si license: - mit --- This repo contains data and source code for the paper Nanayakkara, P., & Ranathunga, S. (2018, May). Clustering Sinhala News Articles Using Corpus-Based Similarit