Dataset of randomized sentences in Irish
收藏资源简介:
This file is a TXT containing one sentence per line, compressed using ZIP This dataset contains neither creative works nor copyrighted works that are not available on the Web A subset of data from Phase 1 of the National Corpus of Irish Project (Tionscadal Chorpas Náisiúnta na Gaeilge) that satisfied the criteria were selected for this task The contents of these data were shuffled 1000 times using a Python program, deduplication was not done on this dataset, The final dataset contains 1951584 setences or 76578449 words (counted using WC) This dataset was created with a view to sharing a subset of the data collected for the National Corpus of Irish project without infringing on copyright, and without going against both the terms of use and the trust of the people who shared data with the project - noting this also only applies to a subset of the people who shared data with the aforementioned project. The downstream benefits of publishing this dataset are numerous, helping to support Irish-language NLP research in universities and not-for-profits was chief among our motivations. It has come to our attention that researchers in these types of groups or organisations do not always have high-quality data to hand, nor are these data widely available on the Web, nor do they necessarily have the time or resources to collect a comparable dataset. This dataset can be requested by contacting the Project Leader by email or gaois@dcu.ie.



