遇见数据集

Applying Topic Modelling to the Folk Song Corpus of the Podillia Region

收藏
Zenodo2024-12-06 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is based on Ukrainian folk songs from the Podillia region collections (Dei 1965; Yefremova & Dmytrenko 2014; Myshanych 1976). The text corpus comprises 2,762 songs, containing a total of 52,004 lines and 209,075 tokens. The corpus of Podillia folk songs is in Ukrainian. It was analysed using the R programming language (version 4.4.1) along with RStudio (version: 2024.04.2+764). The text analysis code was developed at the Estonian Literary Museum. This dataset includes the following files: 1. TM_Ukr.folk_songs.R: R script for the topic modeling analysis of the corpus. It covers text preprocessing, stopwords removal, tokenization, lemmatization, part-of-speech (PoS) tagging, document-term matrix (DTM) construction, latent Dirichlet allocation (LDA) analysis, coherence evaluation metrics, word embeddings using the GloVe algorithm, and K-means clustering. 2. corpus_Podillia_folk_songs.csv: A file containing the text data of Podillia region folk songs.

提供机构:
Zenodo
创建时间:
2024-12-06
二维码
社区交流群
二维码
科研交流群
商业服务