遇见数据集

Dataset of Middle Dutch lexical stress patterns and syllabifications

收藏
Zenodo2020-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This dataset consists of <strong>48.219 Middle Dutch words</strong> taken from in total 205 rhymed texts of the <em>Cd-rom Middelnederlands </em>(1998). All of these words have been <strong>assigned a syllabification and lexical stress pattern</strong>. E.g.: <em>proevede</em> is syllabified as <em>proe-ve-de</em> and has a stress index set at -3, which means that – counting from the rightmost syllable – the third syllable receives stress. This upload contains the following files: The <strong>JSON-file</strong> (compressed), which was used as input data for a machine learning algorithm trained for the automatic syllabification and stress assignment of Middle Dutch polysyllabic words (for the code of this experiment, see GitHub) An <strong>Excel-file</strong>, containing the same data as the JSON (for more convenient reference) A <strong>split file </strong>(compressed), used in the training proces of the above-mentioned experiment A pdf-file with some <strong>insightful illustrations</strong> about the contents of the dataset This dataset is part of the research of Wouter Haverals (FWO, University of Antwerp), carried out under the supervision of prof. Mike Kestemont and em. prof. Frank Willaert.

提供机构:
Zenodo
创建时间:
2019-03-04
二维码
社区交流群
二维码
科研交流群
商业服务