遇见数据集

Data and models for automatic scansion experiment Dutch Song Database

收藏
Zenodo2020-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This release contains the <strong>data </strong>used in an experiment on automatic scansion for historical Dutch song texts. Aside form the data, two <strong>models </strong>are included in this release as well. One model is essential for running the code that is part of this experiment (model_s); while the other model is an example of an acquired automatic scansion model (best_model). <strong>Item descriptions</strong>: <em>meertens-meter-songs.zip</em> → collection of 23,197 historic Dutch songs (xml-format). These files (and the gathered meta-data) stems from a collaboration project between the <em>Dutch Song Database </em>and the <em>Digital Library for Dutch Literature</em>. All files contain meta-data on the number of beats that is present in individual verse lines. Snippet: <pre><code class="language-xml">&lt;lg&gt; &lt;l id="s1:l1" met="4" type="-+"&gt; Een Meysken op een Rivierken &lt;rhyme label="a" type="m"&gt;sadt&lt;/rhyme&gt;,&lt;/l&gt; &lt;l id="s1:l2" met="2" type="-+"&gt; So schoon zy &lt;rhyme label="b" type="m"&gt;was&lt;/rhyme&gt;,&lt;/l&gt; &lt;l id="s1:l3" met="4" type="-+"&gt; Sy sadt en verbeyde haer soete &lt;rhyme label="c" type="m"&gt;Lief&lt;/rhyme&gt;,&lt;/l&gt; &lt;l id="s1:l4" met="2" type="-+"&gt; Int groene &lt;rhyme label="b" type="m" corresp="#s1:l2"&gt;gras&lt;/rhyme&gt;.&lt;/l&gt; &lt;/lg&gt;</code></pre> <em>model_s</em> → model used for syllabification and assignment of lexical stress of (historic) Dutch words. The development of this model was part of a previous project. <em>stress_xml.zip</em> → collection of 23,197 historic Dutch songs (xml-format). These are the same songs a the <em>meertens-meter-songs</em>, yet now their individual words are syllabified and annotated for lexical stress. The songs in this folder are used as input during the training process. Snippet: <pre><code class="language-xml">&lt;l id="s1:l1" met="4" type="-+"&gt; &lt;w token="een"&gt; &lt;s word-stress="1" line-stress="0"&gt;een&lt;/s&gt; &lt;/w&gt; &lt;w token="meysken"&gt; &lt;s word-stress="1" line-stress="0"&gt;meys&lt;/s&gt; &lt;s word-stress="0" line-stress="0"&gt;ken&lt;/s&gt; &lt;/w&gt; &lt;w token="op"&gt; &lt;s word-stress="1" line-stress="0"&gt;op&lt;/s&gt; &lt;/w&gt; &lt;w token="een"&gt; &lt;s word-stress="1" line-stress="0"&gt;een&lt;/s&gt; &lt;/w&gt; &lt;w token="rivierken"&gt; &lt;s word-stress="0" line-stress="0"&gt;ri&lt;/s&gt; &lt;s word-stress="1" line-stress="0"&gt;vier&lt;/s&gt; &lt;s word-stress="0" line-stress="0"&gt;ken&lt;/s&gt; &lt;/w&gt; &lt;rhyme label="a" type="m"&gt; &lt;w token="sadt"&gt; &lt;s word-stress="1" line-stress="0"&gt;sadt&lt;/s&gt; &lt;/w&gt; &lt;/rhyme&gt; &lt;/l&gt;</code></pre> <em>gold_scan.zip</em> → 198 Dutch song files (xml-format). These files have been annotated by an expert for line stress. <em>eval_splits.zip </em>→ contains the splits made from <em>gold_scan. </em>These are the splits used in the automatic scansion experiment: a development set of 98 songs (used during training), and a test set of 99 songs (used for evaluating the best model after training). <em>best_model.zip</em> → contains the files of an acquired model for automatic Dutch song scansion.

提供机构:
Zenodo
创建时间:
2019-07-01
二维码
社区交流群
二维码
科研交流群
商业服务