遇见数据集

Datasets and Models for Historical Newspaper Article Segmentation

收藏
Zenodo2023-01-22 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This record contains the datasets and models used and produced for the work reported in the paper "<em>Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers</em>" (link). Please cite this paper if you are using the models/datasets or find it relevant to your research: <pre><code>@article{barman_combining_2020, title = {{Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers}}, author = {Raphaël Barman and Maud Ehrmann and Simon Clematide and Sofia Ares Oliveira and Frédéric Kaplan}, journal= {Journal of Data Mining \&amp; Digital Humanities}, volume= {HistoInformatics} DOI = {10.5281/zenodo.4065271}, year = {2021}, url = {https://jdmdh.episciences.org/7097}, }</code></pre> <br> <strong>Please note that this record contains data under different licenses.</strong><br> <br> <strong>1. DATA</strong> <strong>Annotations (json files)</strong>: JSON files contains image annotations, with one file per newspaper containing region annotations (label and coordinates) in VIA format. The following licenses apply: luxwort.json: those annotations are under a CC0 1.0 license. Please refer to the right statement specified for each image in the file. GDL.json, IMP.json and JDG.json: those annotations are under a CC BY-SA 4.0 license. <strong>Image files: </strong>The archive images.zip contains the Swiss titles image files (GDL, IMP, JDG) used for the experiments described in the paper. Those images are under copyright (property of the journal <em>Le Temps </em>and of <em>ArcInfo</em>) and can be used <em>for academic research or educational purposes only</em>. Redistribution, publication or commercial use are not permitted. These terms of use are similar to the following right statement: http://rightsstatements.org/vocab/InC-EDU/1.0/ <strong>2. MODELS</strong> Some of the best models are released under a CC BY-SA 4.0 license (they are also available as assets of the current Github release). <strong>JDG_flair-FT</strong>: this model was trained on JDG using french Flair and FastText embeddings. It is able to predict the four classes presented in the paper (<code>Serial</code>, <code>Weather</code>, <code>Death notice</code> and <code>Stocks</code>). <strong>Luxwort_obituary_flair-bpemb</strong>: this model was trained on Luxwort using multilingual Flair and Byte-pair embeddings. It is able to predict the <code>Death notice</code> class. <strong>Luxwort_obituary_flair-FT_indomain</strong>: this model was trained on Luxwort using in-domain Flair and FastText embeddings (trained on Luxwort data). It is also able to predict the <code>Death notice</code> class. Those models can be used to predict probabilities on new images using the same code as in the original dhSegment repository. One needs to adjust three parameters to the <code>predict</code> function: 1) <code>embeddings_path</code> (the path to the embeddings list), 2) <code>embeddings_map_path</code>(the path to the compressed embedding map), and 3) <code>embeddings_dim</code> (the size of the embeddings). Please refer to the paper for further information or contact us. <strong>3. CODE: </strong> https://github.com/dhlab-epfl/dhSegment-text <br> <strong>4. ACKNOWLEDGEMENTS</strong><br> We warmly thank the journal Le Temps (owner of <em>La Gazette de Lausanne</em> and the <em>Journal de Genève</em>) and the group ArcInfo (owner of <em>L'Impartial</em>) for accepting to share the related datasets for academic purposes. We also thank the National Library of Luxembourg for its support with all steps related to the <em>Luxemburger Wort</em> annotation release.<br> This work was realized in the context of the <em>impresso</em> - Media Monitoring of the Past project and supported by the Swiss National Science Foundation under grant CR- SII5_173719.<br> <br> <strong>5. CONTACT</strong><br> Maud Ehrmann (EPFL-DHLAB)<br> Simon Clematide (UZH)

提供机构:
Zenodo
创建时间:
2021-01-30
二维码
社区交流群
二维码
科研交流群
商业服务