遇见数据集

Dataset Snickars Scandia

收藏
Zenodo2021-02-16 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Data for the article "Från chiffer till klartext? Temamodellering av statliga offentliga utredningar 1945–1989", <em>Scandia</em> 2021, forthcoming. In 2015 the National Library of Sweden finished digitising all Governmental Official Reports (SOU) from 1922 to 1999. Traditionally, SOU reports – and work performed within different governmental committees – had the task of preparing the Swedish government for apt and rational decision-making. The range of subjects covered by governmental committees and SOU reports basically includes every area of the Swedish welfare state, from issues centered on migration and the environment to cultural policy and media politics. The article departs from an analysis of all SOU-reports from 1945–89 as one massive dataset; in all 3,154 SOU-reports that contain 87 million tokens. Research has been performed within a Jupyter Lab environment, a web application with executable Python code which can be run to perform data analysis. The Jupyter Lab environment has been developed at the digital humanities hub, Humlab at Umeå University, and research is related to the project, Welfare State Analytics. Text Mining and Modeling Swedish Politics, Media &amp; Culture, 1945–89. It is a digital humanities and digital history project that will digitise literature, curate already digitised collections, and perform research via probabilistic methods and text mining models. If all SOU-reports are considered as one single text written by the state, what themes in this vast text can software read and perceive? It is possible to answer such a broad question by way of topic modeling, a computational method to study themes in texts by accentuating words that tend to co-occur and together create different topics. Via co-occurrence, topic modeling creates topics in the form of clusters of similar words (topics); a term or a word may be a part of several topics with different degrees of probability. Topics also occur in relation to each other, and clusters and networks can be visualised by using software as Gephi. The article focuses on topics related to media and media policy. Depending on how many topics a topic model displays – in the article models of 50, 100, 200 and 500 topics are used – different media topics can be detected. In the 50 model, one media topic was found, whereas in the 500 model there were several, with more specific traits as for example film censorship or daily press subsidies. One finding is that film was the single medium that the SOU-genre between 1945–89 devoted most attention, another is that archival issues were closely linked to media topics during the same period. Governmental committees and SOU reports on media were primarily focused on future oriented policies, above all how media should be supported or regulated. Yet, archiving the same media forms was also something that the state was repeatedly interested in. In conclusion, the article in general explains what topic modeling is, how the method can be used in digital historical research – not the least in relation to close reading – and how statistical analysis of the distribution of words in the form of topics can generate interesting results. The SOU data is rich; topics can be traced with many different themes. As a researcher, however, one must learn to work with data; to load different models in the Jupyter Lab environment, to compute various input values, change parameters and often cure outcomes in a way that differs from traditional historical research practices. Keywords: digital humanities, digital history, topic modeling, media history, Swedish Governmental Official Reports (SOU)

本文为发表于《Scandia》2021年(待刊)的论文《从密码到明文?1945–1989年瑞典政府官方调查报告的主题建模》所用数据集。 2015年,瑞典国家图书馆完成了1922年至1999年全部瑞典政府官方调查报告(Government Official Reports,缩写SOU)的数字化工作。长期以来,SOU报告以及各类政府委员会的工作成果,均承担着为瑞典政府提供恰当且合理的决策参考的职能。政府委员会与SOU报告所涵盖的研究主题,基本覆盖瑞典福利国家的所有领域:从移民与环境议题,到文化政策与媒介政治均有涉及。 本论文将1945年至1989年间全部3154份SOU报告作为一个大型数据集展开分析,总文本量达8700万Token。本研究依托Jupyter Lab环境开展:该环境是一款支持可执行Python代码的网页应用,可用于执行数据分析工作。Jupyter Lab环境由于默奥大学Humlab数字人文中心开发,本研究隶属于「福利国家分析:1945–1989年瑞典政治、媒介与文化文本挖掘与建模」项目。该项目属于数字人文与数字历史研究范畴,旨在实现文献数字化、对已有数字化馆藏进行整理,并通过概率方法与文本挖掘模型开展相关研究。 若将全部SOU报告视作瑞典官方撰写的一部整体文本,那么软件能够读取并识别出这部巨幅文本中的哪些主题?要解答这一宽泛的研究问题,可借助主题建模(topic modeling)技术:该计算方法通过聚焦于高频共现词汇以构建不同主题,从而实现文本主题分析。基于词汇共现关系,主题建模以相似词汇簇的形式生成主题:单个术语或词汇可依据不同概率权重隶属于多个主题。主题之间亦存在关联,可借助Gephi等可视化软件呈现主题簇与主题网络。 本论文聚焦于与媒介及媒介政策相关的主题。主题模型所生成的主题数量可灵活调整——本论文分别采用了50、100、200及500个主题的模型,可识别出不同精细程度的媒介相关主题。在50主题模型中仅识别出1个媒介相关主题,而在500主题模型中则可识别出多个具备更具体特征的主题,例如电影审查、日报补贴等。研究发现:1945年至1989年间,电影是SOU报告体裁重点关注的单一媒介;同时,档案议题与同期的媒介主题存在紧密关联。 涉及媒介议题的政府委员会报告与SOU报告,核心关注面向未来的政策制定,尤其是媒介的扶持与监管路径。但与此同时,对同类媒介载体进行档案留存,亦是瑞典政府长期关注的议题。 综上,本论文系统阐释了主题建模技术的原理、其在数字历史研究(尤其是与细读研究结合的场景)中的应用路径,以及基于主题分布的词汇统计分析如何产出富有价值的研究成果。SOU数据集内容丰富,可依托其开展多主题的主题追踪研究。但研究者需掌握数据处理的相关技能:在Jupyter Lab环境中加载不同模型、计算各类输入参数、调整模型超参数,且往往需要以区别于传统历史研究的方式对分析结果进行修正与优化。 关键词:数字人文、数字历史、主题建模、媒介史、瑞典政府官方调查报告(SOU)

提供机构:
Zenodo
创建时间:
2021-02-16
二维码
社区交流群
二维码
科研交流群
商业服务