遇见数据集

SUBTLEX-SR: A subtitle-based frequency norm for Serbian

收藏
Zenodo2026-05-01 更新2026-05-26 收录
官方服务:

资源简介:

SUBTLEX-SR is a subtitle-based frequency norm for Serbian, the Serbian member of the SUBTLEX family of psycholinguistic frequency resources (following Brysbaert & New, 2009). It provides word-form and lemma frequencies, contextual diversity, and dispersion measures derived from the Serbian portion of OpenSubtitles v2018. **Contents.** The resource consists of two lexical tables: - **Wordform table** (2,198,809 entries): one row per surface form, with frequency from a 50-million-token lemmatized subsample, frequency from the full 64,842-film cleaned corpus, contextual diversity, three dispersion measures (Gries DP, DPnorm, Juilland D over decade buckets), and POS distribution.- **Lemma table** (330,535 entries): one row per lemma, with subsample-derived frequency, contextual diversity, and POS distribution. Both tables are provided in two scripts (Latin and Cyrillic) and in two formats (CSV and Apache Parquet). Methodology JSON files documenting all construction decisions, and a deterministic per-film manifest of the lemmatized subsample, are included for reproducibility. Four figures from the accompanying paper (baseline correlations, register divergence, Zipf distribution, CD vs. frequency) are also included. **Headline numbers.** 64,842 distinct films (the contextual-diversity base); 287.6 million alphabetic tokens in the cleaned corpus; 50.5 million classla-tokens in the lemmatized subsample; year coverage 1902–2020 with frequency data restricted to films from 1950 onwards. **Construction summary.** Source: OpenSubtitles v2018 raw Serbian (178,596 XML files). Cleaning included repair of a CP-1250-as-CP-1252 mojibake encoding error affecting 90.6% of source documents, two-phase deduplication consolidating 178,494 cleaned uploads into 64,842 distinct films, language-script normalization, and contamination filtering. Lemmatization performed with classla 2.2.1 (standard model variant) on a stratified subsample of 8,933 films. **Validation.** SUBTLEX-SR correlates strongly with OPUS's pre-computed Serbian frequency table (Pearson *r* = 0.97 on 498,489 forms; internal-consistency check) and moderately with the web-derived srLex baseline (*r* = 0.68 on 216,260 forms; *r* = 0.69 on 35,968 lemmas). The reduced srLex correlation reflects a register difference between subtitle dialogue and web prose; the divergent vocabulary sorts coherently into dialogue-characteristic classes (negated future-tense auxiliaries, vocatives, interjections) on one side and news/government-characteristic classes (country and region names, politicians' surnames, formal connectives) on the other. **Limitations.** No behavioral validation against lexical decision RT data has been performed for this release; this is being pursued through collaboration with the Laboratory for Experimental Psychology, University of Novi Sad. classla's Serbian model lemmatizes a small set of Serbian forms to their Croatian variants (e.g., *šta → što*, *koga → tko*); this affects all classla-sr users and is documented in the README. The corpus contains translation residue from non-Serbian source films (predominantly English-language Hollywood and BBC content) and some Bosnian/Croatian-orthography forms typical of the BCMS continuum. See the README and methodology JSON files for full documentation. **Companion paper.** Popović, M. (in preparation). *SUBTLEX-SR: A subtitle-based frequency norm for Serbian, with attention to register coverage in existing Serbian frequency resources.* Submitted to Language Resources and Evaluation. **License.** CC BY-SA 4.0. Source subtitle text is not redistributed with this deposit; see the README for source-corpus access via OPUS. **References.** - Brysbaert, M., & New, B. (2009). Moving beyond Kučera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English. *Behavior Research Methods*, 41(4), 977–990.- Lison, P., & Tiedemann, J. (2016). OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles. *Proceedings of LREC 2016*, 923–929.- Ljubešić, N., & Dobrovoljc, K. (2019). What does neural bring? Analysing improvements in morphosyntactic annotation and lemmatisation of Slovenian, Croatian and Serbian. *Proceedings of BSNLP 2019*, 29–34.

SUBTLEX-SR是面向塞尔维亚语的基于字幕的词频规范,属于心理语言学词频资源SUBTLEX家族的塞尔维亚语分支(遵循Brysbaert & New, 2009的研究范式)。该资源基于OpenSubtitles v2018的塞尔维亚语部分,提供词形、词元(lemma)频率、语境多样性以及分散度指标。 **内容**。该资源包含两张词汇表: - **词形表**(2,198,809条条目):每一行对应一个表层形式,包含5000万Token经过词元化后的子样本频率、完整64,842部影片清洗语料的频率、语境多样性、三种分散度指标(Gries DP、DPnorm、十年区间下的Juilland D)以及词性(Part-of-Speech, POS)分布。 - **词元表**(330,535条条目):每一行对应一个词元,包含子样本导出的频率、语境多样性以及词性分布。 两张词汇表均提供拉丁字母与西里尔字母两种脚本版本,以及CSV和Apache Parquet两种格式。为保证研究可复现性,资源附带了记录所有构建决策的方法论JSON文件,以及词元化子样本的逐影片确定性清单。此外还包含配套论文中的四张图表:基线相关性、语域差异、齐普夫(Zipf)分布、语境多样性与词频关系图。 **核心统计数据**。涵盖64,842部独立影片(作为语境多样性的计算基准);清洗后语料共包含2.876亿字母Token;词元化子样本包含5050万classla Token;时间覆盖范围为1902年至2020年,其中词频数据仅包含1950年及以后的影片。 **构建流程概述**。数据源为OpenSubtitles v2018原始塞尔维亚语语料(178,596个XML文件)。清洗流程包括:修复90.6%源文档存在的CP-1250被误识为CP-1252的乱码问题;两阶段去重,将178,494条清洗后的上传内容整合为64,842部独立影片;语言-脚本规范化以及污染过滤。使用classla 2.2.1(标准模型变体)对分层抽取的8,933部影片进行词元化处理。 **有效性验证**。SUBTLEX-SR与OPUS预计算的塞尔维亚语词频表相关性极强(基于498,489个词形的皮尔逊相关系数r=0.97,作为内部一致性检验),与网络来源的srLex基线相关性中等(基于216,260个词形的r=0.68;基于35,968个词元的r=0.69)。srLex相关性偏低反映了字幕对话与网络散文之间的语域差异:差异词汇可清晰分为两类,一类为对话特征类(否定式将来时助动词、呼语、感叹词),另一类为新闻/政府文本特征类(国家与地区名称、政客姓氏、正式连接词)。 **局限性说明**。本版本尚未针对词汇决策反应时数据开展行为验证,该工作正与诺威萨德大学实验心理学实验室合作推进。classla的塞尔维亚语模型会将少量塞尔维亚语词形映射为克罗地亚语变体(例如*šta → što*,*koga → tko*),该问题影响所有使用classla-sr的用户,相关说明已收录在README文件中。语料包含非塞尔维亚语源影片的翻译残留(主要为英语好莱坞与BBC内容),以及部分属于BCMS语域的波斯尼亚/克罗地亚正写法形式。完整文档请参阅README与方法论JSON文件。 **配套论文**。Popović, M.(待刊)。《SUBTLEX-SR:面向塞尔维亚语的基于字幕的词频规范,兼论现有塞尔维亚语词频资源的语域覆盖情况》。已提交至《Language Resources and Evaluation》。 **授权协议**。CC BY-SA 4.0。本存档不附带源字幕文本,请参阅README了解通过OPUS获取源语料的方式。 **参考文献**: 1. Brysbaert, M., & New, B. (2009). 超越Kučera和Francis:现有词频规范的批判性评估与面向美式英语的新型改进词频测度。《Behavior Research Methods》, 41(4), 977–990. 2. Lison, P., & Tiedemann, J. (2016). OpenSubtitles2016:从电影与电视字幕中提取大规模平行语料。《Proceedings of LREC 2016》, 923–929. 3. Ljubešić, N., & Dobrovoljc, K. (2019). 神经网络带来了什么?斯洛文尼亚语、克罗地亚语与塞尔维亚语的形态句法标注和词元化改进分析。《Proceedings of BSNLP 2019》, 29–34.

提供机构:
Zenodo
创建时间:
2026-05-01
二维码
社区交流群
二维码
科研交流群
商业服务