softcatala/optimot-linguistic-data
收藏资源简介:
该数据集包含从加泰罗尼亚政府语言政策部门的公共Optimot语言咨询服务中提取的4,011个条目。每条记录都涉及一个加泰罗尼亚语问题或语言主题,并包括解释、来源元数据以及可用的直接来源URL。数据以JSON Lines格式提供,每个行包含字段如Fitxa(Optimot卡片标识符)、Darrera versió(最后更新日期)、Títol(条目标题)、Categoria(分类类别)、Resposta(答案或解释文本)、source(来源归属)、source_id(来源页面标识符)、source_url(直接URL)和download_date(数据集提取日期)。数据集适用于检索、搜索、加泰罗尼亚语语言辅助和语言资源实验,但不应作为独立基准使用,因为来源内容是公开的,可能已出现在模型预训练数据中。
This dataset contains 4,011 entries extracted from the public Optimot language advisory service of the Catalan Government's Language Policy Department. Each entry covers a Catalan language question or linguistic topic, and includes explanations, source metadata, and a direct source URL when available. The dataset is provided in JSON Lines format, with each line containing fields such as Fitxa (Optimot card identifier), Darrera versió (last update date), Títol (entry title), Categoria (category), Resposta (answer or explanatory text), source (source attribution), source_id (source page identifier), source_url (direct URL), and download_date (dataset extraction date). This dataset is suitable for retrieval, search, Catalan language assistance, and linguistic resource experiments, but should not be used as a standalone benchmark, as the source content is publicly available and may have been included in model pre-training data.




