遇见数据集

Data for: Benchmarking the Sentence-Level Simplification of Dutch Municipal Text

收藏
NIAID Data Ecosystem2026-05-01 收录
数据链接:
官方服务:

资源简介:

The corpus consists of 1311 automatically aligned complex-simple sentence pairs. Original documents For the creation of the dataset, we have used ~50 documents provided by the Communications Department of the City of Amsterdam. The documents have diverse sources and purposes (e.g. reports, citizen letters, newsletters, etc.) and cover a variety of topics (legal, medical, urban planning, etc.). The documents were reviewed by an expert and contain edits related to simplification, but also tone of voice, spelling corrections and other improvements. For the alignment of the sentence, we have used the original and final version of the documents. Alignment The alignment processed consists of 2 steps. For each document, create candidate complex-simple pairs by: aligning paragraphs (based on TF-IDF similarity) aligning sentences within the aligned paragraphs (based on TF-IDF similarity) Post-processeding After the initial alignment of sentences we: filter the candidate pairs where differences are only in capitalization or punctuation drop duplicates for every complex sentence where there are multiple possible simple versions, create a new entry by merging all simple versions (under the assumption that a complex sentence was simplified by splitting it into 2 or more simple sentences) for every complex sentence, preserve the simple version with the lowest Levenshtein edit-distance Anonymization In order to publish the dataset, the following changes have been performed to the dataset: We have removed: names (including those of people with public functions such as the gemeentesecretaris) -> substituted with [NAME] organizations (with the exception of public organization such as Gemeente Amsterdam and GGD, RIVM) -> substituted with [ORGANIZATION] addresses (whenever they included an exact street and number) -> substituted with [ADDRESS] phone numbers -> substituted with [NUMBER] a handful of sentence posing privacy or information security risks, or containing otherwise sensitive information -> replaced by "xxx xxx xxx" for transparency Finally, the anonymized version of the dataset does not contain further information about the source documents (e.g. their names), however, a document ID has been added in order to provide context information about sentences stemming from the same document. Acknowledgements This dataset was created by Amsterdam Intelligence for the City of Amsterdam. We owe a special thank you to the Communications Department of the City of Amsterdam for providing us with the original 48 documents. We also thank Daniel Vlantis for providing feedback during the dataset creation and for extensive experiments with it. License This data is licensed under the terms of the European Union Public License 1.2 (EUPL-1.2).

本语料库包含1311对自动对齐的复合句-简单句配对。 ## 原始文档 本数据集的构建采用了阿姆斯特丹市通信部门提供的约50份文档。这些文档来源与用途多样(例如报告、市民来信、通讯简报等),涵盖法律、医疗、城市规划等多个主题领域。 经专家审阅后的文档包含了与句式简化、语气调整、拼写修正及其他优化相关的编辑内容。为实现句子对齐,我们采用了文档的原始版本与终版进行比对。 ## 对齐流程 本次对齐流程共分为两个步骤。针对每份文档,我们通过以下方式生成候选复合句-简单句配对: 1. 基于词频-逆文档频率(TF-IDF)相似度对齐段落; 2. 在已对齐的段落内,同样基于TF-IDF相似度对齐句子。 ## 后处理流程 完成句子的初始对齐后,我们执行以下操作: - 过滤掉仅存在大小写或标点差异的候选配对; - 移除重复项; - 针对存在多个候选简单句版本的复合句,将所有简单句版本合并为一条新条目(假设该复合句通过拆分为2个或更多简单句完成简化); - 针对每一条复合句,保留与原句莱文斯坦编辑距离(Levenshtein edit-distance)最小的简单句版本。 ## 匿名化处理 为发布本数据集,我们对其进行了如下匿名化处理: 具体移除与替换内容如下: 1. 姓名(包括诸如市政秘书(gemeentesecretaris)等公职人员姓名)均替换为[NAME]; 2. 机构(阿姆斯特丹市政当局(Gemeente Amsterdam)、荷兰公共卫生服务中心(GGD)、荷兰国家公共卫生与环境研究所(RIVM)等公共机构除外)均替换为[ORGANIZATION]; 3. 包含精确街道与门牌号的地址均替换为[ADDRESS]; 4. 电话号码均替换为[NUMBER]; 5. 少量存在隐私或信息安全风险、或包含其他敏感信息的句子,为保证透明度已替换为"xxx xxx xxx"。 最终,匿名化后的数据集未保留源文档的额外信息(例如文档名称),但新增了文档ID,以说明来自同一文档的句子之间的上下文关联。 ## 致谢 本数据集由阿姆斯特丹智能(Amsterdam Intelligence)为阿姆斯特丹市打造。我们特别感谢阿姆斯特丹市通信部门提供的48份原始文档。同时感谢丹尼尔·弗兰蒂斯(Daniel Vlantis)在数据集构建过程中提供的反馈意见,以及基于本数据集开展的大量实验工作。 ## 授权许可 本数据集采用欧洲联盟公共许可证1.2版(EUPL-1.2)进行授权。

创建时间:
2024-03-25
二维码
社区交流群
二维码
科研交流群
商业服务