遇见数据集

Dataset for Doctoral Dissertation: Artificial Intelligence-Based Sentencing Prediction and Judgment Text Generation

收藏
Zenodo2025-07-31 更新2026-05-26 收录
官方服务:

资源简介:

Judgments related to guilty verdicts in public insult cases were collected from three Taiwan district courts (Taipei, New Taipei, and Shilin) for the period from January 1, 2015, to April 30, 2025, using the Taiwan Judicial Yuan Law and Regulations Retrieving System. To ensure data consistency, judgments were excluded if any of the following conditions applied: the main text indicated recidivism; the judgment or related prosecutorial documents lacked a crime-facts paragraph; the judgment lacked a sentencing paragraph; the conviction resulted from plea bargaining; the case involved co-defendants or additional criminal offenses; or the same defendant was sentenced separately for multiple counts of public insult. After all filters were applied, 1,858 judgments remained for analysis. After collecting the judgments, the following information was extracted: ・Crime Facts for Sentencing Prediction The crime facts paragraph includes the core descriptive elements needed for the sentencing prediction task, such as the time and location of the offense, the identities of the complainant and defendant, and the motive, purpose, and method of the crime. Personal names were anonymized to protect privacy and ensure consistency suitable for the model. Specifically, the complainant’s name was replaced with “Complainant,” the defendant’s (or victim’s) name with “Defendant,” and all other individuals’ names with “Third person,” or, when more precision was needed, with a role-based description such as “Defendant’s mother,” “Defendant’s son,” or “Defendant’s boyfriend.” If the complainant was a legal entity, the original corporate name was kept. When multiple complainants appeared in the same case, sequential labels like “Complainant 1,” “Complainant 2,” etc., were assigned in this study. Address information was normalized by replacing the white circle placeholder character “○” with the Chinese character “某,” while any digits in house or floor numbers that had been masked with the digit “0” to anonymize the address were retained, thus preserving the original numeric format while concealing the true location. Formulaic legal expressions—such as “基於公然侮辱之犯意” (with intent to publicly insult), “不特定多數人得以共見共聞” (observable to an indeterminate public), and “足以貶損告訴人之人格及社會評價” (sufficient to damage the complainant’s reputation)—were removed because, as highlighted in recent sentencing prediction research, the input should only include information available before the court pronounces judgment. These expressions appear exclusively in the written reasons of an adjudicated decision and would therefore leak post-decision language into the model, undermining the validity of pre-judgment prediction. Text referring to other criminal charges (whether dismissed or prosecuted separately), errata paragraphs regarding applications for summary judgment, and descriptions of how the complainant discovered the offense or which agency handled the investigation were also removed. Finally, the cleaned and standardized narratives were stored in the column “Crime Facts for Sentencing Prediction.” This preprocessing guarantees that each entry includes only the factual information relevant to predicting the type and severity of punishment, excluding identifying details and legally redundant phrasing. ・Objective Facts Influencing Sentencing Objective information was extracted from the sentencing paragraph while excluding any subjective judicial assessments. The retained details include the defendant’s criminal record, confession, apology, any reconciliation or mediation with the complainant, and background factors such as education, occupation, income, marital status, and family circumstances. When the sentencing paragraph contained the personal names of the parties or third persons, those names were anonymized following the same rules used in “Crime Facts for Sentencing Prediction”—that is, replacing them with “Complainant,” “Defendant,” “Third person,” or a role-specific label (e.g., Defendant’s mother). All processed data were stored in the column “Objective Facts Influencing Sentencing.” ・Judge’s Name Each judge’s name was extracted and stored in the “Judge’s Name” column. ・Sentence Type The sentence type (detention or fine) was extracted from the judgment texts and stored in the “Sentence Type” column. The category “fine” was then encoded as integer 0, and “detention” as integer 1. Subsequently, the “Crime Facts for Sentencing Prediction,” “Objective Facts Influencing Sentencing,” and “Judge’s Name” columns were concatenated, in that order, to create a single “Input Text” column. Combined with the “Sentence Type” column as the class label, this configuration yielded the dataset for the sentencing-prediction task, which is designed as a binary classification (fine versus detention). For the factual paragraph generation task, the “Crime Facts for Sentencing Prediction” column was used as the input text. The corresponding target text was extracted from each judgment’s crime facts paragraph after applying the same cleaning and anonymization rules, but retaining formulaic legal expressions (e.g., “基於公然侮辱之犯意” (with intent to publicly insult), “不特定多數人得以共見共聞” (with intent to publicly insult), and “足以貶損告訴人之人格及社會評價” (sufficient to damage the complainant’s reputation)). Merging these input–output pairs created the dataset for the factual paragraph generation task, enabling the model to transform objective crime facts into formally worded crime facts paragraphs that reflect the conventional drafting style of Taiwanese criminal judgments. For the sentencing paragraph generation task, the input text was created by combining the “Objective Facts Influencing Sentencing” column with the brief statement of crime facts that precedes the reasoning section in the original sentencing paragraph. The corresponding target text was derived from the same sentencing paragraph after replacing the complainant’s name with “Complainant,” the defendant’s with “Defendant,” and any other individuals’ names with “Third person” or role-based labels (e.g., “Defendant’s mother”), and removing documentary cross-references to exhibits, page numbers, and similar citations—such as “(見臺灣高等法院被告前案紀錄表)”(see High Court prior offense record), “(見本院卷第25頁)”(see case file, p. 25), or “(見偵字卷第7頁)”(see investigation dossier, p. 7). These input–output pairs were then combined to create the dataset used for the sentencing paragraph generation task. Given the limited number of imprisonment sentences for public insult cases, such judgments were excluded, thereby restricting the task to a binary classification between fines and detention sentences.

提供机构:
Zenodo
创建时间:
2025-07-31
二维码
社区交流群
二维码
科研交流群
商业服务