遇见数据集

AHID-CN Dataset v0.2

收藏
Zenodo2026-08-11 更新2026-08-13 收录
官方服务:

资源简介:

AHID-CN is the Chinese-language corpus of the Animal Harm Incident Database (AHID), an open-source project that discovers, archives, deduplicates, and cross-checks publicly available information about animal harm events. This release contains two sub-corpora that address different questions and are deliberately kept separate rather than merged into a single table. The first sub-corpus covers 30 incidents reported in Chinese-language public sources, collected via manual URL-driven backfill (not automated platform scraping) and processed through a rule-based pipeline: archiving, source-dependency analysis, claim extraction, and evidence-sufficiency scoring. A rule engine, not a language model, decides what gets published. It has broad coverage of whatever becomes publicly visible, but no denominator. The second sub-corpus, added in this version, is a systematic census of 490 animal-harm criminal judgments drawn from CAIL2018 — a frozen, downloadable research corpus of 2.68 million Chinese criminal judgments. Because the search procedure and source corpus are both fixed and public, this sub-corpus has an explicit denominator: anyone can rerun the same search against the same data and get the same candidate set. It covers only what happened to result in a criminal prosecution, which the incident table does not guarantee, and its single largest finding is that in the absence of a dedicated anti-cruelty statute, intentional animal harm never enters the criminal record under its own name — it is prosecuted under ordinary charges, theft and robbery first among them. The project's methodological lineage follows the AI Incident Database (AIID): both address the same underlying problem — public information that is scattered, prone to deletion, easily misattributed, and costly to get wrong — with a reproducible, source-traceable, uncertainty-disclosing structure. What's included Five tables, distributed as CSV files: - incidents_public.csv (30 rows) — one row per incident, including date/location precision, animal category, harm type, an automated evidentiary status (A1-A4/AX/AF), and a 0-100 evidence sufficiency score- sources_public.csv (53 rows) — every archived source, including ones later found unavailable; source tier, independence status, and archival status are retained rather than filtered out- claims_public.csv (151 rows) — every extracted, individually checkable factual claim, including claims marked contradicted or claimed-only- responses_public.csv (34 rows) — institutional responses (police, courts, schools, agencies) decomposed from claims_public- judgments_census.csv (490 rows) — one row per judgment, including the full court-established fact text, charge, and per-row evidence-transparency flags distinguishing court-documented harm from presumed or unverified cases The first four tables share an incident_id key and belong to the reported-incident pipeline; judgments_census.csv is an independent flat table with its own census_id key and does not run the evidence-sufficiency scoring engine because a single authoritative court judgment does not need the multi-source independence scoring built for socially sourced claims. All method documentation ships with the package in both Chinese and English — the four Chinese-language documents include post-edited English translations, with the Chinese files authoritative when the two diverge. Known limitations (full details in documentation/known_limitations.md) The incident table is a 30-incident pilot release, not a stable-rate corpus, and should not be used to infer relative incidence rates across regions or groups. Evidence-scoring weights are an uncalibrated v0. A single-rater gold-standard audit found zero content errors across all 30 incidents but has not yet produced a publishable inter-rater reliability statistic. For minors, the dataset records involvement and the institutional response but does not publish identifiers that would locate the specific child (name, face, school, class, or sub-district address); adults are identified only to the extent the original primary source itself already did. The judgment census retains CAIL2018's own name-masking and inherits its limitations: no case number, court, or date; coverage ends at 2018; single-defendant cases only. Roughly 17% of its rows are the same underlying case appearing in more than one CAIL2018 competition split — kept deliberately unmerged, so counts require deduplication first. License All five CSV tables are licensed CC BY-SA 4.0. Third-party source article text and media referenced by the incident tables are not redistributed and remain with their original rights holders. The judgment census's full-text column is republished from CAIL2018's own public release of official court documents, which are not third-party media. Links Source code, pipeline, and documentation: https://github.com/nanyi-deng/animal-harm-incident-databaseThe concept DOI 10.5281/zenodo.21462311 always resolves to the latest version.

提供机构:
Zenodo
创建时间:
2026-08-11
二维码
社区交流群
二维码
科研交流群
商业服务