遇见数据集

Identifying Judicially-Challenged Patents Held by Chinese Listed Companies: A Machine Learning Benchmark with Partial Judicial Labels and Positive-Unlabelled Re-evaluation — v2

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

China Listed-Firm Patent Invalidation Benchmark (v2 — Knowledge-Aware Multi-Classifier Ensemble) A leakage-aware, PU-aware machine-learning benchmark for predicting which patents held by Chinese A-share listed companies are vulnerable to judicial invalidation by the China National Intellectual Property Administration (CNIPA). The v2 release adds a knowledge-aware multi-classifier ensemble framework, an IPC knowledge-graph GraphSAGE embedding, the non-negative PU (nnPU) loss with a group-stratified extension (GS-nnPU), and full five-seed paired-bootstrap inference. Headline result. A stacked ensemble with a gradient-boosted meta-learner over eight base classifiers reaches a pooled patent-level area under the ROC curve of 0.7473 ± 0.0065 (five seeds, StratifiedGroupKFold-by-firm), statistically significantly above the strongest single classifier (Random Forest with the IPC knowledge-graph embedding at 0.7104 ± 0.0077); the paired-bootstrap 95% confidence interval on the gap excludes zero by +0.0173 on the lower bound. Sample. 1,174,359 patents × 122 base features × 32 IPC-knowledge-graph features = 154 features total; 4,035 distinct A-share listed firms; grant years 2008–2022; 545 patents carry a positive judicial-invalidation label. Authors. Tianyi Li (Party School of the Central Committee of C.P.C., Beijing); Xuanyi Zhang (Editorial Department, Management World, Beijing); Xiaoyu Chen (corresponding, Nanyang Business School, NTU, Singapore, ORCID 0000-0002-6442-5142). What is new in v2 (vs. v1): Component v1 (May 2026) v2 (this release) Headline AUC RF at 0.7081 (single seed, leakage-naive) Stacked-XGB ensemble at 0.7473 (5-seed, leakage-aware, paired-bootstrap CI) Feature set 122 structured 122 structured + 32-dim IPC knowledge-graph GraphSAGE embedding Loss function Imbalance-reweighted BCE nnPU (Kiryo et al. 2017) + GS-nnPU group-stratified extension Classifiers reported 7 tabular (single seed) 7 tabular + 3 MLP-loss + 3 ensemble aggregators (logistic-stacked, XGBoost-stacked, agentic) Statistical inference Single-seed point estimates 5-seed × 5-fold StratifiedGroupKFold + paired-bootstrap 95% CI on 9 headline comparisons GS-nnPU theory — Theorem 1 (unbiasedness under SCAR-per-group) + Theorem 2 (variance bound) + full proofs in manuscript Appendix A Package contents. manuscript/ (the v2 paper PDF + LaTeX source + cover letter + highlights + CoI); code/v5f_ensemble/ (15 new Python/shell scripts implementing the ensemble framework); results/v5f_ensemble/ (all five-seed OOF predictions + bootstrap CI JSONs + the 985,904-row IPC knowledge-graph embedding in Parquet format); plus all v1 code, data, and results carried over unchanged for backward compatibility. Licence. MIT for code, CC-BY-4.0 for derived data; CSMAR commercial licence governs raw upstream patent fields. Reproducing the headline. Full pipeline ~18 hours of single-machine CPU; lightweight reproduction (cached artefacts → summary tables) ~5 minutes. See README.md inside the zip for the step-by-step.

提供机构:
Zenodo
创建时间:
2026-05-09
二维码
社区交流群
二维码
科研交流群
商业服务