遇见数据集

Empathy on a Budget: measurement harness, scenario suite, and results dataset

收藏
Zenodo2026-08-25 更新2026-10-01 收录
官方服务:

资源简介:

A reproducible measurement instrument that captures the human-centred quality of language-model decision support and the system cost of producing it over the same generation, in a single pass: an Empathy–Trust–Accountability rubric scored by a fixed automated judge (gpt-4o-mini), alongside per-inference latency, output throughput and resident memory. The deposit contains the measurement harness; the judge prompts and the full ETA rubric verbatim; twelve high-stakes health-decision scenarios parameterized with published NFHS-5 (2019–21) all-India survey indicators, with sources; the 144-row evaluation dataset (12 scenarios × 4 quantized Llama-3.2 variants × 3 runs, none excluded) from which everyreported number derives, and the analysis script that produces every statistic, table, and figure in the associated article. Reproducibility is bounded, and the bounds are stated in the README: the analysis is deterministic from the released dataset, but the dataset is not deterministic from the models — re-running the benchmark re-runs live generation and live judging, neither of which is seeded. The energy column is present and zero-filled. Model tags are not pinned to manifest digests. The harness is pinned by commit rather than by a release tag. Code is licensed GPL-2.0-or-later; data are licensed CC BY 4.0.

提供机构:
Zenodo
创建时间:
2026-08-25
二维码
社区交流群
二维码
科研交流群
商业服务