遇见数据集

Fine-Tuned LLMs for Turkish Postpartum Care — Benchmark Dataset and Evaluation Data

收藏
Zenodo2026-06-28 更新2026-08-01 收录
官方服务:

资源简介:

Provider-neutral, held-out benchmark of eight fine-tuned large language models (5 OpenAI GPT + 3 Google Gemini) for Turkish postpartum (fourth-trimester) care support. This deposit is the citable Data Availability record for the study. It contains the fine-tuning datasets (one per platform), the held-out gold reference test set, model answers with dual-LLM-judge scores, platform-reported evaluation scores, the summarized analysis results, and the figures reported in the study. Methods (in brief): The primary metric is judge-independent objective accuracy on a single common held-out reference test set, scored against guideline-derived criteria. Statistics: 95% Wilson confidence intervals, omnibus Cochran's Q, pairwise McNemar exact + Holm correction; secondary dual-LLM-judge quality (Friedman + Wilcoxon; non-normal by the ±1.5 skewness/kurtosis rule), inter-judge Cohen's kappa, and platform-vs-local concordance. Self-reported training loss/accuracy are not used to compare vendors (not comparable across tokenizers/loss definitions). Privacy: The deposit contains no API keys or service-account credentials. Account- and project-level identifiers have been removed (OpenAI eval/run IDs replaced with REDACTED; Vertex project number / endpoint IDs with PROJECT / ENDPOINT). Fine-tuned model family names are retained as part of the scientific result. All clinical data are de-identified. Ethics: The system gives no diagnoses, treatments, or doses and refers danger signs to a health facility. It is a complement to, not a replacement for, professional care.

提供机构:
Zenodo
创建时间:
2026-06-28
二维码
社区交流群
二维码
科研交流群
商业服务