Paired human and LLM-as-judge annotations of base-model continuations in English, Chinese and Czech (4,860 outputs)
收藏资源简介:
Paired human and LLM-as-judge annotations of base-model continuations in English, Chinese and Czech. The dataset contains all 4,860 rated outputs (18 sentences × 3 languages × 3 models × 3 temperatures × 10 repetitions), each with the input prompt, the raw model continuation, the paired human and automatic ratings, and the raw rationales of the automatic judge. The coding books (the full test battery in all three languages) are included as a separate PDF. Columns of the sheet "annotations": language – Language of the prompt (English / Chinese / Czech) model – Base model (DeepSeek = deepseek-ai/deepseek-67b-base; Llama = meta-llama/Meta-Llama-3.1-405B; Qwen = Qwen/Qwen2.5-72B) sentence_id – Sentence of the test battery (1–18) with language code (A = English, Z = Chinese, C = Czech); see the coding books temperature – Sampling temperature (0.0 / 0.1 / 0.3) output_id – Output code: (L,Q,D)(1–18)g(A,Z,C)(repetition 1–10)t(0,I,III) prompt – The input sentence exactly as sent to the model, in the language of the battery model_continuation – The raw generated continuation of the model (unedited; may contain LaTeX, dialogue or translation artefacts) cohesion_human – Linguistic cohesion, human annotator (1 = cohesive, 0 = not cohesive) coherence_human – Cultural coherence, human annotator (1 = coherent, 0 = not coherent) cohesion_sonnet – Linguistic cohesion, Claude Sonnet 4.6 / LLM-as-judge (1/0) coherence_sonnet – Cultural coherence, Claude Sonnet 4.6 / LLM-as-judge (1/0) cohesion_rationale_sonnet – Raw, unedited rationale of the automatic judge for the cohesion rating coherence_rationale_sonnet – Raw, unedited rationale of the automatic judge for the coherence rating; for 2 outputs (L16gC4t0, L16gC6t0) no rationale was returned, marked "[no rationale returned by the judge]" The model outputs were generated in July 2025 (Replicate: DeepSeek, Qwen; Hyperbolic: Llama); the automatic annotation was performed in May 2026. All statistics reported in the associated article (Cohen's kappa, observed agreement, MAD, Wilson confidence intervals, Wilcoxon signed-rank tests) can be reproduced from this dataset. The written rationales of the human annotator are available from the contact person on reasonable request.



