遇见数据集

Prediction-Level Benchmark Metadata and Model Scores for Multi-Provider AI-Text Detection in Turkish Scholarly Abstracts

收藏
Mendeley Data2026-08-09 收录
官方服务:

资源简介:

This dataset contains prediction-level experimental metadata and derived model scores from a frozen benchmark used to evaluate AI-text detection in Turkish scholarly abstracts. The release comprises 1,678 benchmark records grouped into 60 independent source clusters: 60 human records and metadata for 1,618 AI-assisted variants. The benchmark covers seven academic domains, four generative-AI providers (OpenAI, Gemini, Claude, and DeepSeek), six prompt families, four contribution levels, and 179 adversarial transformations. Each record includes provenance identifiers, source-cluster identifiers, binary labels, contribution metadata, provider and prompt metadata, domain and document-type variables, adversarial-operation metadata, text hashes, word counts, and three frozen model-score fields. The repository also provides raw and processed tabular releases, a data dictionary, source registry, quality-control reports, descriptive tables, authoritative results registries, deterministic validation code, statistical-analysis code, figure-generation code, software requirements, licenses, citation metadata, and SHA-256 checksums. The release is designed to support reproducible investigation of source leakage, cross-domain reliability, human false-positive rates, provider and prompt variation, calibration, threshold sensitivity, and prediction-level information quality. No full copyrighted scholarly abstracts and no full LLM-generated texts are redistributed. The repository contains experimental metadata, numerical prediction outputs, provenance identifiers, truncated content hashes, and reproducibility materials. Dataset accuracy is verified through deterministic Python checks and cryptographic checksums rather than by generative AI.

创建时间:
2026-07-28
二维码
社区交流群
二维码
科研交流群
商业服务