LLM Adversarial Evaluation Dataset (Zhang et al., 2025)
收藏资源简介:
This dataset accompanies the paper “Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities” (Zhang et al., 2025).It contains behavioral and simulated data from experiments evaluating the decision-making robustness of large language models (LLMs) under adversarial and dynamic conditions. The dataset includes results from two canonical paradigms: Two-Armed Bandit Task: tests exploration-exploitation balance across different models and decoding settings (e.g., temperature, top-p). Multi-Round Trust Task (MRTT): examines cooperative and adaptive decision-making in social exchange between LLMs and adversarial agents. Each file records trial-level choices and rewards, with aggregated performance metrics for human and model comparisons.The dataset supports reproducible behavioral analyses and provides a foundation for studying model-specific susceptibilities to manipulation, rigidity, and fairness recognition.



