Tendem Evaluation Dataset
收藏资源简介:
该数据集包含94个真实世界任务的评估数据,涵盖运营、营销、分析和销售四个领域,包含任务描述、系统输出结果、质量评分和时间成本等信息,用于比较Tendem混合AI+人类系统与ChatGPT Agent和Upwork自由职业者的性能表现
This dataset contains evaluation data for 94 real-world tasks across four domains: operations, marketing, analytics, and sales. It includes task descriptions, system output results, quality scores, and time costs, and is designed to compare the performance of the Tendem hybrid AI+human system, ChatGPT Agent, and Upwork freelancers.
Tendem Evaluation 数据集概述
数据集基本信息
- 数据集名称:Tendem Evaluation
- 数据来源:Tendem混合AI+人类系统评估
- 任务数量:94个真实世界任务
系统架构
- AI代理:执行常规任务(网页浏览、数据处理、文件操作)
- 人类专家:验证结果、处理模糊案例、确保质量
- 多层质量保证:在交付客户前验证每个可交付成果
评估结果对比
性能指标
| 系统 | 质量(良好率) | 中位数时间(小时) | 中位数价格(美元) |
|---|---|---|---|
| Tendem | 74.5% | 16.4 | $32 |
| Upwork | 53.2% | 35.0 | $50 |
| ChatGPT代理 | 40.4% | 0.13 | 订阅制 |
关键发现
- 质量提升:比Upwork高21.3个百分点
- 交付速度:比Upwork快53%
- 成本效益:中位数成本比Upwork低36%
数据集结构
tendem-benchmark/ ├── input_tasks.jsonl # 94个任务描述 ├── output_results.jsonl # 包含质量评级和时间的输出结果 ├── input_files/ # 按任务ID组织的输入文件 │ └── {task_id}/ └── output_files/ # 系统输出文件 ├── chatgpt_agent/ # ChatGPT代理输出 ├── tendem/ # Tendem系统输出 └── upwork/ # Upwork自由职业者输出
质量评估标准
- 良好(可直接交付客户)
- 中等(需要编辑)
- 差(需要返工)
- 拒绝(拒绝处理)
任务分布
94个任务涵盖4个领域:
- 运营(28个):数据收集、格式转换、自动化
- 营销(24个):内容创作、竞争研究
- 分析(22个):数据分析、仪表板、研究
- 销售(20个):联系人数据、数据丰富
相关资源
- 产品官网:https://tendem.ai
- 完整论文:https://toloka.ai/files/tendem_whitepaper.pdf




