OmegaUse-OfficeVal
收藏资源简介:
OmegaUse-OfficeVal是一个基准数据集,用于评估大语言模型智能体在真实办公套件长时程任务上的表现。其核心特点是经济基础性,每个任务都标注了人力劳动时间和任务价格代理,以衡量任务完成度并分析人力成本和经济价值。数据集包含100个从真实办公场景收集的任务,涵盖文字处理文档、电子表格、演示文稿及跨文件工作流,任务来源包括从业者真实需求和自由职业平台。每个任务提供高层级用户指令和输入工件,要求智能体生成最终交付物(如Word文档、电子表格或演示文稿),评估聚焦于最终交付物的质量和正确性。数据规模约216个输入文件(平均每个任务约2.2个)和110个最终输出文件,输入为多模态,包括文本指令、Office文档、图像和视频,任务指令和材料主要为中文,元数据和文档为英文。任务按输出格式分为五类:Word(37个)、PowerPoint(39个)、Excel(19个)、PDF(4个)和跨文件(Word + Excel,1个)。评估采用基于代码的确定性验证器进行两阶段评分,包括可用性维度和任务完成维度。数据构建经过多阶段任务适应流程,确保隐私保护、可执行性和真实性。
OmegaUse-OfficeVal is a benchmark dataset for evaluating the performance of Large Language Model (LLM) AI Agents on long-duration real-world office suite tasks. The core characteristic of this dataset lies in its economic grounding: each task is annotated with two complementary signals, namely man-hour labor time and task price proxy, enabling the evaluation to not only measure task completion but also analyze the associated labor costs and economic value of the task. The dataset contains 100 tasks collected from real office scenarios, covering word processing documents, spreadsheets, presentations, and cross-file workflows. These tasks originate from real office requirements proposed by practitioners and freelance platforms, ensuring the benchmark is grounded in realistic economic demands. Each task provides high-level user instructions and input artifacts, requiring the AI Agent to generate final deliverables (such as Word documents, spreadsheets, or presentations). The evaluation focuses on the quality and correctness of the final deliverables, rather than the specific execution trajectories. In terms of data scale, the dataset contains approximately 216 input files (averaging about 2.2 per task) and 110 final output files. The inputs are multimodal, including text instructions, Office documents, images, and videos, as many critical details in real-world office requests are explicitly conveyed through embedded media rather than plain text. Task instructions and materials are primarily in Chinese, while metadata and supporting documents are in English. Tasks are categorized into five types based on the required output format: Word (37), PowerPoint (39), Excel (19), PDF (4), and cross-file (Word + Excel, 1). The distribution of input files and golden standard artifacts covers various formats such as video/audio, images, DOCX, PPTX, XLSX, and PDF. The evaluation uses a code-based deterministic validator to perform two-stage scoring on final output files: first, conduct availability checks (including file format, openability, content integrity, etc.), then carry out fine-grained, user-centric scoring based on task completion dimensions (covering content, format, structure, numerical accuracy, layout, instruction adherence, etc.). The dataset contains a total of approximately 381 availability check items and 2,236 task completion scoring items. The dataset construction follows a multi-stage task adaptation workflow, including task collection, funnel screening, privacy-preserving adaptation, input reconstruction and de-identification, scoring criterion generation and code-based validator creation, and final acceptance review. This ensures that the tasks have no privacy risks, can be executed normally, and conform to the language, structure, and artifact formats of real office scenarios.
OmegaUse-OfficeVal 数据集概述
基本信息
- 数据集名称: OmegaUse-OfficeVal
- 许可证: Apache-2.0
- 语言: 中文和英文
- 规模: 100 个样本(少于 1K)
- 领域: Office 套件生产力任务(Word、PowerPoint、Excel、PDF 及跨文件工作流)
- 任务类别: 其他(办公代理、LLM 代理、计算机使用、经济接地、文档处理、电子表格、演示文稿)
数据集内容
该基准测试用于评估 LLM 代理在长周期、真实世界办公套件任务上的表现,涵盖文字处理文档、电子表格、演示文稿及跨文件生产力工作流。任务来源于从业者提出的真实办公需求和自由职业平台,具有真实经济需求基础。每个任务提供高层次的用户指令和输入工件,要求代理生成最终交付物(如 Word 文档、电子表格或演示文稿)。评估聚焦于最终交付物的质量和正确性,而非具体执行轨迹。
核心特性
- 经济接地: 每个任务标注了人类劳动时间和任务价格代理两个补充信号,可从任务完成度、人类努力和经济价值三方面分析代理性能。
- 数据规模: 100 个长周期任务,220 个输入文件(平均每个任务 2.2 个文件),115 个金色工件。
- 模态: 文本指令、Office 文档、PDF、图像、音频和视频。
- 评估方式: 基于最终交付工件的确定性、基于代码的验证器。
- 双语: 中文和英文的任务定义与评估标准。
数据分布
根据所需输出格式,任务分为五类:
| 输出类型 | 任务数 | 占比 |
|---|---|---|
| Word | 37 | 37% |
| PowerPoint | 39 | 39% |
| Excel | 19 | 19% |
| 4 | 4% | |
| 跨文件(Word + Excel) | 1 | 1% |
| 总计 | 100 | 100% |
输入文件和金色工件分布:
| 文件类型 | 输入文件数 | 金色工件数 |
|---|---|---|
| 视频/音频 | 10 | 0 |
| 图像 | 77 | 0 |
| DOCX | 63 | 48 |
| PPTX | 31 | 40 |
| XLSX | 25 | 24 |
| 14 | 3 | |
| 总计 | 220 | 115 |
任务组成
每个任务包含五个组成部分:
- 指令: 用户请求的文本描述及具体要求。
- 输入文件: 完成任务所需的文件集合(托管在数据集中)。
- 人类劳动时间: 人类工人在无 LLM 辅助下完成任务的记录时间(以分钟为单位)。
- 任务价格代理: 任务级别的价格信号,估算完成该任务的市场价格(以人民币元为单位)。
- 代码验证器: 由指令和评估标准派生出的可执行评估代码,为候选工件分配分数。
经济接地机制
- 人类劳动时间: 每个任务由至少两名招募的标注者在质量把控协议下完成;当两次有效完成时间差异较大时,增加第三名标注者。报告的人类劳动时间为最短两次有效完成时间的平均值,以减少异常慢速尝试的影响。
- 任务价格代理: 采用混合策略。具有明确价格信号的任务直接使用从业者提供的价格(
price_source = explicit_price);无明确价格的任务由三名领域专家独立估算,并通过一致性基础程序聚合估计值,剔除明显异常值(price_source = estimated_price)。
评估协议
- 评估方法: 使用确定性、基于代码的验证器对最终输出文件进行评估,不强制固定执行轨迹。代理可通过 GUI 操作、脚本、API 或混合策略完成任务。
- 评估维度:
- 可用性评估(Dim-1): 检查文件格式是否正确、文件可正常打开、内容/布局未被严重损坏、工件保持可编辑性。若任一可用性项目失败,任务得分为零,不再评估完成度。
- 任务完成度评估(Dim-2): 细粒度、以用户为中心的评分点,涵盖内容、格式、结构、数值准确性、布局和指令遵循。正向项目奖励正确完成的要求(权重 +1/+3/+5);负向项目惩罚非预期更改或可避免的损坏(对称惩罚 -1/-3/-5)。
- 评分计算: 工件必须首先通过所有可用性项目;然后基于加权正负评估项目计算归一化任务完成度分数,原始分数下限截断为零。
评估标准统计:
| 评估维度 | 总项目数 | 平均每任务 |
|---|---|---|
| 可用性检查(Dim-1) | 219 | 2.19 |
| 任务完成度项目(Dim-2) | 2009 | 20.09 |
数据构建与隐私
数据集通过多阶段任务适配流程构建:
- 任务收集: 从业者提出基于日常工作流程的真实办公套件任务及代表性样例材料。
- 漏斗筛选: 初始任务池逐步筛选,仅保留基于真实需求、交付物明确且具有足够非平凡性和长周期性的任务。
- 隐私保护适配: 重写指令以移除敏感或可识别信息,同时保留原始用户意图、约束和任务关键细节。移除主观、不可验证的要求,保留自然口语化表达。
- 输入重建与去标识化: 输入文件在 LLM 辅助下重建并人工修订,消除隐私/版权风险,修复布局和一致性问题,保持工件与预期任务一致。
- 评估标准生成与代码验证器: 评估标准在 LLM 辅助下生成并由专家迭代精炼,转换为可执行代码验证器,解决人类与代码判断差异。
- 最终验收审查: 三名高级专家确认每个任务无隐私风险、可正常执行、符合真实办公场景的语言、结构和工件格式;仅当三位专家一致同意时任务才被纳入。
筛选阶段统计:
| 漏斗阶段 | 保留任务数 |
|---|---|
| 初始池 | 1,715 |
| 筛选阶段 1 | 595 |
| 筛选阶段 2 | 282 |
| 最终基准 | 100 |
预期用途与限制
OmegaUse-OfficeVal 旨在作为可复现、经济接地的大规模测试平台,用于衡量 LLM 代理在日常办公生产力方面取得可靠进展的程度。它不用于衡量 LLM 代理能否替代整个职业;而是评估工人在有能力代理协助下的实际工作场景。由于负面失败模式无法穷举,分数下限为零,应将其解读为交付物质量的相对度量。
引用
bibtex @inproceedings{omegause_officeval_2026, title = {OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding}, booktitle = {tech report}, year = {July, 2026} }





