mssense Evaluation Benchmark for Closed-Vocabulary Action Trace Generation
收藏资源简介:
Evaluation-only benchmark (1865 samples, 9 task families) for closed-vocabulary action trace generation in conversational Robotic Process Automation (RPA) authoring: clarification policy, LAT audit, semantic judgment, workflow creation, business-rule extraction, visual grounding/governance, modification intent, audit, and interaction regression. Each sample pairs a conversational request with oracle labels for judging whether a generated action trace is executable against a closed, typed, channel-specific action catalogue. Includes a JSON schema, a Datasheet for Datasets, the evaluation protocol, and provenance/licensing documentation. This public release is privacy-sanitized. Supports the manuscript "Closed-Vocabulary Action Trace Generation for Conversational RPA Authoring" (Novelis).



