Collaborative AIaaS Composition Dataset
收藏资源简介:
Collaborative AIaaS Composition Dataset是由澳大利亚科廷大学构建的专为协作式人工智能即服务(AIaaS)组合研究设计的大规模数据集。该数据集整合了来自5个提供商的25,900个AIaaS服务,覆盖分类、回归、检测、聚类和自然语言处理等12个AI任务族,并包含10,000个具有多样化目标、QoS约束和偏好的协作服务请求。数据集创建过程包括:从主流平台收集预训练和微调模型,利用基准数据集评估服务的功能与QoS属性(如准确率、延迟、可靠性),并通过扰动-拒绝算法生成真实的服务需求,最后采用多臂老虎机(MAB)算法确定服务组合的可组合性并生成组合解决方案。该数据集旨在解决现有数据集缺乏AI特定属性和协作组合支持的问题,可用于服务推荐、选择、QoS预测以及协作AIaaS组合方法的评估与基准测试。
Collaborative AIaaS Composition Dataset is a large-scale dataset constructed by Curtin University of Australia, specifically designed for research on collaborative Artificial Intelligence as a Service (AIaaS) composition. This dataset integrates 25,900 AIaaS services from 5 providers, covering 12 AI task families including classification, regression, detection, clustering, natural language processing and other common AI tasks. It also contains 10,000 collaborative service requests with diverse objectives, QoS constraints and preferences. The dataset construction process includes: collecting pre-trained and fine-tuned models from mainstream platforms, evaluating the functional and QoS attributes of the services (such as accuracy, latency, reliability) using benchmark datasets, generating realistic service requirements via the perturbation-rejection algorithm, and finally determining the composability of service compositions and generating composite solutions using the Multi-Armed Bandit (MAB) algorithm. This dataset aims to address the gap that existing datasets lack AI-specific attributes and support for collaborative composition, and can be applied to service recommendation, service selection, QoS prediction, as well as the evaluation and benchmark testing of collaborative AIaaS composition methods.
Collaborative AIaaS Composition Benchmark 数据集概述
数据集来源
- 数据集详情页面地址:https://github.com/deepakkanneganti9/CAIaaS
- 关联论文:A Collaborative Artificial Intelligence as a Service Composition Dataset
数据集内容结构
数据文件
datasets/AIaaS_Service_Dataset.csv:包含基准测试所使用的 AIaaS 服务记录。datasets/update_signatures/:包含用于计算 AIaaS 服务之间接口/模型兼容性的更新签名向量。service_request/AIaaS_Collaborative_Service_Request_Dataset.json:包含基准测试中使用的协作服务请求。
基准测试代码
benchmark/canonical_full_benchmark.py:包含完整的基准测试实现,涵盖目标计算和实验中使用的优化技术。benchmark/plot_final_average_quality.py:从归档的基准测试结果文件中重新生成平均服务质量图。
输出文件
outputs/AIaaS_Composition_Benchmark_Results.csvoutputs/AIaaS_Composition_Quality_by_Length.csvoutputs/AIaaS_Composition_Quality_Comparison.pdfoutputs/AIaaS_Composition_Quality_Comparison.png
基准测试优化技术
基准测试包含以下优化技术:
- Random Search
- Greedy
- Epsilon-Greedy
- Genetic Algorithm (GA)
- DAAGA
- MWOA
- CSSA
- SDFGA
- BPSC-GA
- PK-IDPSO
其中,MAB 方法被用作所提出的解生成策略,在最终平均质量比较图中不作为基线绘制。
可组合性目标
基准测试使用每个服务请求中的偏好权重计算加权协作可组合性分数,所实现的项包括:
quality score:归一化任务性能,对应准确性或任务质量组件。response time:使用请求延迟阈值和所选服务延迟之和的有界延迟分数。tail latency:使用请求尾部延迟阈值和所选服务尾部延迟之和的有界尾部延迟分数。resource cost score:请求权重中使用的资源/成本组件。interface compatibility:使用余弦相似度从更新签名向量计算的成对兼容性。update signature availability:所选服务的更新签名信息的可用性。
最终分数是服务请求 preference_weights 中列出的活跃项的加权和。服务数据集还包含平均计算时间和可靠性分数等附加属性,这些属性保留在数据集中,但除非包含在服务请求权重中,否则不作为此基准测试的活跃项。
复现与重运行
重新生成最终图
在包含 AIaaS_Collaborative_Composition_Benchmark/ 的目录下运行:
bash
python3 -m AIaaS_Collaborative_Composition_Benchmark.benchmark.plot_final_average_quality
重新生成:
outputs/AIaaS_Composition_Quality_by_Length.csvoutputs/AIaaS_Composition_Quality_Comparison.pdfoutputs/AIaaS_Composition_Quality_Comparison.png
重运行完整优化
在包含 AIaaS_Collaborative_Composition_Benchmark/ 的目录下运行:
bash
python3 -m AIaaS_Collaborative_Composition_Benchmark.benchmark.canonical_full_benchmark
重新计算基于签名的完整协作基准测试并写入:
outputs/canonical_recomputed_full_benchmark.csvoutputs/canonical_recomputed_full_benchmark_summary.json
完整优化运行较慢,因为它评估所有协作服务请求在所有基准技术上的表现,并重新计算基于签名的接口兼容性。执行过程中会打印进度消息,例如:processed 100/2000 collaborative requests
注意事项
datasets/update_signatures/目录是自包含完整优化重运行所必需的。移除此目录可减小目录大小,但完整基准测试将无法计算基于签名的接口兼容性。




