HAPI
收藏资源简介:
HAPI(API历史)是一个包含1,761,417个实例的纵向数据集,记录了从2020年到2022年商业ML API的应用情况,涉及亚马逊、谷歌、IBM、微软等提供的API。该数据集覆盖了图像标记、语音识别、文本挖掘等多种任务。HAPI是首个系统研究ML API应用的大规模数据集,为ML-as-a-service领域的研究提供了宝贵资源。通过分析HAPI,研究者可以深入了解API性能随时间的变化,以及不同API在处理相同任务时的性能差异,从而为ML API的选择和使用提供科学依据。
HAPI (API History) is a longitudinal dataset containing 1,761,417 instances, which documents the deployment and usage status of commercial machine learning (ML) APIs between 2020 and 2022, involving APIs provided by Amazon, Google, IBM, Microsoft and other technology vendors. This dataset encompasses a wide range of tasks such as image tagging, speech recognition, text mining and more. As the first large-scale dataset for systematically investigating the applications of ML APIs, HAPI serves as a valuable resource for research in the ML-as-a-service domain. By analyzing HAPI, researchers can gain in-depth insights into the temporal changes of API performance and the performance differences between different APIs when handling the same task, thereby providing scientific basis for the selection and utilization of ML APIs.

- 1HAPI: A Large-scale Longitudinal Dataset of Commercial ML API Predictions斯坦福大学 · 2022年



