遇见数据集

A Distributed Data Processing Scheme Based on Hadoop for Synchrotron Radiation Experiments

收藏
DataCite Commons2025-04-27 更新2025-05-18 收录
官方服务:

资源简介:

This research presents a case study on synchrotron radiation biomolecular crystallography to illustrate a beamline distributed data processing scheme based on the Hadoop ecosystem.We build a distributed file storage system for experimental crystallography data based on Hadoop HDFS. Additionally, we develop a resource scheduling system for the cluster using Hadoop YARN. Furthermore, we design and develop a distributed automated data processing pipeline(Spark-DIALS) for crystallography by combining Hadoop Spark and DIALS. The dials_spot_finder.py and dials_integrate_run.py contain the source code transforming the spots finding and integrate of original DIALS.Moreover, the solution utilizes FastAPI to deploy each functional module in a distributed microservice architecture.There are primarily microservices related to Spark distributed automatic processing jobs and HBase data table operations(sparkJobApi.py&hbaseApi.py).

本研究以同步辐射生物分子晶体学为案例,阐释了一种基于Hadoop生态系统的光束线分布式数据处理方案。我们基于Hadoop HDFS搭建了面向实验晶体学数据的分布式文件存储系统。此外,我们借助Hadoop YARN开发了集群资源调度系统。进一步地,我们结合Hadoop Spark与DIALS,设计并开发了面向晶体学研究的分布式自动化数据处理流水线(Spark-DIALS)。dials_spot_finder.py与dials_integrate_run.py为封装了原始DIALS斑点识别与数据积分功能的源代码文件。本方案采用FastAPI将各功能模块部署为分布式微服务架构,核心微服务主要涵盖Spark分布式自动处理任务(sparkJobApi.py)与HBase数据表操作(hbaseApi.py)两类。

提供机构:
Science Data Bank
创建时间:
2023-11-06
搜集汇总
数据集介绍
A Distributed Data Processing Scheme Based on Hadoop for Synchrotron Radiation Experiments 数据集图片
背景与挑战
背景概述
该数据集提供了一种基于Hadoop生态系统的分布式数据处理方案,专门针对同步辐射生物分子晶体学实验,集成了HDFS存储、YARN资源调度和Spark-DIALS自动处理管道,并采用FastAPI微服务架构实现模块化部署。其核心内容包括相关源代码,涉及数据处理、Hadoop等技术,适用于物理和计算机科学领域的研究。数据集规模较小,包含4个文件,总数据量为10.20 KB,发布于2023年11月6日。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务