http://azkaban.github.io/azkaban/docs/2.5/

最新推荐文章于 2023-03-17 11:22:27 发布

转载最新推荐文章于 2023-03-17 11:22:27 发布 · 319 阅读

azkaban 专栏收录该内容

7 篇文章

订阅专栏

Azkaban是一款用于解决Hadoop作业依赖问题的工作流调度系统。它最初为LinkedIn内部需求而开发，随着Hadoop用户的增加，Azkaban逐渐发展成为一个更为强大的解决方案。该系统由三个关键组件构成：关系型数据库(MySQL)、Azkaban Web Server和Azkaban Executor Server。

Azkaban 2.5 Documentation

Overview
Getting Started
Configuration
User Manager
Creating Flows
Using Azkaban
AJAX API
Plugins
Job Types

Overview

Azkaban was implemented at LinkedIn to solve the problem of Hadoop job dependencies. We had jobs that needed to run in order, from ETL jobs to data analytics products.

Initially a single server solution, with the increased number of Hadoop users over the years, Azkaban has evolved to be a more robust solution.

Azkaban consists of 3 key components:

Relational Database (MySQL)
AzkabanWebServer
AzkabanExecutorServer

Relational Database (MySQL)

Azkaban uses MySQL to store much of its state. Both the AzkabanWebServer and the AzkabanExecutorServer access the DB.

How does AzkabanWebServer use the DB?

The web server uses the db for the following reasons:

Project Management - The projects, the permissions on the projects as well as the uploaded files.
Executing Flow State - Keep track of executing flows and which Executor is running them.
Previous Flow/Jobs - Search through previous executions of jobs and flows as well as access their log files.
Scheduler - Keeps the state of the scheduled jobs.
SLA - Keeps all the sla rules