Pipelines
Versioned DAGs, server-derived execution plans, task attempts and code-backed pipeline projects.
Pipelines are the general orchestration surface in VegaFlow. They cover data transformations, ELT graphs, ML jobs and mixed workflows that do not fit the one-source/one-destination QuickFlow model.
The important boundary is between a definition version and a run. The version contains the authored DAG plus a server-derived execution plan. The run points to that immutable version and records what actually happened.
Definition lifecycle
draft → validated version → published version → run(s)A mutation creates a new version rather than editing a version already used by a run. Publication validates graph structure, component configuration, connection/cluster references and the principal's permissions. The server derives the execution plan; callers cannot submit a trusted plan blob alongside an unverified graph.
DAG rules
- Node IDs are unique inside a version.
- Edges reference existing nodes and valid ports.
- The executable graph is acyclic.
- Required component configuration is validated against the component type.
- Disabled or design-only nodes do not become executable tasks.
- Every execution target and connection is resolved inside the same organization and workspace.
The server owns the plan
A UI can send positions, labels and authored edges. VegaFlow validates the authoritative task order and dependency graph before publishing a runnable version.
Runs, tasks and attempts
A pipeline run owns task runs. A task run owns attempts. This three-level model preserves the difference between “the pipeline failed,” “this task failed,” and “the first attempt failed but retry two worked.”
Pipeline run
├── extract_orders
│ └── attempt 1 / succeeded
├── build_features
│ ├── attempt 1 / failed
│ └── attempt 2 / succeeded
└── train_model
└── attempt 1 / runningTransitions are guarded. A terminal attempt cannot return to running, an older attempt cannot overwrite a newer one, and the pipeline status is derived from its current task states.
Code-backed projects
Pipeline projects can bind a Git integration and a revision. A published version records the resolved commit, entry point and build identity so a retry does not suddenly execute the head of a moved branch. Git push webhooks can trigger a validation or release workflow according to the integration policy.
Use code-backed pipelines for real transformation or ML projects. Keep passwords and platform credentials in managed connections/secrets, not in the repository or pipeline parameter defaults.
Scheduling and events
Pipelines can start from a cron schedule, an explicit run request or a configured external event/webhook. The trigger creates a normal durable run. It does not bypass version selection, permission checks or cluster placement.
Operating a failure
- Open the parent run and find the first failed task on the critical path.
- Inspect its latest attempt, structured error, selected environment and logs.
- Decide whether the failure is deterministic (bad SQL/config) or transient (capacity/network/service).
- Publish a new definition version for deterministic changes. Do not mutate history.
- Retry only the failed work when dependency outputs remain valid; otherwise create a new run.
An orchestration system earns its keep during this part, not when every box is green.