Pipelines

Versioned DAGs, server-derived execution plans, task attempts and code-backed pipeline projects.

Pipelines are the general orchestration surface in VegaFlow. They cover data transformations, ELT graphs, ML jobs and mixed workflows that do not fit the one-source/one-destination QuickFlow model.

The important boundary is between a definition version and a run. The version contains the authored DAG plus a server-derived execution plan. The run points to that immutable version and records what actually happened.

Definition lifecycle

draft → validated version → published version → run(s)

A mutation creates a new version rather than editing a version already used by a run. Publication validates graph structure, component configuration, connection/cluster references and the principal's permissions. The server derives the execution plan; callers cannot submit a trusted plan blob alongside an unverified graph.

DAG rules

  • Node IDs are unique inside a version.
  • Edges reference existing nodes and valid ports.
  • The executable graph is acyclic.
  • Required component configuration is validated against the component type.
  • Disabled or design-only nodes do not become executable tasks.
  • Every execution target and connection is resolved inside the same organization and workspace.

The server owns the plan

A UI can send positions, labels and authored edges. VegaFlow validates the authoritative task order and dependency graph before publishing a runnable version.

Runs, tasks and attempts

A pipeline run owns task runs. A task run owns attempts. This three-level model preserves the difference between “the pipeline failed,” “this task failed,” and “the first attempt failed but retry two worked.”

Pipeline run
  ├── extract_orders
  │     └── attempt 1 / succeeded
  ├── build_features
  │     ├── attempt 1 / failed
  │     └── attempt 2 / succeeded
  └── train_model
        └── attempt 1 / running

Transitions are guarded. A terminal attempt cannot return to running, an older attempt cannot overwrite a newer one, and the pipeline status is derived from its current task states.

Code-backed projects

Pipeline projects can bind a Git integration and a revision. A published version records the resolved commit, entry point and build identity so a retry does not suddenly execute the head of a moved branch. Git push webhooks can trigger a validation or release workflow according to the integration policy.

Use code-backed pipelines for real transformation or ML projects. Keep passwords and platform credentials in managed connections/secrets, not in the repository or pipeline parameter defaults.

Scheduling and events

Pipelines can start from a cron schedule, an explicit run request or a configured external event/webhook. The trigger creates a normal durable run. It does not bypass version selection, permission checks or cluster placement.

Operating a failure

  1. Open the parent run and find the first failed task on the critical path.
  2. Inspect its latest attempt, structured error, selected environment and logs.
  3. Decide whether the failure is deterministic (bad SQL/config) or transient (capacity/network/service).
  4. Publish a new definition version for deterministic changes. Do not mutate history.
  5. Retry only the failed work when dependency outputs remain valid; otherwise create a new run.

An orchestration system earns its keep during this part, not when every box is green.

Vegalake, VegaDB and VegaFlow are trademarks or registered trademarks of Vegalake Inc.

On this page