Platform overview

How VegaDB, VegaFlow and VegaGraph share data, execution history and context.

Vegalake combines a distributed database, data and AI workflows, and an organizational knowledge graph.

Three products, one system

ProductPrimary jobWhat it contributes to the platform
VegaDBDistributed SQL warehouses over managed and open lakehouse datagoverned data products, query history, schemas, materializations and serving endpoints
VegaFlowConnections, QuickFlows and versioned data/AI pipelinesmovement, transformation, schedules, runs, environments and operational lineage
VegaGraphTyped graph across data, software, infrastructure, people and business meaningownership, context, policy, evidence, dependency, discovery and impact analysis
Clouds + applications + databases + events
                    │
                  VegaFlow
            integrate · transform · run AI
                    │
                  VegaDB
          federate · model · query · serve
                    │
                 VegaGraph
     connect assets · people · software · policy · KPIs

The arrows are not one-way. VegaGraph can show which flow produces a table, which model and service consume it, which team owns the system, which policy applies in production and which KPI is exposed by a proposed change.

A representative path

Connect

VegaFlow creates governed connections to operational systems and discovers their datasets without copying credentials into flow definitions.

Integrate and operate

A QuickFlow continuously syncs selected streams, or a pipeline combines SQL, Python, quality and ML tasks in a versioned DAG.

Store or federate

VegaDB manages analytical tables and queries open DuckLake, Iceberg or Delta data across Amazon, Google Cloud and Azure storage.

Serve

Separate VegaDB warehouses serve BI, applications, transformations and exploration over the same authorized data.

Understand

VegaGraph connects sources, flows, tables, fields, repositories, services, deployments, owners, policies and business concepts.

Questions you can answer with the graph

VegaGraph links pipelines and tables to their dependencies, owners and policies. Use those relationships to answer questions such as:

  • Which customer-facing services and KPIs depend on this field?
  • Which production context is affected, and is staging different?
  • Who owns the source, flow, model and business definition?
  • Which classification or policy applies, and what evidence supports it?
  • What changed between the pipeline version that succeeded and the one that failed?

These relationships let teams investigate a VegaFlow run or VegaDB dataset alongside its owners, dependencies and business impact.

Shared operating model

Organizations contain workspaces. Users and service principals receive scoped permissions. Secrets stay in managed secret objects. Data, connections, warehouses, compute environments, flows and graph entities remain workspace-scoped, with history and provenance preserved across product boundaries.

Vegalake, VegaDB and VegaFlow are trademarks or registered trademarks of Vegalake Inc.

On this page