Platform overview
How VegaDB, VegaFlow and VegaGraph share data, execution history and context.
Vegalake combines a distributed database, data and AI workflows, and an organizational knowledge graph.
Three products, one system
| Product | Primary job | What it contributes to the platform |
|---|---|---|
| VegaDB | Distributed SQL warehouses over managed and open lakehouse data | governed data products, query history, schemas, materializations and serving endpoints |
| VegaFlow | Connections, QuickFlows and versioned data/AI pipelines | movement, transformation, schedules, runs, environments and operational lineage |
| VegaGraph | Typed graph across data, software, infrastructure, people and business meaning | ownership, context, policy, evidence, dependency, discovery and impact analysis |
Clouds + applications + databases + events
│
VegaFlow
integrate · transform · run AI
│
VegaDB
federate · model · query · serve
│
VegaGraph
connect assets · people · software · policy · KPIsThe arrows are not one-way. VegaGraph can show which flow produces a table, which model and service consume it, which team owns the system, which policy applies in production and which KPI is exposed by a proposed change.
A representative path
Connect
VegaFlow creates governed connections to operational systems and discovers their datasets without copying credentials into flow definitions.
Integrate and operate
A QuickFlow continuously syncs selected streams, or a pipeline combines SQL, Python, quality and ML tasks in a versioned DAG.
Store or federate
VegaDB manages analytical tables and queries open DuckLake, Iceberg or Delta data across Amazon, Google Cloud and Azure storage.
Serve
Separate VegaDB warehouses serve BI, applications, transformations and exploration over the same authorized data.
Understand
VegaGraph connects sources, flows, tables, fields, repositories, services, deployments, owners, policies and business concepts.
Questions you can answer with the graph
VegaGraph links pipelines and tables to their dependencies, owners and policies. Use those relationships to answer questions such as:
- Which customer-facing services and KPIs depend on this field?
- Which production context is affected, and is staging different?
- Who owns the source, flow, model and business definition?
- Which classification or policy applies, and what evidence supports it?
- What changed between the pipeline version that succeeded and the one that failed?
These relationships let teams investigate a VegaFlow run or VegaDB dataset alongside its owners, dependencies and business impact.
Shared operating model
Organizations contain workspaces. Users and service principals receive scoped permissions. Secrets stay in managed secret objects. Data, connections, warehouses, compute environments, flows and graph entities remain workspace-scoped, with history and provenance preserved across product boundaries.