VegaFlow troubleshooting
Diagnose connection, discovery, preparation, execution, scheduling and deployment failures.
Start with the highest lifecycle layer that is wrong, then move inward. A failed pipeline does not automatically mean its connector or compute environment is unhealthy.
| Symptom | Inspect first | Common causes |
|---|---|---|
| Connection validation fails | Validation field errors and network context | DNS, TLS trust, expired secret, firewall, source privilege |
| Discovery is empty | Allowed namespaces and source principal | Namespace filter, missing schema/table access, wrong database |
| Environment will not start | Environment status and capacity event | Node profile unavailable, quota, region/network policy |
| Run stays queued | Queue reason and environment readiness | Active-run policy, stopped compute, capacity, permission change |
| Run fails during preparation | Execution timeline | Version not published, deployment stale, revoked dependency, missing secret |
| One QuickFlow stream fails | Stream attempt and current source schema | Cursor/key changed, incompatible type, destination constraint |
| Pipeline task fails | Run history and latest attempt logs | Code/config error, resource limit, dependency output, external service |
| Schedule did not run | Schedule state and next run time | Disabled/retired, time zone, no active deployment, deduplication |
| Event was ignored | Trigger event record | Duplicate source ID, no matching enabled trigger, invalid source |
| Deployment awaits approval | Deployment and source approval records | Environment approval policy, changed candidate generation |
Connection diagnosis
- Read the connector property schema again; the enabled connector version may have changed.
- Validate from the same compute/network context used by the flow.
- Confirm the secret is active without attempting to read its value.
- Test the narrow source/destination privileges required by the selected mode.
- Repeat catalog discovery and compare stream identity, field types, cursor and primary key.
Execution diagnosis
- Open the execution timeline and identify the first failing transition.
- Confirm pipeline version, source revision, deployment generation and environment.
- Inspect child run history before reading raw logs.
- Check whether logs were truncated or have expired.
- Decide whether retry is safe or a new version is required.
Schema changes
Do not automatically broaden schema-evolution policy to clear a failure. Determine whether the source change is additive, destructive, type-widening or identity-changing. Update mappings and publish a new version when downstream meaning changes.
Partial external effects
A canceled or failed task can have committed data or called an external API before stopping. Reconcile the destination with the execution/run identity. Retry only when the connector operation and your task design make the repeated effect safe.
Support bundle
When escalating, include organization/workspace, resource IDs, immutable version/revision IDs, execution and run IDs, event or schedule identity, timestamps, structured error code and relevant timeline entries. Remove access tokens, credentials and sensitive payloads. Do not paste an unrestricted connection document or full production log archive into a ticket.