Recover: JarvisOS Control Plane
No signal criticalDispatches tasks to worker nodes and owns workspace state.
- Owner
- homelab-operator
- Target RTO
- 20m
- Last heartbeat
- never
- Dependent entities
- 1
Recovery readiness70/100
Procedure
Read fully before acting.- Never run two control planes against one database — stop the primary before promoting the standby.
- Promote on VIN, then point it at the active Postgres endpoint.
- Drain in-flight worker tasks before failing back; tasks are at-least-once and will re-run.
Bring-up order
Dependencies first — starting out of order is a common second incident.- Home1 Node
- Postgres (primary) Database
- VIN / VPS Node
- ISP Uplink Dependency
- Google OAuth Dependency
- Redis (cache / queue) Database
- JarvisOS Auth (SSO) Auth
- JarvisOS Control Plane Service target
Failover targets
mode: manual-
Standby unit on VIN; needs a reachable Postgres before it will accept work.
Restore source
- Destination
- nas-backups
- Schedule
- hourly
- Last success
- never
- RPO
- 1h
- Verified by
- weekly
What this affects
Impacted if this stays down, or while you restart it.
Downstream
JarvisOS Worker (VIN)
Readiness breakdown
- Documented runbook 20 pts
- Backup stale or missing 30 pts
- Restore verification scheduled 20 pts
- Healthy failover target available 30 pts