refactor(iac): cluster split — 1 cluster per plane, breakpilot-* naming
ci / shared (pull_request) Successful in 23s
ci / validate (pull_request) Successful in 4s

Restructures the draft to reflect the 2026-06-30 cluster decision:

Three Orca clusters, each becoming its own Gitea repo at migration time:
- breakpilot-edge   → vm-edge       (Identity + Infra: KC, Gitea, Infisical, PowerDNS, Orca-Proxy)
- breakpilot-control → vm-control   (Portal, tenant-registry, ERPNext, MariaDB, Stalwart)
- breakpilot-app    → vm-app-prod + vm-app-stage (CERTifAI, compliance-*, Mongo, MinIO, Qdrant, LiteLLM)

Key model points encoded:
- Identity (Keycloak) co-tenant with Infra on vm-edge (1 VM core), per
  INFRASTRUCTURE.md §6 — heap pinned so it cannot starve PowerDNS/Infisical
- Stage and prod live in the same breakpilot-app cluster on different
  VMs. Stage authenticates via prod Keycloak under tenant.kind = "stage";
  no duplicated identity, no duplicated control plane (per §5).
- The existing CERTifAI Keycloak will be repurposed for breakpilot-edge
  rather than standing up a fresh one — same realm, same users.
- Multi-VM rollout gated on legal entity being established so we can sign
  SysEleven / Hetzner business contracts. Until then, single-VM ops
  continues via ~/workspace/orca-infra; this repo is design-only.

Mechanical changes:
- manifests/{vm-edge,vm-control,vm-data,stage}/ → clusters/{breakpilot-edge,breakpilot-control,breakpilot-app/services/{prod,stage}}/services/
- vm-data → vm-app-prod, stage → vm-app-stage in node references and headers
- overlays/{stage,prod}/overlay.toml point at the new cluster paths
- scripts/validate.sh now enforces a per-cluster node whitelist
  (breakpilot-edge → vm-edge, breakpilot-control → vm-control,
  breakpilot-app → {vm-app-prod, vm-app-stage}) instead of dir-name equality
- New READMEs at clusters/, clusters/breakpilot-edge/,
  clusters/breakpilot-control/, clusters/breakpilot-app/ documenting
  scope, SLA targets, co-tenant notes, and the future-repo split
- Top README rewritten to lead with the cluster-split decision and the
  legal-entity gate; per-milestone fill-in table re-pathed

Validation:
- make validate → 38 files OK (35 manifests + 3 overlays)
- make plan ENV=stage → 11 resolved manifests in .orca-out/stage/
- make plan ENV=prod  → 24 resolved manifests in .orca-out/prod/
This commit is contained in:
Sharang Parnerkar
2026-06-30 22:14:51 +02:00
parent f1c3fd14b9
commit 8971152da0
44 changed files with 428 additions and 125 deletions
+83
View File
@@ -0,0 +1,83 @@
# breakpilot-app
App plane (formerly "Data plane" in the May 18 doc). Two VMs in one cluster:
- **`vm-app-prod`** — production workloads, customer data, real keys
- **`vm-app-stage`** — staging workloads, demo data, sandbox tenant
Becomes its own Gitea repo `platform/breakpilot-app` at migration time.
## Services
### Prod ([`services/prod/`](./services/prod/), 9 services)
| Service | Purpose |
|---|---|
| `admin-compliance.toml` | Compliance back-office admin |
| `ai-compliance-sdk.toml` | LLM-mediated compliance SDK |
| `backend-compliance.toml` | Compliance scanner backend |
| `certifai-dashboard.toml` | CERTifAI dashboard (first product) |
| `litellm.toml` | LLM gateway |
| `minio.toml` | S3-compatible object store |
| `mongodb.toml` | Compliance + CERTifAI data |
| `pg-app.toml` | Per-tenant application Postgres (SPOF — RISK-1 in `§7`) |
| `qdrant.toml` | Vector store (rebuildable) |
### Stage ([`services/stage/`](./services/stage/), 11 services)
| Service | Purpose |
|---|---|
| `admin-compliance.toml` | Stage admin |
| `ai-compliance-sdk.toml` | Stage SDK |
| `backend-compliance.toml` | Stage backend |
| `certifai-dashboard.toml` | Stage CERTifAI |
| `customer-portal.toml` | Stage portal (calls prod KC + prod tenant-registry with `tenant.kind = "stage"`) |
| `litellm.toml` | Stage LLM gateway (may reuse prod) |
| `mongodb-stage.toml` | Stage Mongo |
| `orca-proxy.toml` | Per-node ingress on `vm-app-stage` |
| `pg-app-stage.toml` | Stage Postgres |
| `qdrant-stage.toml` | Stage vector store |
| `tenant-registry.toml` | (placeholder; stage calls prod registry — manifest exists for parity) |
## SLA targets (per `INFRASTRUCTURE.md §6`)
```
Owns: DATA_RESIDENCY — all customer data (MongoDB, pg-app, MinIO) must stay EU
RPO (product) — compliance records ≤ 6h; chat history ≤ 24h
DATA_ISOLATION — every query scoped by org_id / tenant_id
AUDIT_TRAIL — product-level actions
AVAILABILITY — CERTifAI ≥ 99.5%; compliance ≥ 99.5%
```
## Why one cluster, two VMs
- **Same orca-platform config, different physical workloads.** Stage and
prod don't drift on infra config because they're in the same cluster.
- **No "oops touched prod" accidents.** Per-service `placement.node`
pins which VM each container lands on; prod manifests live in `prod/`,
stage in `stage/`.
- **Blast radius is physical.** A stage misconfiguration cannot exhaust
`vm-app-prod`'s memory because they share nothing but the Orca control
socket.
## Stage isolation contract
Stage services never:
- email real customers (Stalwart accept-rule on `breakpilot-control` drops
recipients not matching `*+stage@*`)
- trigger real Polar charges (`POLAR_API_URL` points at sandbox)
- carry real customer data (sandbox tenant resets nightly)
Stage services always:
- authenticate via prod Keycloak with `tenant.kind = "stage"`
- read tenant config from prod `tenant-registry` (read-only for stage tenants)
- expose `*.stage.breakpilot.com` (or whatever the staging subdomain is)
See `INFRASTRUCTURE.md §5` for the full stage-prod sharing contract.
## Sandbox tenant
The `sandbox` tenant in `stage/` is the only place that accepts writes
from public demo visitors. A nightly cron (lands as `sandbox-reset.toml`
when seed-data fixtures are wired in M13.1) resets its Mongo collections
and Keycloak user attributes from a versioned seed.
@@ -0,0 +1,15 @@
# admin-compliance stub — full config lands in M7.x.
# Host: vm-app-prod. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "admin-compliance"
image = "registry.breakpilot.com/admin-compliance:placeholder"
port = 3002
depends_on = ["backend-compliance", "ai-compliance-sdk"]
[service.placement]
node = "vm-app-prod"
[service.resources]
memory = "512Mi"
cpu = 0.25
@@ -0,0 +1,15 @@
# ai-compliance-sdk stub — full config lands in M7.x.
# Host: vm-app-prod. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "ai-compliance-sdk"
image = "registry.breakpilot.com/ai-compliance-sdk:placeholder"
port = 3001
depends_on = ["pg-app", "qdrant", "litellm"]
[service.placement]
node = "vm-app-prod"
[service.resources]
memory = "1Gi"
cpu = 0.5
@@ -0,0 +1,15 @@
# backend-compliance stub — full config lands in M7.x.
# Host: vm-app-prod. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "backend-compliance"
image = "registry.breakpilot.com/backend-compliance:placeholder"
port = 3000
depends_on = ["pg-app", "minio"]
[service.placement]
node = "vm-app-prod"
[service.resources]
memory = "1Gi"
cpu = 0.5
@@ -0,0 +1,15 @@
# certifai-dashboard stub — full config lands in M6.x.
# Host: vm-app-prod. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "certifai-dashboard"
image = "registry.breakpilot.com/certifai:placeholder"
port = 3000
depends_on = ["mongodb", "litellm"]
[service.placement]
node = "vm-app-prod"
[service.resources]
memory = "1Gi"
cpu = 0.5
@@ -0,0 +1,18 @@
# litellm stub — full config lands in M6.x.
# Host: vm-app-prod. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "litellm"
image = "ghcr.io/berriai/litellm:main-stable"
port = 4000
[service.placement]
node = "vm-app-prod"
[service.resources]
memory = "1Gi"
cpu = 0.5
[service.env]
LITELLM_MASTER_KEY = "${secrets.LITELLM_MASTER_KEY}"
LITELLM_SALT_KEY = "${secrets.LITELLM_SALT_KEY}"
@@ -0,0 +1,23 @@
# minio stub — full config lands in M7.x.
# Host: vm-app-prod. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "minio"
image = "minio/minio:latest"
port = 9000
extra_ports = ["9001:9001"]
cmd = ["server", "/data", "--console-address", ":9001"]
[service.placement]
node = "vm-app-prod"
[service.resources]
memory = "1Gi"
cpu = 0.5
[service.volume]
path = "/data"
[service.env]
MINIO_ROOT_USER = "${secrets.MINIO_ROOT_USER}"
MINIO_ROOT_PASSWORD = "${secrets.MINIO_ROOT_PASSWORD}"
@@ -0,0 +1,21 @@
# mongodb stub — full config lands in M6.x.
# Host: vm-app-prod. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "mongodb"
image = "mongo:7"
port = 27017
[service.placement]
node = "vm-app-prod"
[service.resources]
memory = "2Gi"
cpu = 1.0
[service.volume]
path = "/data/db"
[service.env]
MONGO_INITDB_ROOT_USERNAME = "${secrets.MONGO_ADMIN_USER}"
MONGO_INITDB_ROOT_PASSWORD = "${secrets.MONGO_ADMIN_PASSWORD}"
@@ -0,0 +1,23 @@
# pg-app stub — full config lands in M4.1.
# Host: vm-app-prod. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
# RISK-1 (§12): single instance owns tenant_registry + compliance schemas. Split into pg-registry + pg-compliance at Tier B.
[[service]]
name = "pg-app"
image = "postgres:16-alpine"
port = 5432
[service.placement]
node = "vm-app-prod"
[service.resources]
memory = "3Gi"
cpu = 1.0
[service.volume]
path = "/var/lib/postgresql/data"
[service.env]
POSTGRES_DB = "platform"
POSTGRES_USER = "platform"
POSTGRES_PASSWORD = "${secrets.PG_APP_PASSWORD}"
@@ -0,0 +1,17 @@
# qdrant stub — full config lands in M7.x.
# Host: vm-app-prod. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "qdrant"
image = "qdrant/qdrant:v1.10.0"
port = 6333
[service.placement]
node = "vm-app-prod"
[service.resources]
memory = "1Gi"
cpu = 0.5
[service.volume]
path = "/qdrant/storage"
@@ -0,0 +1,14 @@
# admin-compliance stub — full config lands in M7.x.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "admin-compliance"
image = "registry.breakpilot.com/admin-compliance:env-stage"
port = 3002
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "256Mi"
cpu = 0.25
@@ -0,0 +1,14 @@
# ai-compliance-sdk stub — full config lands in M7.x.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "ai-compliance-sdk"
image = "registry.breakpilot.com/ai-compliance-sdk:env-stage"
port = 3001
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "512Mi"
cpu = 0.25
@@ -0,0 +1,14 @@
# backend-compliance stub — full config lands in M7.x.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "backend-compliance"
image = "registry.breakpilot.com/backend-compliance:env-stage"
port = 3000
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "512Mi"
cpu = 0.25
@@ -0,0 +1,14 @@
# certifai-dashboard stub — full config lands in M6.x.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "certifai-dashboard"
image = "registry.breakpilot.com/certifai:env-stage"
port = 3000
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "1Gi"
cpu = 0.5
@@ -0,0 +1,15 @@
# customer-portal stub — full config lands in M5.1.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "customer-portal"
image = "registry.breakpilot.com/portal:env-stage"
port = 3000
domain = "*.stage.breakpilot.com"
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "1Gi"
cpu = 0.5
@@ -0,0 +1,14 @@
# litellm stub — full config lands in M6.x.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
[[service]]
name = "litellm"
image = "ghcr.io/berriai/litellm:main-stable"
port = 4000
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "512Mi"
cpu = 0.25
@@ -0,0 +1,15 @@
# mongodb-stage stub — full config lands in M6.x.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
# Ephemeral.
[[service]]
name = "mongodb-stage"
image = "mongo:7"
port = 27017
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "512Mi"
cpu = 0.25
@@ -0,0 +1,15 @@
# orca-proxy stub — full config lands in M1.2.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
# Stage proxy only routes to stage app containers.
[[service]]
name = "orca-proxy"
image = "orca-managed/orca-proxy:placeholder"
port = 443
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "256Mi"
cpu = 0.5
@@ -0,0 +1,15 @@
# pg-app-stage stub — full config lands in M4.1.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
# Ephemeral; no backup, no volume; reset on each release.
[[service]]
name = "pg-app-stage"
image = "postgres:16-alpine"
port = 5432
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "1Gi"
cpu = 0.5
@@ -0,0 +1,15 @@
# qdrant-stage stub — full config lands in M7.x.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
# Ephemeral, tiny corpus.
[[service]]
name = "qdrant-stage"
image = "qdrant/qdrant:v1.10.0"
port = 6333
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "512Mi"
cpu = 0.25
@@ -0,0 +1,19 @@
# tenant-registry stub — full config lands in M4.1.
# Host: vm-app-stage. Resource budget per INFRASTRUCTURE.md §6 co-tenant notes.
# Calls PROD Keycloak per §2 "Calls OUT to prod"; audience is stage_client_id.
[[service]]
name = "tenant-registry"
image = "registry.breakpilot.com/tenant-registry:env-stage"
port = 8090
depends_on = ["pg-app-stage"]
[service.placement]
node = "vm-app-stage"
[service.resources]
memory = "512Mi"
cpu = 0.25
[service.env]
KEYCLOAK_ISSUER = "https://auth.breakpilot.com/realms/breakpilot-prod"