Skip to content

DocsEnterprise

High availability

Several API replicas behind a load balancer, on shared stores, with one queue of background index jobs that survives a replica stopping.

What the replicas share

  • Repositories, the graph and vectors, in your Postgres, Neo4j and Qdrant.
  • Principals, sessions and the audit chain, in Postgres; rate limits and the cache, in Redis.
  • Each replica picks up the repositories another indexed, told at once through Redis.
  • With the ha feature, one queue of background index jobs, in Postgres.

The shared job queue

  • Any replica accepts a job, and the first free worker on any replica claims it; each replica runs BBM_ATLAS_EE_HA_JOB_WORKERS at once (2).
  • Every replica shows its progress, and a cancellation reaches it at its next step.
  • A job whose replica stops - no heartbeat for five minutes - is queued again and picked up by another. One that has started three times fails, rather than looping.
  • Finished jobs are kept for a week.

The queue needs the Postgres store (BBM_ATLAS_STORAGE_BACKEND=all). BBM_ATLAS_EE_HA_SHARED_JOBS=false keeps each replica's jobs to itself.

On Kubernetes

terminal
kubectl create secret generic atlas-license --from-file=license=./license.keyhelm install atlas oci://<registry>/charts/bbm-atlas \  --set enterprise.licenseSecret=atlas-license \  --set ha.enabled=true --set storage.backend=all \  --set credentials.existingSecret=atlas-credentials \  --set stores.neo4jUri=bolt://graph:7687 --set stores.qdrantUrl=http://vectors:6333

The ha profile runs three replicas by default (ha.replicas), with rolling updates that never drop below them, a disruption budget that keeps all but one running through node drains, and pods spread across nodes. The credentials Secret holds api_keys, postgres_dsn, neo4j_password, qdrant_api_key and redis_url; see Production.