cb9aefd49b
## Summary Self-hosted deployments now run ClickHouse from the official [`clickhouse/clickhouse-server`](https://hub.docker.com/r/clickhouse/clickhouse-server) image instead of `bitnamilegacy/clickhouse`. Bitnami's free image catalog is EOL and the frozen legacy archive tops out at ClickHouse 25.7.5, below the 25.8 minimum the platform requires since v4.5.0, which broke every ClickHouse insert on chart-bundled deployments. Both stacks now default to 26.2, the same version the platform is developed and tested against. Existing deployments keep their ClickHouse data with no manual migration. Fixes #4197. ## Details **Docker Compose**: the `clickhouse` service uses the official image with its native env vars, plus the recommended `nofile` ulimits. It reuses the same named volume as before: a `data-paths.xml` config override points ClickHouse at the `data/` subdirectory of the volume, which is exactly the layout the Bitnami image used, so old volumes work in place (including SQL-created users) and fresh installs get the identical layout. The service follows the required-secrets model: `CLICKHOUSE_PASSWORD` must be set, matching the other services. **Helm chart**: the Bitnami ClickHouse subchart is replaced by a chart-owned single-node StatefulSet and Service running the official image (non-root, HTTP `/ping` probes, config overrides mounted into `config.d`, and the same `data-paths.xml` layout compatibility). On upgrade, the chart automatically adopts the data PVC left behind by the old subchart (`data-<release>-clickhouse-shard0-0`) via `lookup`, and `fsGroup` relabeling handles the uid change on first mount. Both the ClickHouse server and the webapp read the password from the same chart-managed datastore secret (auto-generated and retained across upgrades), so the server credential and the app's connection URL always match. Existing `clickhouse.*` values keep working: `auth` (including `existingSecret`/`existingSecretKey`), `persistence` (including `global.storageClass`), `resources`, `secure`, `external.*`, `configdFiles`, and now `nodeSelector`/`tolerations`/`affinity`. Bitnami-only keys (`shards`, `replicaCount`, `keeper`, `resourcesPreset`) are gone; default `resources` requests/limits match what the old preset applied. The docs state the 25.8 minimum for bring-your-own ClickHouse. ## Upgrade caveats An adversarial review of the upgrade path found a few cohorts that need awareness (all documented): - **GitOps tools that render with `helm template`** (no cluster access): PVC auto-detection can't run, so `clickhouse.persistence.existingClaim` must be set to the old PVC name or ClickHouse starts on a fresh empty volume. Documented in the values file and the Kubernetes self-hosting docs. Tools that run real helm installs (e.g. Flux) adopt automatically. - **A pinned `CLICKHOUSE_IMAGE_TAG`** pointing at a Bitnami tag must be updated to an official image tag; documented in the Docker self-hosting docs. - **Storage without `fsGroup` support** (NFS, hostPath): set `clickhouse.volumePermissions.enabled: true` for a one-time ownership-fixing init container. - **Rollback is not automatic**: once the official image has run, file ownership changes and the Bitnami image can no longer read the volume without a manual chown, and ClickHouse does not support downgrades across the version gap. ## Verification - Full upgrade simulation for Compose, twice (before and after rebasing onto the required-secrets release): booted the ClickHouse service from the old compose file on `main` (Bitnami), wrote thousands of rows, then brought the same project up with this branch's compose file. The official 26.2 server came up healthy on the same volume with all rows intact, SQL-created users working, and writes succeeding. - Adoption scenarios tested against real containers: old volume + root entrypoint (Compose), old volume owned by the Bitnami uid + non-root 101 with fsGroup-style group permissions (Kubernetes), and fresh volumes for both. - `helm lint`, `helm template` (default values, `existingClaim` set, external ClickHouse, volumePermissions/scheduling toggles, and the production example) and kubeconform all pass, mirroring the release CI steps. The rendered webapp Deployment and ClickHouse StatefulSet resolve to the same datastore secret key. - Inserts using `input_format_json_infer_array_of_dynamic_from_array_of_different_types` (the setting that fails on 25.7.5) succeed on the upgraded volume. ## Upgrade preflight and docs A production upgrade report on this branch surfaced two hazards that predate this PR — both landed in chart 4.5.6 (#4316) — so they are fixed here rather than left for the next person to hit. **`secrets.existingSecret` gained two required keys.** The webapp started reading `PROVIDER_SECRET` and `COORDINATOR_SECRET`, and when `existingSecret` is set the chart generates nothing, so a missing key only surfaced as a `CreateContainerConfigError` partway through the webapp rollout. The pre-install/pre-upgrade validation now looks the Secret up and fails with the complete list of missing keys, leaving the running release untouched. It is skipped under `helm template` and client-side dry-run, where `lookup` cannot read the cluster. **Bundled datastore credentials moved into the chart-managed Secret** (`<release>-clickhouse`/`admin-password` → `trigger-datastore`/`clickhouse-admin-password`). The chart wires both ends itself, but consumers outside it — maintenance CronJobs, Grafana datasources, secret syncs — have to be repointed. A new `## Upgrading` section in the Kubernetes docs carries the old→new mapping, the two new keys, and a pointer to the ClickHouse image notes. The existingSecret key list in the docs also named `OBJECT_STORE_ACCESS_KEY_ID`/`OBJECT_STORE_SECRET_ACCESS_KEY`, which are env var names rather than keys the chart reads; corrected to the real key names and the condition under which they apply. Verified on a throwaway kind cluster with `--dry-run=server`: a pre-4.5.6 Secret fails with both key names listed, the documented `kubectl patch` clears it, and default values, `existingClaim`, external ClickHouse, volumePermissions/scheduling and the production example all still render. A real `helm install` followed by an upgrade against an incomplete Secret aborts with the release still at revision 1 and `deployed`. `helm lint`, the CI render and kubeconform (59 resources, 0 invalid) pass. --------- Co-authored-by: nicktrn <55853254+nicktrn@users.noreply.github.com>
312 lines
11 KiB
YAML
312 lines
11 KiB
YAML
name: trigger
|
|
|
|
x-logging: &logging-config
|
|
driver: ${LOGGING_DRIVER:-local}
|
|
options:
|
|
max-size: ${LOGGING_MAX_SIZE:-20m}
|
|
max-file: ${LOGGING_MAX_FILES:-5}
|
|
compress: ${LOGGING_COMPRESS:-true}
|
|
|
|
services:
|
|
webapp:
|
|
image: ghcr.io/triggerdotdev/trigger.dev:${TRIGGER_IMAGE_TAG:-latest}
|
|
restart: ${RESTART_POLICY:-unless-stopped}
|
|
logging: *logging-config
|
|
ports:
|
|
- ${WEBAPP_PUBLISH_IP:-0.0.0.0}:8030:3000
|
|
depends_on:
|
|
- postgres
|
|
- redis
|
|
- clickhouse
|
|
- s2
|
|
networks:
|
|
- webapp
|
|
- supervisor
|
|
volumes:
|
|
- shared:/home/node/shared
|
|
# Only needed for bootstrap
|
|
user: root
|
|
# Only needed for bootstrap
|
|
command: sh -c "chown -R node:node /home/node/shared && exec ./scripts/entrypoint.sh"
|
|
healthcheck:
|
|
test:
|
|
[
|
|
"CMD",
|
|
"node",
|
|
"-e",
|
|
"http.get('http://localhost:3000/healthcheck', res => process.exit(res.statusCode === 200 ? 0 : 1)).on('error', () => process.exit(1))",
|
|
]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 5
|
|
start_period: 10s
|
|
environment:
|
|
APP_ORIGIN: ${APP_ORIGIN:-http://localhost:8030}
|
|
LOGIN_ORIGIN: ${LOGIN_ORIGIN:-http://localhost:8030}
|
|
API_ORIGIN: ${API_ORIGIN:-http://localhost:8030}
|
|
ELECTRIC_ORIGIN: http://electric:3000
|
|
# Realtime streams v2, backed by the bundled s2-lite service. This powers
|
|
# AI-agent token streaming and run streams. Point the endpoint at a hosted
|
|
# S2 (https://s2.dev) instead by overriding these and setting an access token.
|
|
# To fall back to the Redis-backed v1 streams, set the version to `v1`.
|
|
REALTIME_STREAMS_DEFAULT_VERSION: ${REALTIME_STREAMS_DEFAULT_VERSION:-v2}
|
|
REALTIME_STREAMS_S2_BASIN: ${REALTIME_STREAMS_S2_BASIN:-trigger-realtime}
|
|
REALTIME_STREAMS_S2_ENDPOINT: ${REALTIME_STREAMS_S2_ENDPOINT:-http://s2/v1}
|
|
REALTIME_STREAMS_S2_SKIP_ACCESS_TOKENS: ${REALTIME_STREAMS_S2_SKIP_ACCESS_TOKENS:-true}
|
|
REALTIME_STREAMS_S2_ACCESS_TOKEN: ${REALTIME_STREAMS_S2_ACCESS_TOKEN:-}
|
|
DATABASE_URL: ${DATABASE_URL:-postgresql://postgres:${POSTGRES_PASSWORD}@postgres:5432/main?schema=public&sslmode=disable}
|
|
DIRECT_URL: ${DIRECT_URL:-postgresql://postgres:${POSTGRES_PASSWORD}@postgres:5432/main?schema=public&sslmode=disable}
|
|
SESSION_SECRET: ${SESSION_SECRET}
|
|
MAGIC_LINK_SECRET: ${MAGIC_LINK_SECRET}
|
|
ENCRYPTION_KEY: ${ENCRYPTION_KEY}
|
|
PROVIDER_SECRET: ${PROVIDER_SECRET}
|
|
COORDINATOR_SECRET: ${COORDINATOR_SECRET}
|
|
MANAGED_WORKER_SECRET: ${MANAGED_WORKER_SECRET}
|
|
REDIS_HOST: redis
|
|
REDIS_PORT: 6379
|
|
REDIS_TLS_DISABLED: true
|
|
APP_LOG_LEVEL: info
|
|
DEV_OTEL_EXPORTER_OTLP_ENDPOINT: ${DEV_OTEL_EXPORTER_OTLP_ENDPOINT:-http://localhost:8030/otel}
|
|
DEPLOY_REGISTRY_HOST: ${DOCKER_REGISTRY_URL:-localhost:5000}
|
|
DEPLOY_REGISTRY_NAMESPACE: ${DOCKER_REGISTRY_NAMESPACE:-trigger}
|
|
OBJECT_STORE_BASE_URL: ${OBJECT_STORE_BASE_URL:-http://minio:9000}
|
|
OBJECT_STORE_ACCESS_KEY_ID: ${OBJECT_STORE_ACCESS_KEY_ID}
|
|
OBJECT_STORE_SECRET_ACCESS_KEY: ${OBJECT_STORE_SECRET_ACCESS_KEY}
|
|
GRACEFUL_SHUTDOWN_TIMEOUT: 1000
|
|
NODE_MAX_OLD_SPACE_SIZE: ${NODE_MAX_OLD_SPACE_SIZE}
|
|
# Bootstrap - this will automatically set up a worker group for you
|
|
# This will NOT work for split deployments
|
|
TRIGGER_BOOTSTRAP_ENABLED: 1
|
|
TRIGGER_BOOTSTRAP_WORKER_GROUP_NAME: bootstrap
|
|
TRIGGER_BOOTSTRAP_WORKER_TOKEN_PATH: /home/node/shared/worker_token
|
|
# ClickHouse configuration
|
|
CLICKHOUSE_URL: ${CLICKHOUSE_URL:-http://${CLICKHOUSE_USER:-default}:${CLICKHOUSE_PASSWORD}@clickhouse:8123?secure=false}
|
|
CLICKHOUSE_LOG_LEVEL: ${CLICKHOUSE_LOG_LEVEL:-info}
|
|
# Run replication
|
|
RUN_REPLICATION_ENABLED: ${RUN_REPLICATION_ENABLED:-1}
|
|
RUN_REPLICATION_CLICKHOUSE_URL: ${RUN_REPLICATION_CLICKHOUSE_URL:-http://${CLICKHOUSE_USER:-default}:${CLICKHOUSE_PASSWORD}@clickhouse:8123}
|
|
RUN_REPLICATION_LOG_LEVEL: ${RUN_REPLICATION_LOG_LEVEL:-info}
|
|
# Limits
|
|
# TASK_PAYLOAD_OFFLOAD_THRESHOLD: 524288 # 512KB
|
|
# TASK_PAYLOAD_MAXIMUM_SIZE: 3145728 # 3MB
|
|
# BATCH_TASK_PAYLOAD_MAXIMUM_SIZE: 1000000 # 1MB
|
|
# TASK_RUN_METADATA_MAXIMUM_SIZE: 262144 # 256KB
|
|
# DEFAULT_ENV_EXECUTION_CONCURRENCY_LIMIT: 100
|
|
# DEFAULT_ORG_EXECUTION_CONCURRENCY_LIMIT: 100
|
|
# Internal OTEL configuration
|
|
INTERNAL_OTEL_TRACE_LOGGING_ENABLED: ${INTERNAL_OTEL_TRACE_LOGGING_ENABLED:-0}
|
|
|
|
postgres:
|
|
image: postgres:${POSTGRES_IMAGE_TAG:-14}
|
|
restart: ${RESTART_POLICY:-unless-stopped}
|
|
logging: *logging-config
|
|
ports:
|
|
- ${POSTGRES_PUBLISH_IP:-127.0.0.1}:5433:5432
|
|
volumes:
|
|
- postgres:/var/lib/postgresql/data/
|
|
networks:
|
|
- webapp
|
|
command:
|
|
- -c
|
|
- wal_level=logical
|
|
environment:
|
|
POSTGRES_USER: ${POSTGRES_USER:-postgres}
|
|
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?Set POSTGRES_PASSWORD in .env - run ./generate-secrets.sh}
|
|
POSTGRES_DB: ${POSTGRES_DB:-postgres}
|
|
healthcheck:
|
|
test: ["CMD", "pg_isready", "-U", "postgres"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 10s
|
|
|
|
redis:
|
|
image: redis:${REDIS_IMAGE_TAG:-7}
|
|
restart: ${RESTART_POLICY:-unless-stopped}
|
|
logging: *logging-config
|
|
ports:
|
|
- ${REDIS_PUBLISH_IP:-127.0.0.1}:6389:6379
|
|
volumes:
|
|
- redis:/data
|
|
networks:
|
|
- webapp
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 10s
|
|
|
|
electric:
|
|
image: electricsql/electric:${ELECTRIC_IMAGE_TAG:-1.2.4}
|
|
restart: ${RESTART_POLICY:-unless-stopped}
|
|
logging: *logging-config
|
|
depends_on:
|
|
- postgres
|
|
networks:
|
|
- webapp
|
|
environment:
|
|
DATABASE_URL: ${DATABASE_URL:-postgresql://postgres:${POSTGRES_PASSWORD}@postgres:5432/main?schema=public&sslmode=disable}
|
|
ELECTRIC_INSECURE: true
|
|
ELECTRIC_USAGE_REPORTING: false
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:3000/v1/health"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 10s
|
|
|
|
clickhouse:
|
|
image: clickhouse/clickhouse-server:${CLICKHOUSE_IMAGE_TAG:-26.2}
|
|
restart: ${RESTART_POLICY:-unless-stopped}
|
|
logging: *logging-config
|
|
ports:
|
|
- ${CLICKHOUSE_PUBLISH_IP:-127.0.0.1}:9123:8123
|
|
- ${CLICKHOUSE_PUBLISH_IP:-127.0.0.1}:9090:9000
|
|
ulimits:
|
|
nofile:
|
|
soft: 262144
|
|
hard: 262144
|
|
environment:
|
|
CLICKHOUSE_USER: ${CLICKHOUSE_USER:-default}
|
|
CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD:?Set CLICKHOUSE_PASSWORD in .env - run ./generate-secrets.sh}
|
|
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1
|
|
volumes:
|
|
# The same volume works across upgrades from the previous Bitnami-based
|
|
# setup: data-paths.xml keeps the on-disk layout compatible.
|
|
- clickhouse:/var/lib/clickhouse
|
|
- ../clickhouse/data-paths.xml:/etc/clickhouse-server/config.d/data-paths.xml:ro
|
|
- ../clickhouse/override.xml:/etc/clickhouse-server/config.d/override.xml:ro
|
|
networks:
|
|
- webapp
|
|
healthcheck:
|
|
test:
|
|
[
|
|
"CMD",
|
|
"clickhouse-client",
|
|
"--host",
|
|
"localhost",
|
|
"--port",
|
|
"9000",
|
|
"--user",
|
|
"${CLICKHOUSE_USER:-default}",
|
|
"--password",
|
|
"${CLICKHOUSE_PASSWORD}",
|
|
"--query",
|
|
"SELECT 1",
|
|
]
|
|
interval: 5s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 10s
|
|
|
|
registry:
|
|
image: registry:${REGISTRY_IMAGE_TAG:-2}
|
|
restart: ${RESTART_POLICY:-unless-stopped}
|
|
logging: *logging-config
|
|
ports:
|
|
- ${REGISTRY_PUBLISH_IP:-127.0.0.1}:5000:5000
|
|
networks:
|
|
- webapp
|
|
volumes:
|
|
# registry-user:very-secure-indeed
|
|
- ../registry/auth.htpasswd:/auth/htpasswd:ro
|
|
environment:
|
|
REGISTRY_AUTH: htpasswd
|
|
REGISTRY_AUTH_HTPASSWD_REALM: Registry Realm
|
|
REGISTRY_AUTH_HTPASSWD_PATH: /auth/htpasswd
|
|
healthcheck:
|
|
test: ["CMD", "wget", "--spider", "-q", "http://localhost:5000/"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 10s
|
|
|
|
minio:
|
|
image: bitnamilegacy/minio:${MINIO_IMAGE_TAG:-latest}
|
|
restart: ${RESTART_POLICY:-unless-stopped}
|
|
logging: *logging-config
|
|
ports:
|
|
- ${MINIO_PUBLISH_IP:-127.0.0.1}:9000:9000
|
|
- ${MINIO_PUBLISH_IP:-127.0.0.1}:9001:9001
|
|
networks:
|
|
- webapp
|
|
volumes:
|
|
- minio:/bitnami/minio/data
|
|
environment:
|
|
MINIO_ROOT_USER: ${MINIO_ROOT_USER:-${OBJECT_STORE_ACCESS_KEY_ID:-admin}}
|
|
MINIO_ROOT_PASSWORD: ${MINIO_ROOT_PASSWORD:-${OBJECT_STORE_SECRET_ACCESS_KEY:?Set OBJECT_STORE_SECRET_ACCESS_KEY in .env - run ./generate-secrets.sh}}
|
|
MINIO_DEFAULT_BUCKETS: packets
|
|
MINIO_BROWSER: "on"
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:9000/minio/health/live"]
|
|
interval: 5s
|
|
timeout: 10s
|
|
retries: 5
|
|
start_period: 10s
|
|
|
|
# Generates the s2-lite basin spec from REALTIME_STREAMS_S2_BASIN so the bundled
|
|
# basin always matches the basin the webapp targets. s2-lite is distroless (no
|
|
# shell), so a tiny init container writes the spec into a shared volume first.
|
|
s2-init:
|
|
image: busybox:${BUSYBOX_IMAGE_TAG:-1.37}
|
|
restart: "no"
|
|
logging: *logging-config
|
|
environment:
|
|
REALTIME_STREAMS_S2_BASIN: ${REALTIME_STREAMS_S2_BASIN:-trigger-realtime}
|
|
command:
|
|
- sh
|
|
- -c
|
|
- |
|
|
cat > /config/s2-spec.json <<JSON
|
|
{
|
|
"basins": [
|
|
{
|
|
"name": "$$REALTIME_STREAMS_S2_BASIN",
|
|
"config": { "create_stream_on_append": true, "create_stream_on_read": true }
|
|
}
|
|
]
|
|
}
|
|
JSON
|
|
volumes:
|
|
- s2-config:/config
|
|
networks:
|
|
- webapp
|
|
|
|
# s2-lite: the open-source, self-hostable S2 (https://s2.dev) server that backs
|
|
# realtime streams v2. It's distroless (no shell), so a wget/curl healthcheck
|
|
# always reports unhealthy even when the API is up — hence no healthcheck.
|
|
# Runs as root so it can write to the named volume; storage lives in SlateDB
|
|
# under /data, and the basin is created on startup from the generated spec.
|
|
s2:
|
|
image: ${S2_IMAGE:-ghcr.io/s2-streamstore/s2:latest@sha256:d6ded5ca7dd619fa7c946f06e39a98f9c95c6883c8bb884e5eaa129f232c920c}
|
|
restart: ${RESTART_POLICY:-unless-stopped}
|
|
logging: *logging-config
|
|
depends_on:
|
|
s2-init:
|
|
condition: service_completed_successfully
|
|
command: ["lite", "--init-file", "/config/s2-spec.json", "--local-root", "/data"]
|
|
user: "0:0"
|
|
volumes:
|
|
- s2-config:/config
|
|
- s2:/data
|
|
networks:
|
|
- webapp
|
|
|
|
volumes:
|
|
clickhouse:
|
|
postgres:
|
|
redis:
|
|
shared:
|
|
minio:
|
|
s2:
|
|
s2-config:
|
|
|
|
networks:
|
|
docker-proxy:
|
|
name: docker-proxy
|
|
supervisor:
|
|
name: supervisor
|
|
webapp:
|
|
name: webapp
|