19 KiB
Deployment — Coolify
This document covers deploying the Polaris Task Force to Coolify (self-hosted PaaS). The app is built as a single Docker image (multi-stage, Bun + Node) and deployed once per environment, each backed by its own Coolify-managed PostgreSQL service.
Local full-stack dev with the same image shape is described at the end (docker compose up).
1. Architecture at a glance
Coolify cluster
┌──────────────────────────────────────────────────────────┐
│ │
dev env │ ┌────────────┐ ┌─────────────┐ │
(polaris-dev) │ │ App svc │ ← │ Postgres │ (Coolify-managed, │
│ │ (Docker- │ │ svc │ per-env) │
│ │ file) │ │ + volume │ │
│ │ + media │ └─────────────┘ │
│ │ volume │ │
│ └─────┬──────┘ │
│ │ scheduled jobs (docker exec → bun run payload) │
└────────┼──────────────────────────────────────────────────┘
│
\________ builds pushed on git push (webhook)
- One image, parameterized by env vars per environment. No image bake per env required.
- PostgreSQL is provisioned by Coolify as a separate service in each environment (
dev,stg,prd). The app talks to it viaDATABASE_URIinjected by Coolify. - Payload uploads (
media/collection) are written to a Coolify Persistent Storage Volume mounted at/app/mediain the app container, so uploads survive container restarts, rollbacks, and rebuilds. - Scheduled one-shot jobs (game-tick, market-tick) run via
docker execCoolify Scheduled Tasks against the running app container — same image, same env, no second service required. - The Discord bot ships inside the image (
src/bot/) and activates per-environment whenDISCORD_TOKENis set — prod only. It shares the same Postgres DB via the container env and inherits the single-instance constraint: 2 replicas = two bot instances (duplicate command registration, double embeds/DMs). Notifications generated while the container is down are not backfilled (the Discord notification bridge polls with an in-memory cursor).
⚠️ Single-instance only. The realtime SSE pipeline (
gameTick → /api/game-tick/notify → in-process bus → SSE /api/realtime → client router.refresh()) uses an in-memory subscriber bus (src/lib/realtime/bus.ts) that is not shared across processes. Coolify must run exactly 1 replica of the app container per environment for SSE to work. If you scale the app horizontally, SSE stops working (live shipment/event log updates silently fail). Everything else continues to work; you'll need a distributed bus before you can scale. Seesrc/lib/realtime/bus.tsandAGENTS.mdfor context.
2. Prerequisites in Coolify
- A working Coolify instance (you said yours is at
https://coolify.onyxsimple.com). - A destination server registered with Coolify (Docker engine reachable).
- This repo reachable by Coolify (Git source — GitHub/GitLab/etc.) so it can build on push. A manual image-build flow is also fine if you prefer.
- A Postgres-compatible SMTP relay if you want password-reset email (the app already supports Brevo; otherwise leave the
EMAIL_*env vars blank).
3. Provisioning a per-environment PostgreSQL service
For each environment (dev, stg, prd):
- In Coolify, create (or open) the Project that represents the environment (e.g.
Polaris → dev). - New Resource → Database → PostgreSQL.
- Pick a sensible DB name, user, password. Coolify creates the database, exposes the connection string, and injects
POSTGRES_*env vars into the same project automatically. - Note the Connection String Coolify generates (e.g.
postgres://ptf_app:<password>@<coolify-internal-host>:5432/ptf-app-dev). You'll wire this into the app service next.
Backups, version upgrades, and per-env resource limits are all handled by Coolify on this service — that's the whole point of hosting the DB individually.
4. Creating the app service (per environment)
- Inside the same Coolify Project, New Resource → Service / Application → From a Dockerfile (or "From a Git repository" and point at the repo root — Coolify will detect
Dockerfile). - Configure the build:
- Port:
3000(alreadyEXPOSE'd). - Health Check Path:
/api/health— Coolify will hit this to decide "healthy". The DockerHEALTHCHECKdirective does the same thing inside the container. - Build Pack: Dockerfile.
- Don't use the
docker-compose.ymlfor the production app service — that file is for local dev (it bundles its own Postgres). Coolify manages the Postgres service for you.
- Port:
- Set the environment variables (Section 5).
- Attach a Persistent Storage entry mapping
/app/media(read/write). This is where Payload stores uploaded media. The directory is created and chowned inside the image; just mount the volume on top. - Deploy. First build will take a few minutes (Bun install + Next.js standalone build,
--max-old-space-size=8000— make sure the build host has ≥ 4 GB free RAM, ideally 8).
Build args you may want to override
The Dockerfile pins bun@1.3.11 and node:22-alpine. If you want to bump either without editing the Dockerfile, Coolify doesn't pass build args by default — just edit the Dockerfile lines (RUN npm install -g bun@<version> and FROM node:<tag>-alpine).
5. Environment variables (per environment)
| Var | Required | Example / Notes |
|---|---|---|
DATABASE_URI |
✅ | Coolify auto-injects this from the per-env Postgres service if linked. Format: postgres://<user>:<pwd>@<host>:5432/<db> |
PAYLOAD_SECRET |
✅ | Random 32+ chars. openssl rand -hex 16. Unique per environment. |
APP_URL |
✅ | Public base URL for this environment, e.g. https://dev.ptf.example.com |
GAME_TICK_NOTIFY_SECRET |
✅ | Shared secret for /api/game-tick/notify. Must match the value used by your scheduled job (Sec. 6). |
EMAIL_FROM_ADDRESS |
— | Outbound email "from" address (e.g. arma@onyxsimple.com) |
EMAIL_FROM_NAME |
— | Display name (e.g. Arma) |
EMAIL_HOST |
— | SMTP host (e.g. smtp-relay.brevo.com) |
EMAIL_PORT |
— | 587 |
EMAIL_USERNAME |
— | SMTP user |
EMAIL_PASSWORD |
— | SMTP password / API key |
DISCORD_TOKEN |
iff bot enabled in this env | Set in prod only — presence activates the bot (activation gate below). |
DISCORD_GUILD_ID |
iff DISCORD_TOKEN set |
Config throws at import if the token is set and this is missing. |
DISCORD_OPS_CHANNEL_ID |
— | Ops channel for attendance RSVP embeds. |
DISCORD_ANNOUNCE_CHANNEL_ID |
— | Channel for /announce. |
DISCORD_STAFF_ROLE_IDS |
— | Comma-separated staff role IDs (staff-only /announce). |
DISCORD_ATTENDANCE_POLL_MS |
— | Attendance reconcile poll interval; default 60000. |
DISCORD_NOTIFICATION_POLL_MS |
— | Notification bridge poll interval; default 20000. |
Never copy the dev
PAYLOAD_SECRETinto prod. Never copyGAME_TICK_NOTIFY_SECRET. Generate fresh per environment.
Discord bot activation gate
The bot ships inside the same app image and starts only when DISCORD_TOKEN is present:
- Presence gate — the container entrypoint checks for
DISCORD_TOKEN; if it's unset the bot never starts and the web app is unaffected. - Fail-fast on missing guild id — if
DISCORD_TOKENis set butDISCORD_GUILD_IDis missing,src/bot/config.tsthrows at import time and the container fails fast. Set both together. - Prod preflight checklist — configure
DISCORD_OPS_CHANNEL_IDandDISCORD_STAFF_ROLE_IDSbefore adding the token: on startup the bot immediately posts attendance embeds and starts DMing users per their notification preferences. - Same-token-across-envs rationale — the token is ONE Discord application credential, not a per-env secret. The value is identical across environments only because the presence gate restricts activation to prod; this does not conflict with the "never copy secrets between envs" guidance above, which governs per-env secrets like
PAYLOAD_SECRET/GAME_TICK_NOTIFY_SECRET. - Session-flapping warning — never run the bot with the prod token from two places at once (e.g. local dev + prod): Discord force-disconnects one of the sessions.
- Env-clone warning — a Coolify "clone environment" copy would carry the token along. Remove it from non-prod copies, or dev/stage would silently activate the bot.
- Recovery — the entrypoint does not supervise/restart the bot. If the bot crashes, the web app stays healthy (announced subordinate-uptime default); recovery = Coolify Restart or redeploy. Troubleshooting: "token set but no bot login" → check the container logs for the bot's fail-fast line and verify the token/guild are valid.
⚠️ The bot shares the single-instance constraint (Sec. 1): running 2 replicas = two bot instances — duplicate command registration, double embeds/DMs. Notifications generated while the container is down are not backfilled (in-memory cursor).
6. Scheduled jobs (game-tick, market-tick)
The game-tick and market-tick Payload bins are registered on payload.config.ts and expected to run on a regular cadence (cron). In Coolify:
Option A — Coolify Scheduled Tasks (recommended)
Coolify has a Scheduled Tasks feature per service (<service> → Scheduled Tasks → Add). Each scheduled task runs a command inside the running service container — equivalent to docker exec. Add:
| Job | Schedule (cron) | Command |
|---|---|---|
| Process shipments | */5 * * * * (every 5 min, tune to match gameTickIntervalMinutes in your Game Rules global) |
bun run payload game-tick |
| Refresh market | */10 * * * * (every 10 min — adjust as needed) |
bun run payload market-tick |
The commands run inside the app container with the app's env (incl. DATABASE_URI, APP_URL, GAME_TICK_NOTIFY_SECRET), so the bins talk to the same DB and post the notify callback correctly.
Option B — external cron calling docker exec over SSH
If you prefer existing infra:
*/5 * * * * ssh -i /path/to/key root@coolify-host \
docker exec polaris-task-force-app-<env> bun run payload game-tick
*/10 * * * * ssh -i /path/to/key root@coolify-host \
docker exec polaris-task-force-app-<env> bun run payload market-tick
The container name comes from Coolify's per-service container name (visible on the service dashboard).
Confirming a tick ran
Each tick:
- Updates
Shipmentsrows in DB (status, fuel, arrival). POSTs to/api/game-tick/notify(internal SSE bus push → clientsrouter.refresh()). If no connected clients, this just 204s.- Emits
GameEventLogentries you'll see in the Payload admin → Game Event Logs.
7. Running migrations after a schema change
If you add a new collection / field and create a new Payload migration under src/migrations/, run it once after deploying the new image:
docker exec <container-name> bun run payload migrate
Coolify Scheduled Tasks support ad-hoc runs of the same command — fire it manually, watch the logs, then commit. Never run drizzle-kit (bun run db migrate) in prod; bun run db push fails on this DB because of the PostGIS spatial_ref_sys ownership issue (see AGENTS.md Gotchas). Use Payload's migration runner exclusively.
For dev (where push: true in payload.config.ts), starting the dev server auto-applies schema changes — that's intentional and fine for dev databases.
8. Persistent media volume — important for rollbacks
Payload's Media collection (src/collections/Media.ts) uses upload: true with no storage adapter, so files land on local disk under <projectRoot>/media (i.e. /app/media inside the container, since WORKDIR /app and process.cwd() = /app).
In Coolify:
- On the app service → Persistent Storage → Add Path.
- Host path or named volume → mapped to
/app/mediain the container. - Coolify preserves this volume across deploys, including rollbacks.
Without this, every redeploy wipes user uploads.
Note on the local-disk default: this works fine for a self-hosted single-instance deployment, but it does not scale beyond one replica (same caveat as the SSE bus). If you outgrow the single-instance architecture (more replicas, multi-host), switch to an S3-compatible storage adapter — add it to
payload.config.tsunderplugins(the project already has a// storage-adapter-placeholderhook for this) and drop the/app/mediavolume.
9. Health checks
- Container
HEALTHCHECK(inside the Dockerfile): every 30 s,wget http://127.0.0.1:3000/api/health. 3 consecutive failures → unhealthy. - Coolify itself should also probe
GET /api/health(configure the service Health Check Path to/api/health). Coolify will route traffic only to "healthy" containers, so it doubles as your readiness gate after deployment. - The health endpoint is deliberately DB-free — a transient database outage never marks the container unhealthy, so Coolify won't force a useless restart cycle while the Postgres service is bumping.
Response shape:
{ "ok": true, "uptime": 1243, "version": "0.1.6", "commit": "a1b2c3d", "buildTime": "2026-08-12T20:30:00.000Z" }
10. Deploy flow (git push → live)
Recommended flow:
- Bump version locally:
bun pm version patch(or your own flow). Editpackage.json#versionand push the commit. - Push to your environment's branch (e.g.
main→ prod,staging→ stg,develop→ dev), or whatever branch Coolify watches. - Coolify's webhook fires, builds the image, and rolls the container.
- Coolify waits for
/api/healthto come back 200, then routes traffic. - If a new migration shipped with that deploy, run it via
docker execor a Scheduled Task (Sec. 7).
No SSH, no rsync, no systemctl. The legacy build/deploy.sh (gitignored, the old SSH+systemd flow to jmf-usrv-2404) is no longer used — exclude it from future deploys. bun run deploy (the npm script) will still try to run it (it calls postversion → bash ./build/deploy.sh); don't run it from your local machine anymore, or remove that script entry from package.json once you've moved onto Coolify for good.
11. Rollbacks
Coolify supports two rollout modes:
- Automatic rollback on failed health: if the new container never becomes healthy (3 failed
/api/healthprobes within thestart_period), Coolify keeps the old container running and doesn't route traffic to the new one. Brings you back to the previous working version with no action. - Manual rollback: in the service dashboard,/files history → pick a previous image tag → "Rollback". This redeploys that exact image with the existing
/app/mediavolume intact (because the volume is mounted, not tied to the image).
Rollbacks do not roll the database schema. If a deploy includes a non-reversible migration that has already run, rolling back the image but keeping the DB can break the older image's assumptions. Mitigations:
- Only ship additive migrations (don't drop columns in the same deploy that removes their consumer code — drop them in a follow-up deploy instead).
- For destructive changes, take a Coolify DB snapshot first.
12. Local full-stack dev with docker compose
For developers who want the production image shape locally (no Coolify required):
cp .env.example .env
# Edit at minimum PAYLOAD_SECRET, GAME_TICK_NOTIFY_SECRET.
docker compose up --build
This boots:
postgres— Postgres 17 alpine, persisted to thepgdatanamed volume.app— the production Dockerfile image, talking to the bundled Postgres viaDATABASE_URIrewritten by compose fromPOSTGRES_USER/_PASSWORD/_DB.
App on http://localhost:3000, Postgres on localhost:5432.
Useful commands:
docker compose up --build # rebuild + boot
docker compose down # stop
docker compose down -v # stop + wipe the Postgres volume
docker compose exec app bun run payload game-tick # fire a tick manually
docker compose exec app bun run payload migrate # apply migrations
docker compose exec app sh # shell inside the app container
src/app/(frontend)/layout.tsxgates on auth — first deploy is a clean DB; you'll need to seed an admin user. The existing seed scripts undersrc/tools/seed/run via bun, e.g.docker compose exec app bun run src/tools/seed/<file>.ts(refer to those files for the intended entry point).
13. Quick troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Container restart loops, never healthy | DB connection unreachable | Verify DATABASE_URI resolves from inside the container (docker compose exec app wget -qO- $DATABASE_URI:5432 or nc -zv ...). Coolify: confirm the Postgres and app services are in the same Project so Coolify's network reaches it. |
| 401 loop after deploy | PAYLOAD_SECRET changed without rotating user sessions |
Keep PAYLOAD_SECRET constant across deploys of the same env. If you must rotate, users will need to re-log in (cookies are JWT-signed against this secret). |
| Realtime updates stopped working | More than 1 replica running, OR scheduled tick stopped | Limit replicas to 1. Check the scheduled tasks in Coolify are firing and GameEventLogs show fresh gameTick events. |
| First build OOMs | Host has < 8 GB RAM during Next build | The bun run build script sets --max-old-space-size=8000. Give the Coolify build host at least 8 GB, or lower this in package.json#scripts.build (memory will become the limiter earlier). |
bun run db push fails PostGIS error |
This is expected (must be owner of table spatial_ref_sys) |
Use bun run payload migrate only. See AGENTS.md Gotchas. |
| Token set but bot never logs in / silent bot death | Entrypoint gate, config throw, or revoked/invalid token | Check the container logs for the bot's fail-fast line. Confirm DISCORD_GUILD_ID is set (the config throws at import if missing) and the token is valid/not revoked. The web app is unaffected — the bot is an in-process subordinate. |
Last updated: 2026-08-12. Verify against the live repo if anything looks stale.