# AMS-Shiva — updating a deployed server

`DEPLOYMENT.md` is for the **first** install. This is for every time after it:
new code has been written, the folder has been uploaded, and the running site
has to pick the change up.

The one rule behind everything below: **uploading a file changes nothing.** The
containers run a copy of the source baked into an image. Until the image is
rebuilt, the site keeps serving exactly what it served before.

---

## 1. Before you upload

| Check | Why |
|---|---|
| **Exclude `.env` from the upload** | It holds this server's passwords, keys, ports and tenant. Overwriting it with the repo's copy (or with `.env.server.example`) is the single fastest way to destroy a working deployment. |
| **Exclude `storage/`** | Uploaded documents live there. An upload can delete or re-own them. |
| **Take a backup if the change carries migrations** | `bash "default server details/scripts/backup.sh"`. A migration cannot be rolled back automatically — see §7. |
| **Let the upload finish** | A half-transferred file looks identical to a complete one, and gets built. |

Upload everything else — `backend/`, `ams-frontend/`, `infra/`, `ai/`,
`print-helper/`, `default server details/` — over the existing folder.

---

## 2. One command

```bash
cd /var/www/<your-domain>
sudo bash "default server details/scripts/update.sh"
```

It asks nothing, keeps the `.env` already on the server, and stops at the first
real failure rather than leaving a half-updated stack serving traffic.

`sudo` because it repairs storage ownership and reads master's `.env`.

### Flags

| Flag | Effect |
|---|---|
| `--no-uat` | Skip the end-to-end test. Faster, and blind. |
| `--seed` | Also re-run the role/user seeder. **Resets every seeded password** to its `.env` tier value — off by default for that reason. |
| `--printers` | Re-seed the label printers from `backend/printer/` without asking. |
| `--rollback` | Restore the images from before the last update and restart. |
| `--yes` | Never prompt. For unattended runs. |

---

## 3. What the eight steps do

**1/8 — Checking what we are updating.**
Refuses to run if `.env` has no `APP_DOMAIN` or no `POSTGRES_PASSWORD`, because
that is what the *template* looks like, not a deployed config — the surest sign
the upload overwrote it. Also strips Windows line endings (CRLF survives an
upload from Windows and breaks bash and `.env` parsing in ways that surface much
later as random failures), confirms Docker is reachable, and requires ~3GB free.

**2/8 — Saving a rollback point.**
Copies `.env` into `.deploy-state/env-backups/` (last 10 kept — that directory
holds every secret the site has used, so it is not left to grow) and records the
**image IDs** the running containers use. The IDs, not the tags: Compose reuses
the same tag on every build, so a tag is not a rollback point.

**3/8 — New settings introduced by this code.**
Diffs `.env.server.example` against the live `.env` and fills in any key the new
code reads but the server does not have. Nothing else warns you about this — the
container reads an unset variable and behaves as if the feature is simply off.
Keys with a template default are added silently; keys without one are asked for.

**4/8 — Re-reading master's shared token.**
`INTERNAL_SERVICE_TOKEN` belongs to master. If it was rotated there, a stale copy
here keeps working until the Redis cache expires and then fails **every** login
with "Invalid tenant" — minutes later, with every container still healthy.

**5/8 — Building images, and the signing keys.**
Also checks `backend/keys/` before anything restarts. The JWT keypair is bind-
mounted from the host, never baked into the image, so an upload that misses it —
or leaves it owned by root — takes the API down completely once the new
containers start. Generated if absent, and always re-owned to uid 1001.

`build api web`. This is the step that actually picks up the new code. Unchanged
Docker layers are cached, so a small change is fast. Storage ownership is
re-asserted to uid 1001 afterwards, because an upload can leave the bind mount
owned by root, after which every in-app upload fails with what looks like an
application bug.

**6/8 — Applying database migrations.**
`alembic upgrade head`, run **before** the new containers take over. That
ordering is deliberate: if a migration fails, the **old** code is still running
against the **old** schema — a working site — and the script stops without
restarting anything.

**7/8 — Restarting on the new images.**
`up -d`, then polls `/api/v1/health` until the API answers. If it never does, the
script fails loudly and tells you how to roll back.

**7b/8 — Label printers.**
Asks whether to re-seed them from `backend/printer/`. New printer files do
nothing until this runs. Re-running updates by name rather than duplicating, so
answering yes every time is safe, and a failure here is never fatal — a missing
printer list is a nuisance, a failed update over one is not. `--printers` skips
the question, `scripts/printers.sh` does it separately.

**8/8 — End-to-end test.**
Runs `uat.sh`: containers, loopback ports, the public URL through Apache, the
master registry, an unknown-tenant probe, tenant data counts, **a real login**,
the pre-login endpoints the browser calls, and document storage.

---

## 4. Doing it by hand

If you would rather not run the script, this is the same sequence.

```bash
cd /var/www/<your-domain>
DC=(docker compose --project-directory . -f "default server details/docker-compose.server.yml")

# 1  Safety
cp .env .env.backup
bash "default server details/scripts/backup.sh"          # if migrations are involved

# 2  Note the current images, in case you need to go back
docker inspect -f '{{.Image}}' "$("${DC[@]}" ps -q api)"
docker inspect -f '{{.Image}}' "$("${DC[@]}" ps -q web)"

# 3  New .env keys this code introduced
diff <(grep -oE '^[A-Z][A-Z0-9_]*=' "default server details/.env.server.example" | sed 's/=$//' | sort -u) \
     <(grep -oE '^[A-Z][A-Z0-9_]*=' .env | sed 's/=$//' | sort -u) | grep '^<'

# 4  Build
"${DC[@]}" build api web

# 5  Databases up, then migrate — before restarting anything
"${DC[@]}" up -d db cache
"${DC[@]}" run --rm --no-deps -T api alembic upgrade head

# 6  Restart on the new images
"${DC[@]}" up -d

# 7  Prove it
curl -sf "http://127.0.0.1:$(sed -n 's/^APP_API_PORT=//p' .env)/api/v1/health" && echo OK
bash "default server details/scripts/uat.sh"
```

---

## 5. What to rebuild for which change

The script always rebuilds both services, which is correct and costs a few
minutes. When you want to do the minimum by hand:

| Changed | Command |
|---|---|
| `backend/**` | `"${DC[@]}" build api && "${DC[@]}" up -d api worker` |
| `ams-frontend/**` | `"${DC[@]}" build web && "${DC[@]}" up -d web` |
| Both | `"${DC[@]}" build api web && "${DC[@]}" up -d` |
| `.env`, backend keys only | `"${DC[@]}" up -d api worker` — no build |
| `.env`, any `VITE_*` / `AES_*` / tenant options / domain | `"${DC[@]}" build web && "${DC[@]}" up -d` |
| New Alembic migration | `"${DC[@]}" run --rm --no-deps api alembic upgrade head` |
| `roles.json` or the seeder | `"${DC[@]}" build api && bash "default server details/scripts/seed.sh" SLUG VERTICAL` |
| `backend/printer/*.json` | `"${DC[@]}" build api && bash "default server details/scripts/printers.sh"` |
| `apache/*.conf` | `sudo cp … /etc/apache2/sites-available/ && sudo systemctl reload apache2` |
| `docker-compose.server.yml` | `"${DC[@]}" up -d` |

Two traps in that table:

- **`worker` runs the same image as `api`.** Restarting only `api` leaves the
  worker on old code — imports, OCR, scheduled reports and expiry checks all keep
  running the previous version, with every container looking healthy.
- **Frontend values are build-time.** `VITE_*` is compiled into the bundle.
  Changing one in `.env` and restarting does nothing; it needs `build web`.

After a frontend rebuild, hard-reload the browser (Ctrl+Shift+R) — a cached
bundle will otherwise hide the change.

---

## 6. Verifying

```bash
"${DC[@]}" ps                        # every service Up / healthy
"${DC[@]}" logs --tail=50 api        # no tracebacks on boot
bash "default server details/scripts/uat.sh"
```

---

## 7. Rolling back

```bash
sudo bash "default server details/scripts/update.sh" --rollback
```

Restores the image IDs recorded in step 2/8 and restarts.

**It restores images only.** A migration that has already run is *not* undone —
this project generates no Alembic downgrades, and inventing one would be worse
than saying so. If an update fails *after* migrating, restore the database from
`backup.sh` instead. That is why the backup in §1 matters for any change
carrying a migration.

`docker image prune` deletes the previous image. Do not prune between an update
and the decision to keep it.

---

## 8. When an update goes wrong

| Symptom | Cause | Fix |
|---|---|---|
| Script stops: "`.env` has no `APP_DOMAIN`" | the upload overwrote `.env` | restore from `.deploy-state/env-backups/`, re-run |
| Script stops at migrations | a migration failed | nothing was restarted — the old site is still up. Read the error, fix, re-run |
| API never becomes healthy | new code fails at boot | `"${DC[@]}" logs --tail=80 api`, then `--rollback` |
| UAT fails on **login rejected (400)** | master cannot resolve the slug | check `MASTER_API_URL`, `INTERNAL_SERVICE_TOKEN`, tenant status in the master console |
| UAT fails on **credentials rejected (401)** | seeded password differs from `SEED_SUPERADMIN_PASSWORD` | `seed.sh SLUG VERTICAL`, or read the real value out of `.env` |
| UAT fails on **the bundle does not carry the slug** | web was not rebuilt | `"${DC[@]}" build --no-cache web` |
| A feature behaves as if it is switched off | new code reads a `.env` key the server does not have | step 3/8 handles this; if you updated by hand, run the `diff` from §4 |
| Background jobs stopped after the update | only `api` was restarted | `"${DC[@]}" up -d worker` |
| Frontend change not visible | cached bundle, or web not rebuilt | hard-reload; then `build web` |
| Uploads fail with permission errors | storage left owned by root by the upload | `sudo chown -R 1001:1001 ./storage/documents` |
| **Every request 500s, log says `JWT key not found at /app/keys/private.pem`** | the keypair is missing from `backend/keys/`, or is owned by root with mode 600 — the container runs as uid 1001 and cannot read it. The keys are **not** in the image; they exist only on the host. | `sudo chown -R 1001:1001 backend/keys && sudo chmod 600 backend/keys/private.pem`, then restart `api worker`. If the files are absent, generate them (step 5/8 now does this automatically). |
| Build fails: no space left | images accumulated | `docker image prune -f` — **not** between an update and a possible rollback |

Logs first, always:

```bash
bash "default server details/scripts/logs.sh" api
```
