docs(06-06): add RUNBOOK.md — env vars, startup, backup, health checks, escalation, failure modes

Covers D-14: all required env vars with examples, Docker Compose startup/shutdown
procedures, PostgreSQL + MinIO backup strategy with restore commands, health check
verification for all 7 services, on-call escalation for 6 alert classes, and
recovery steps for all 7 Phase 6 failure modes (Pitfalls 1, 5, 6, 7 + rate limit
edge cases + Grafana/Loki). docker scout CVE scan gate (D-10) documented in
Section 5 with prerequisites, pass/fail criteria, remediation, and Trivy fallback.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
curo1305
2026-06-04 19:05:05 +02:00
co-authored by Claude Sonnet 4.6
parent 60b621f424
commit 670df192d6
+580
View File
@@ -0,0 +1,580 @@
# DocuVault Operational Runbook
This is the single reference for running DocuVault in production-like environments. It covers
environment variables, startup/shutdown procedures, backup strategy, health checks, on-call
escalation, and common failure-mode recovery. For architecture and rationale see `CLAUDE.md`;
for phase history and roadmap see `.planning/ROADMAP.md`.
> **Last updated:** 2026-06-04
---
## Table of Contents
1. [Environment Variables](#1-environment-variables)
2. [Startup / Shutdown](#2-startup--shutdown)
3. [Backup Strategy](#3-backup-strategy)
4. [Health Checks](#4-health-checks)
5. [Security Gate — docker scout CVE Scan (D-10)](#5-security-gate--docker-scout-cve-scan-d-10)
6. [On-Call Escalation](#6-on-call-escalation)
7. [Common Failure Modes](#7-common-failure-modes)
8. [Phase 6 Deferred Items](#8-phase-6-deferred-items)
---
## 1. Environment Variables
All variables are read from a `.env` file at the repo root (or injected by the container
orchestrator). Copy `.env.example` to `.env` and populate every field before starting services.
> **Security note:** `SECRET_KEY`, `CLOUD_CREDS_KEY`, and all `*_PASSWORD`/`*_SECRET` fields
> must **never** be committed to version control. Rotate them if exposure is suspected.
### PostgreSQL
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `DATABASE_URL` | Yes | Application DSN (`docuvault_app` role — DML only). | `postgresql+psycopg://docuvault_app:****@postgres:5432/docuvault` | `.env` |
| `DATABASE_MIGRATE_URL` | Yes | Migration DSN (`docuvault_migrate` role — DDL). Used by Alembic only. | `postgresql+psycopg://docuvault_migrate:****@postgres:5432/docuvault` | `.env` |
| `POSTGRES_PASSWORD` | Yes | Root postgres superuser password (used by the docker-compose `postgres` service). | `change-me-postgres` | `.env` |
### MinIO
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `MINIO_ENDPOINT` | Yes | Internal Docker hostname:port for MinIO SDK calls. | `minio:9000` | `.env` |
| `MINIO_ACCESS_KEY` | Yes | MinIO application access key. | `docuvault_app` | `.env` |
| `MINIO_SECRET_KEY` | Yes | MinIO application secret key. | `****` | `.env` |
| `MINIO_BUCKET` | Yes | Primary document storage bucket name. | `docuvault` | `.env` |
| `MINIO_PUBLIC_ENDPOINT` | No | Browser-visible hostname for presigned upload/download URLs. Empty → falls back to `MINIO_ENDPOINT`. | `localhost:9000` | `.env` |
| `MINIO_ROOT_USER` | Yes | MinIO root/admin username (used by the `minio` compose service). | `minioadmin` | `.env` |
| `MINIO_ROOT_PASSWORD` | Yes | MinIO root/admin password. | `****` | `.env` |
### Redis / Celery
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `REDIS_URL` | Yes | Redis connection string including password. | `redis://:****@redis:6379/0` | `.env` |
| `REDIS_PASSWORD` | Yes | Redis `requirepass` password (used by the `redis` compose service). | `****` | `.env` |
### Security / Auth
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `SECRET_KEY` | Yes | HMAC/JWT signing secret. Must be a high-entropy random string ≥32 bytes. Rotate to invalidate all active sessions. | `openssl rand -hex 32` output | `.env` |
| `CLOUD_CREDS_KEY` | Yes | Master key for HKDF per-user cloud credential encryption. Must be exactly 32 bytes. | `openssl rand -hex 16` (pad to 32) | `.env` |
| `ADMIN_EMAIL` | Yes | Bootstrap admin account email. Created on first startup if absent. | `admin@example.com` | `.env` |
| `ADMIN_PASSWORD` | Yes | Bootstrap admin account password. Must meet AUTH-01 strength requirements. | `Ch@ngeMe1!` | `.env` |
| `ACCESS_TOKEN_EXPIRE_MINUTES` | No | JWT access token lifetime in minutes. Default: `15`. Do not increase beyond 30. | `15` | `.env` |
| `REFRESH_TOKEN_EXPIRE_DAYS` | No | Refresh token cookie lifetime in days. Default: `30`. | `30` | `.env` |
**Rotation procedure for `SECRET_KEY`:**
1. Generate a new value: `openssl rand -hex 32`
2. Update `.env` and restart the `backend` service.
3. All active JWT sessions are immediately invalidated — users must log in again.
**Rotation procedure for `CLOUD_CREDS_KEY`:**
Requires re-encrypting all stored cloud credentials. Until a migration tool exists, this
rotation is a breaking change — treat the key as long-lived and protect it accordingly.
### CORS / Frontend
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `CORS_ORIGINS` | No | Comma-separated list of allowed CORS origins. Default: `http://localhost:5173`. | `https://app.example.com` | `.env` |
| `FRONTEND_URL` | No | Base URL of the Vue frontend — used in password-reset emails. Default: `http://localhost:5173`. | `https://app.example.com` | `.env` |
| `BACKEND_URL` | No | Externally reachable backend URL — used to construct OAuth callback URLs. Default: `http://localhost:8000`. | `https://api.example.com` | `.env` |
### SMTP (password reset emails)
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `SMTP_HOST` | No | SMTP relay hostname. Leave empty to disable email. | `smtp.sendgrid.net` | `.env` |
| `SMTP_PORT` | No | SMTP port. Default: `587` (STARTTLS). | `587` | `.env` |
| `SMTP_USER` | No | SMTP login username. | `apikey` | `.env` |
| `SMTP_PASSWORD` | No | SMTP login password / API key. | `****` | `.env` |
| `SMTP_FROM` | No | Sender address for outgoing emails. Default: `noreply@docuvault.local`. | `noreply@example.com` | `.env` |
### AI Classification
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `DEFAULT_AI_PROVIDER` | No | Default LLM provider. Default: `ollama`. Options: `ollama`, `openai`, `anthropic`. | `openai` | `.env` |
| `DEFAULT_AI_MODEL` | No | Default model name for the chosen provider. Default: `llama3.2`. | `gpt-4o-mini` | `.env` |
| `SYSTEM_PROMPT` | No | Custom system prompt override for the classifier. Hardcoded fallback lives in `backend/ai/classifier.py`. | `Classify documents concisely.` | `.env` |
### Cloud Storage Providers (Phase 5)
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `GOOGLE_CLIENT_ID` | No | Google OAuth 2.0 client ID for Drive integration. | `1234567890-abc.apps.googleusercontent.com` | Google Cloud Console |
| `GOOGLE_CLIENT_SECRET` | No | Google OAuth 2.0 client secret. | `GOCSPX-****` | Google Cloud Console |
| `ONEDRIVE_CLIENT_ID` | No | Azure AD application (client) ID for OneDrive. | `xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx` | Azure Portal |
| `ONEDRIVE_CLIENT_SECRET` | No | Azure AD client secret value. | `****` | Azure Portal |
| `ONEDRIVE_TENANT_ID` | No | Azure AD tenant. Default: `common` (personal + org accounts). | `common` | `.env` |
### Observability (Phase 6)
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `LOG_LEVEL` | No | structlog minimum level. Default: `INFO`. Options: `DEBUG`, `INFO`, `WARNING`, `ERROR`. | `INFO` | `.env` |
| `LOG_JSON` | No | Emit structured JSON log lines instead of plain text. Default: `false`. Set `true` in production. | `true` | `.env` |
### Load Testing (Phase 6 — dev only)
| Name | Required | Description | Example | Source |
|------|----------|-------------|---------|--------|
| `LOAD_TEST_EMAIL` | No | Email of the pre-created load-test user. Default: `loadtest@example.com`. | `loadtest@staging.example.com` | `.env` (dev only) |
| `LOAD_TEST_PASSWORD` | No | Password for the load-test user. Default: `Loadtest123!@#`. | `MyS3cure!Pass` | `.env` (dev only) |
---
## 2. Startup / Shutdown
### Full stack startup
```bash
# Build images and start all services in the background
docker compose up -d --build
# Tail logs for the first 60 s to confirm everything came up cleanly
docker compose logs -f --tail 50
```
Services start in dependency order:
```
postgres / minio / redis → backend / celery-worker / celery-beat
→ loki → promtail / grafana
→ frontend
```
Access points after a healthy startup:
| Service | URL |
|---------|-----|
| Vue frontend | http://localhost:5173 |
| FastAPI backend (health) | http://localhost:8000/health |
| MinIO console | http://localhost:9001 |
| Grafana (log UI) | http://localhost:3000 |
| Loki (API/ready) | http://localhost:3100/ready |
### Development mode (live reload)
The `docker-compose.yml` already mounts `./backend:/app` with `--reload`, so
Python file changes trigger automatic reload. To force a rebuild after changing
`requirements.txt` or the `Dockerfile`:
```bash
docker compose build backend celery-worker celery-beat
docker compose up -d backend celery-worker celery-beat
```
### Orderly shutdown
```bash
# Stop all containers, preserve volumes (normal dev shutdown)
docker compose down
# Stop and remove named volumes (WARNING: deletes all data)
docker compose down -v
```
> **Warning:** `docker compose down -v` deletes `postgres_data`, `minio_data`,
> `loki_data`, and `grafana_data`. Only use this for a full environment reset.
### celery-beat note (Pitfall 7)
The `celery-beat` service does **not** have `read_only: true` in `docker-compose.yml`
because it writes the `celerybeat-schedule` file to its working directory at startup.
Do not add `read_only:` to the celery-beat block — it will exit immediately with
`Permission denied: 'celerybeat-schedule'`.
---
## 3. Backup Strategy
### PostgreSQL
**Manual backup:**
```bash
docker compose exec -T postgres pg_dump -U postgres docuvault \
> backups/docuvault_$(date +%F).sql
```
**Restore:**
```bash
cat backups/docuvault_YYYY-MM-DD.sql \
| docker compose exec -T postgres psql -U postgres docuvault
```
**Cron pattern (daily at 03:00, keep 7 days):**
```bash
# Add to crontab (crontab -e)
0 3 * * * cd /path/to/docuvault && \
docker compose exec -T postgres pg_dump -U postgres docuvault \
| gzip > backups/$(date +\%F).sql.gz && \
find backups/ -name "*.sql.gz" -mtime +7 -delete
```
### MinIO
**Mirror to local path:**
```bash
docker compose exec minio mc mirror /data /backup/minio-mirror
```
**Mirror to an S3-compatible offsite target:**
```bash
docker compose exec minio mc mirror \
local-minio/docuvault s3://offsite-bucket/docuvault-mirror
```
**Restore from mirror:**
```bash
docker compose exec minio mc mirror /backup/minio-mirror /data
```
### Retention policy (recommendation)
| Frequency | Keep |
|-----------|------|
| Daily | 7 days |
| Weekly | 4 weeks |
| Monthly | 12 months |
Implement via cron + `find … -mtime +N -delete`. Backup automation as a Docker
service is a deferred item (see Section 8).
---
## 4. Health Checks
Run these commands after startup or after an incident to verify each service is
healthy.
| Service | Probe Command | Expected Output |
|---------|--------------|-----------------|
| Backend API | `curl -sf http://localhost:8000/health` | `{"status":"ok",...}` (HTTP 200) |
| Loki | `curl -sf http://localhost:3100/ready` | `ready` |
| Grafana | `curl -sf http://localhost:3000/api/health` | JSON with `"database":"ok"` |
| Postgres | `docker compose exec postgres pg_isready -U postgres -d docuvault` | `accepting connections` (exit 0) |
| MinIO | `docker compose exec minio mc ready local` | exit 0 |
| Redis | `docker compose exec redis redis-cli -a $REDIS_PASSWORD ping` | `PONG` |
| Celery worker | `docker compose exec celery-worker celery -A celery_app inspect ping` | `pong` response |
**All-in-one check script:**
```bash
#!/usr/bin/env bash
set -e
echo "Backend: $(curl -sf http://localhost:8000/health | jq -r .status)"
echo "Loki: $(curl -sf http://localhost:3100/ready)"
echo "Grafana: $(curl -sf http://localhost:3000/api/health | jq -r .database)"
docker compose exec -T postgres pg_isready -U postgres -d docuvault
docker compose exec -T minio mc ready local
docker compose exec -T redis redis-cli -a "$REDIS_PASSWORD" ping
echo "All services healthy."
```
---
## 5. Security Gate — docker scout CVE Scan (D-10)
Run this gate **before every production deploy** and after every base-image rebuild.
Zero critical CVEs is the hard pass requirement.
### Prerequisites
1. Docker Desktop (or Docker Engine with the `scout` plugin) must be installed.
2. Authenticate with Docker Hub — `docker scout` uploads the image manifest to
Docker's CVE analysis service (Pitfall 5):
```bash
docker login
```
Skip if already logged in (check with `docker system info | grep Username`).
### Build the Phase 6 hardened image
The multi-stage `backend/Dockerfile` (Phase 6 D-07) produces a lean runtime image
with `appuser` (uid 1000), no build toolchain, and a read-only-friendly layout:
```bash
cd backend
docker build -t docuvault-backend:phase6 .
```
### Run the CVE scan
```bash
docker scout cves local://docuvault-backend:phase6 --only-severity critical --exit-code
```
| Exit code | Meaning |
|-----------|---------|
| `0` | **PASS** — zero critical CVEs. Safe to deploy. |
| `2` | **FAIL** — one or more critical CVEs found. Do not deploy until resolved. |
A passing scan looks like:
```text
CRITICAL: 0 critical
HIGH: 2 high
MEDIUM: 15 medium
LOW: 40 low
```
### Remediation if the scan fails
1. Pull the latest base image and rebuild:
```bash
docker pull python:3.12-slim
cd backend && docker build --no-cache -t docuvault-backend:phase6 .
docker scout cves local://docuvault-backend:phase6 --only-severity critical --exit-code
```
2. If a Python package is the source of the CVE, bump its pin in
`backend/requirements.txt` to a patched version and rebuild.
3. If the CVE is in the base OS packages (Debian), try pinning to a newer
Python 3.12 patch release (`python:3.12.X-slim`) or, after compatibility
testing, migrate to `python:3.13-slim`.
4. If the CVE cannot be resolved in this release cycle, escalate to the operator
with the full CVE list. The phase must not be marked complete until the
critical count reaches zero.
### Fallback: Trivy (offline / docker scout unavailable)
If `docker scout` is unavailable (offline environment or Docker Hub outage):
```bash
# macOS
brew install trivy
# Debian/Ubuntu
apt install -y trivy
# Run scan
trivy image --severity CRITICAL --exit-code 1 docuvault-backend:phase6
```
Exit code `0` = no critical CVEs. Exit code `1` = critical CVEs found.
### Scan cadence
- **Before every production deploy** — mandatory gate.
- **Weekly scheduled re-scan** — new CVEs are published daily; old images can
become vulnerable after release. CI/CD automation of this cadence is a
deferred item (see Section 8).
---
## 6. On-Call Escalation
DocuVault runs as a **single-operator deployment**. There is no secondary on-call tier.
The operator is the first and only responder. This section documents alert classes and
the self-service recovery action for each so the operator can act at 03:00 without
reading PLAN.md.
### Alert classes and actions
| Alert | Urgency | Immediate action |
|-------|---------|-----------------|
| `backend-down` | **Page immediately** | `docker compose logs --tail 100 backend` — check for `Read-only file system` (Pitfall 1/6) or permission errors. See Section 7 for recovery. |
| `celery-stuck` | Warn within 1 h | `docker compose logs --tail 200 celery-worker` — check for queue depth or task errors. Restart: `docker compose restart celery-worker`. |
| `loki-full-disk` | Warn within 4 h | Loki container >80% disk on `/loki`. Prune old chunks via Loki retention tuning or rotate the `loki_data` volume. |
| `docker-scout-critical-CVE` | Page within 24 h | Rebuild image with `docker pull python:3.12-slim && docker compose build backend`. Re-run Section 5 scan. |
| `quota-violation-spike` | Warn | >10 `quota_exceeded` events/hour from same user → possible abuse. Disable account via admin endpoint. |
| `failed-login-spike` | Warn | >100 `failed_login` events/hour from same IP → verify per-IP rate limiter is firing (10/min). If not, confirm 06-05 deployed and `TRUSTED_PROXY_CIDRS` / proxy topology is correct. |
### Contact
| Role | Contact |
|------|---------|
| Primary on-call | curo1305@curonet.de (project owner) |
| Secondary | *(none — single-operator deployment)* |
---
## 7. Common Failure Modes
### 7.1 `PermissionError: [Errno 13] /tmp/…` in backend logs
**Symptom:** `PermissionError: [Errno 13] Permission denied: '/tmp/tmpXXXXXX'`
**Cause:** Pitfall 1 — `tmpfs` mount missing `mode=1777`. By default, tmpfs mounts
are owned by root and non-root containers (`appuser`, uid 1000) cannot write to them.
**Recovery:**
1. Verify `docker-compose.yml` backend service has:
```yaml
tmpfs:
- "/tmp:mode=1777"
```
2. Recreate the container:
```bash
docker compose up -d --force-recreate backend
```
---
### 7.2 `docker scout` returns "authentication required"
**Symptom:** `docker scout cves` exits with an authentication error.
**Cause:** Pitfall 5 — `docker scout` requires a Docker Hub session to submit the
image manifest to Docker's CVE analysis service.
**Recovery:**
```bash
docker login
# Re-run the scan
docker scout cves local://docuvault-backend:phase6 --only-severity critical --exit-code
```
---
### 7.3 PyMuPDF `Read-only file system` when extracting PDF
**Symptom:** `RuntimeError: Read-only file system` originating from PyMuPDF
(Celery worker or backend during extraction).
**Cause:** Pitfall 6 — PyMuPDF writes font cache data to `/var/cache/fontconfig`
at runtime. Under `read_only: true`, this directory is not writable.
**Recovery:** Add a tmpfs mount for fontconfig in `docker-compose.yml`:
```yaml
services:
backend:
tmpfs:
- "/tmp:mode=1777"
- "/var/cache/fontconfig" # ← add this line
celery-worker:
tmpfs:
- "/tmp:mode=1777"
- "/var/cache/fontconfig" # ← add this line
```
Then recreate:
```bash
docker compose up -d --force-recreate backend celery-worker
```
---
### 7.4 `celery-beat` exits with `Permission denied: 'celerybeat-schedule'`
**Symptom:** `celery-beat` container exits immediately with permission denied on
`celerybeat-schedule`.
**Cause:** Pitfall 7 — `read_only: true` was accidentally added to the `celery-beat`
service block. celery-beat writes its schedule file to its working directory at startup.
**Recovery:** Remove `read_only:` from the `celery-beat` block in `docker-compose.yml`:
```yaml
# celery-beat: NOT hardened — writes celerybeat-schedule (Pitfall 7)
celery-beat:
build: ./backend
# read_only: true ← remove this line if present
...
```
Then recreate:
```bash
docker compose up -d --force-recreate celery-beat
```
---
### 7.5 External clients hit rate limit despite low traffic
**Symptom:** After deploying 06-05, external clients receive `429 Too Many Requests`
at far below the expected threshold.
**Cause:** A reverse proxy (nginx, Caddy, load balancer) sits in front of the backend
but its IP is not in the trusted proxy CIDR list. The rate limiter sees the proxy's IP
as the client IP for all traffic and exhausts the per-IP limit immediately.
**Recovery:** The trusted proxy CIDRs are hardcoded in `backend/deps/utils.py`
as `_TRUSTED_PROXY_NETS`. Add the proxy's CIDR to that list and redeploy:
```python
_TRUSTED_PROXY_NETS = [
ipaddress.ip_network("127.0.0.0/8"),
ipaddress.ip_network("172.16.0.0/12"),
ipaddress.ip_network("192.168.0.0/16"),
ipaddress.ip_network("::1/128"),
ipaddress.ip_network("10.0.0.0/8"), # ← add your proxy CIDR here
]
```
Then rebuild and redeploy the backend.
---
### 7.6 Per-account 429s when expecting per-IP 429s
**Symptom:** A single user receives 429 errors on authenticated endpoints far earlier
than the 100 req/min per-account limit, or all users receive 429s at the same time.
**Cause:** A route handler is missing the first-line assignment
`request.state.current_user = current_user`. Without it, `_account_key` in
`services/rate_limiting.py` falls back to `request.client.host`, making the limiter
key per-IP instead of per-account.
**Recovery:** Find the handler missing the assignment and add it as the first
executable statement inside the function body:
```python
@router.get("/some-endpoint")
@account_limiter.limit("100/minute")
async def some_endpoint(request: Request, current_user: User = Depends(get_regular_user)):
request.state.current_user = current_user # ← first line, always
...
```
---
### 7.7 Grafana loads but no Loki datasource
**Symptom:** Grafana UI loads but the Loki datasource shows "Data source connected and
labels found" or cannot connect.
**Cause:** Loki container not ready when Grafana started, or `loki-config.yaml` is
malformed.
**Recovery:**
1. Check Loki readiness:
```bash
curl -sf http://localhost:3100/ready
docker compose logs --tail 50 loki
```
2. If Loki is not ready, look for YAML parse errors in its logs.
3. Restart the stack after Loki is confirmed healthy:
```bash
docker compose restart grafana
```
---
## 8. Phase 6 Deferred Items
The following improvements are out of scope for Phase 6. They are documented here so the
next operator knows the current runbook acknowledges them:
| Item | Current state | Future path |
|------|---------------|-------------|
| **HTTPS/TLS termination** | Backend serves plain HTTP on port 8000. | Add nginx or Caddy in front; see [nginx reverse proxy pattern](https://docs.nginx.com/nginx/admin-guide/web-server/reverse-proxy/). The runbook will gain a TLS section when this ships. |
| **Horizontal scaling** | Rate limit counters are in-memory (single-instance only). | Switch to Redis-backed counters (slowapi supports this); add uvicorn worker count config. Phase 7+ concern. |
| **CI/CD pipeline** | `docker scout cves` and Locust load tests are run manually. | GitHub Actions workflow for automated scan + load test on every PR. Deferred — no CI setup exists yet. |
| **Backup automation** | `pg_dump` and `mc mirror` run manually or via user-managed cron. | Add a `backup` service to `docker-compose.yml` that runs the cron commands inside the stack. Documented procedure above is the current state. |