generated from john/python-template
This commit is contained in:
@@ -0,0 +1,49 @@
|
||||
# Backup and Restore (V6.0 Phase 4)
|
||||
|
||||
This guide defines operational backup/restore for the Docker PostgreSQL runtime and Synology replication.
|
||||
|
||||
## 1. Backup artifacts
|
||||
|
||||
- Primary local backup location: `./data/backups`
|
||||
- Backup format: PostgreSQL custom dump (`pg_dump -Fc`)
|
||||
- Naming: `postgres-YYYYMMDD-HHMMSS.dump` (UTC timestamp)
|
||||
|
||||
## 2. Creating backups
|
||||
|
||||
Use the scripted command:
|
||||
|
||||
```bash
|
||||
sh deploy/backup/create_postgres_backup.sh
|
||||
```
|
||||
|
||||
Optional environment overrides:
|
||||
|
||||
- `BACKUP_DIR` (default `./data/backups`)
|
||||
- `BACKUP_RETENTION_DAYS` (default `14`)
|
||||
- `SYNOLOGY_BACKUP_DIR` (if set, backup is copied to this mounted path)
|
||||
- `ENV_FILE` (default `.env.production`)
|
||||
- `COMPOSE_FILE` (default `docker-compose.production.yml`)
|
||||
|
||||
Example with Synology mount:
|
||||
|
||||
```bash
|
||||
SYNOLOGY_BACKUP_DIR=/mnt/synology/transcription-backups sh deploy/backup/create_postgres_backup.sh
|
||||
```
|
||||
|
||||
## 3. Restoring from backup
|
||||
|
||||
Restore requires downtime for app + worker writes.
|
||||
|
||||
1. Stop app and worker:
|
||||
- `docker compose --env-file .env.production -f docker-compose.production.yml stop app worker`
|
||||
2. Restore:
|
||||
- `sh deploy/backup/restore_postgres_backup.sh ./data/backups/postgres-YYYYMMDD-HHMMSS.dump`
|
||||
3. Start app and worker:
|
||||
- `docker compose --env-file .env.production -f docker-compose.production.yml start app worker`
|
||||
4. Validate `/healthz` and run one smoke workflow.
|
||||
|
||||
## 4. Retention and recovery targets
|
||||
|
||||
- Retention baseline: keep at least 14 days of backups locally.
|
||||
- Synology copy: replicate each new backup to DS420j mounted path.
|
||||
- Periodic restore drill: run at least once per release cycle to verify recovery.
|
||||
@@ -12,6 +12,7 @@ This runbook is the operational checklist for releasing and monitoring the trans
|
||||
- `OPENROUTER_API_KEY`
|
||||
- `DATABASE__*`
|
||||
- filesystem paths for data/logs/backups.
|
||||
- `CLOUDFLARE_TUNNEL_TOKEN`
|
||||
5. Confirm schema contract alignment is current:
|
||||
- `src/transcription/db/models.py`
|
||||
- `docs/schema.md`
|
||||
@@ -33,6 +34,8 @@ This runbook is the operational checklist for releasing and monitoring the trans
|
||||
4. Verify log flow:
|
||||
- stdout aggregation receives events
|
||||
- file logs are written under `./data/logs`
|
||||
5. Create a fresh PostgreSQL backup after successful deployment:
|
||||
- `sh deploy/backup/create_postgres_backup.sh`
|
||||
|
||||
## 3. Rollback triggers and actions
|
||||
|
||||
@@ -50,6 +53,8 @@ This runbook is the operational checklist for releasing and monitoring the trans
|
||||
4. Preserve incident evidence:
|
||||
- `./data/logs`
|
||||
- relevant DB rows (`job`, `job_source`, `execution_attempt`)
|
||||
5. If persistence regression is confirmed, restore the latest valid DB dump:
|
||||
- `sh deploy/backup/restore_postgres_backup.sh <dump-file>`
|
||||
|
||||
## 4. Post-release monitoring checklist
|
||||
|
||||
@@ -90,10 +95,17 @@ This runbook is the operational checklist for releasing and monitoring the trans
|
||||
### Cloudflare ingress/access failure
|
||||
|
||||
1. Check `cloudflared` container logs for ingress parse, DNS, or auth failures.
|
||||
2. Confirm `deploy/cloudflared/config.yml` tunnel UUID and hostname mappings are correct.
|
||||
3. Confirm `deploy/cloudflared/credentials.json` matches the tunnel configured in Cloudflare.
|
||||
2. Confirm `deploy/cloudflared/config.yml` hostname mappings are correct.
|
||||
3. Confirm `CLOUDFLARE_TUNNEL_TOKEN` in `.env.production` matches the tunnel configured in Cloudflare.
|
||||
4. Confirm Cloudflare Access app policy includes the intended identity/group for that hostname.
|
||||
|
||||
### Backup or restore failure
|
||||
|
||||
1. Verify `postgres` container is healthy and accepting connections.
|
||||
2. Confirm dump file exists and is non-zero size.
|
||||
3. Re-run backup/restore scripts with explicit `ENV_FILE` and `COMPOSE_FILE` if using non-default paths.
|
||||
4. If Synology copy fails, keep local backup and resolve mount/network before next backup cycle.
|
||||
|
||||
## 6. Dependency upgrade policy
|
||||
|
||||
Dependencies are declared in `pyproject.toml` and resolved through the committed
|
||||
|
||||
@@ -60,6 +60,12 @@ Implemented workflow references:
|
||||
2. Validate restore drill from Synology-hosted dump artifacts.
|
||||
3. Document rollback procedure for deployment failure and migration failure scenarios.
|
||||
|
||||
Implemented workflow references:
|
||||
|
||||
- `deploy/backup/create_postgres_backup.sh`
|
||||
- `deploy/backup/restore_postgres_backup.sh`
|
||||
- `docs/backup_restore.md`
|
||||
|
||||
## 3.5 Validation and release gate
|
||||
|
||||
1. `/healthz` confirms app and worker healthy in deployed environment.
|
||||
|
||||
Reference in New Issue
Block a user