# Phase 1 Operations, Rollout, and Rollback

## Safe rollout

1. Deploy migrations through `0102` with capture enabled, publication and activation approval disabled, reverse delivery disabled, and `RB_CONFIRMED_SCHEDULE_ENABLED=False`.
2. Verify request/engine behavior, event version ordering, pending lag, database transaction timing, Kafka max-poll interval, deadlocks/timeouts, and post-commit delivery retries.
3. Seed in bounded batches with `seed_resource_booking_outbox --aggregate ... --after-id ...`; retain the reported cursor.
4. Read the authenticated `/api/integration/resource-booking/v1/snapshot` endpoint. Reuse its `snapshot_watermark` and `next_cursor` for every page; treat the older JSONL exporter as diagnostic only.
5. Deploy the RB-owned adapter disabled; pass contract, security, replacement/tombstone, occurrence, idempotency, load, and reconciliation suites.
6. Mark TT resources mirror-only in Resource Booking and prove conflicting RB writes are rejected.
7. Select/configure the approved transport, enable adapter application, then enable Timetabler publication gradually. Monitor both outbox and adapter application independently.

Staging deployment is enforced by `.github/workflows/deploy.yml` and
`deploy/phase1-publisher-deployment-gate.sh`. The workflow defaults publication and
approval off, stops any old PM2 publisher before the base deploy, writes the protected
staging values atomically, verifies migrations through `0102`, reloads the API, and leaves the
publisher stopped. Only an approved deployment with attested non-empty identities,
Kafka/topic/security settings, snapshot token, publication=true and approval=true may
start exactly one publisher, and it must pass liveness. Any gate failure forces
publication/approval/reverse delivery false again. Do not bypass this with a manual
`pm2 start`.

### Scheduling Engine response consumer deployment

Every repository-managed staging deployment re-registers the existing
`tt_response` PM2 app from `deploy/tt_response.ecosystem.config.cjs`. The
process definition pins the exact deployed checkout, working directory, and
`venv/bin/python`; it also sets `VIRTUAL_ENV` and puts `venv/bin` first in
`PATH`. `deploy/restart_tt_response.sh` first proves the exact Python virtualenv
prefix, compiles the worker, and imports its external dependencies without
initializing Django apps or opening database/Redis connections. It then
restarts the existing app exactly once.

The deployment then verifies the PM2 ID was preserved, the restart counter
advanced exactly once, the live `/proc` command/environment/cwd match the
virtualenv contract, no other PM2 process changed, and the new PID/restart
counter remain stable for 20 seconds. Only after those checks pass is the PM2
process list saved. A mismatch fails the deployment and leaves the Phase 1
publisher/approval rollback gate in force.

## Operator commands

- `integration_outbox_status --errors 20`: state totals, oldest age, attempts, correlation, and sanitized errors.
- `preflight_resource_booking_phase1`: default-off schema/order/configuration probe; add `--require-enabled` only in the managed publisher pre-start check.
- `publish_resource_booking_outbox --once`: bounded publisher execution after activation; omit `--once` for the supervised worker loop.
- `probe_resource_booking_publisher --liveness`: secret-safe configuration and heartbeat liveness probe.
- `publish_post_commit_deliveries --limit 100`: retry durable legacy/microservice Kafka effects.
- `replay_resource_booking_outbox EVENT_ID --reason "..." --actor-id ID`: retry the complete dead-letter source transaction under its same immutable ID/sequence; all members must be dead-lettered together.

Alert on sustained pending/retry age beyond `RB_INTEGRATION_LAG_ALERT_SECONDS`, any dead letter, repeated version gaps/hash conflicts, lock/statement timeout, transaction or max-poll budget breach, post-commit delivery dead letters, and reconciliation divergence. Logs correlate request ID, engine receipt, change set, event/delivery ID, attempts, and replay causation without secrets or full sensitive payloads.

Outbox payloads and applied-response receipts are retained for at least the configured 90/365 days. Automatic deletion is intentionally absent: archive tooling and an approved adapter audit/reconciliation watermark must exist before purging ordering, idempotency, or replay evidence.

## Failure handling

- Before DB commit: the entire mutation, receipt, event, and durable side effects roll back; Kafka redelivery retries.
- After DB commit/before response offset commit: the matching receipt makes redelivery a no-op; outbox and durable delivery workers continue independently.
- Redis/WebSocket failure: database state stays committed and clients refetch canonical state; failure is logged. No database reapplication occurs.
- Downstream Kafka failure: `PostCommitDelivery` retries with a stable idempotency request ID and dead-letters after the configured maximum.
- Adapter outage: disable publication if lag threatens capacity, retain captured change sets, repair the adapter, replay complete dead-letter transactions in source-sequence order, then reconcile before resume.
- Hash mismatch for an existing engine request ID: quarantine operationally; never apply the conflicting body.

## Rollback

Disable `RB_INTEGRATION_PUBLISH_ENABLED` first and remove/false the activation approval. Stop the checked-in PM2 publisher (the systemd file is a reference for non-staging environments), reload the API environment, and persist the stopped supervisor state. Keep outbox capture, receipts, and mirror-only restrictions enabled. Stop the adapter only after recording in-flight acknowledgements. Never renumber/purge source transactions. Reconcile before resume. If the atomic engine consumer itself must be reverted, pause response consumption and prove there is no ambiguous in-flight response first.

The migration is additive. Do not drop integration tables during operational rollback.
