# Phase 1 Staging Bootstrap NO-GO

## Decision

**Phase 1 staging acceptance: NO-GO.**

**Phase 2 production implementation: HARD-BLOCKED.**

Resource Booking staging was reset and the live Staff and Location catalogues both
show `No data available`. This is consistent with the deployed configuration:
Resource Booking ingestion is false and Timetabler publication is disabled. Shared
scope/topic/token values are now reported configured, but runtime bootstrap and
reconciliation have not run. Do not enable either side independently.

## Verified deployment and configuration evidence

| Area | Evidence | Readiness result |
| --- | --- | --- |
| Timetabler baseline code | Deployed SHA `b29ee4c21a04f7f4cb7cc7b36faa2484cf294007`; GitHub Actions run [31023018291](https://github.com/Mayvins/timetabler-be/actions/runs/31023018291) succeeded | Baseline code deployed; deployment success did not activate or attest integration. The v2/gated release needs its own exact SHA/run record |
| Resource Booking code | v2 receiver `a3dcecd0c50daff88bcd030dd5dc039b09d98feb` passed staging run [31099640196](https://github.com/Mayvins/resource-booking-be/actions/runs/31099640196) | Matching receiver deployed; ingestion/repair/reverse remain false pending snapshot reconciliation |
| TT capture | `RB_INTEGRATION_CAPTURE_ENABLED=true` by default | Must be verified on the server and confirmed by outbox rows; default is not runtime proof |
| TT publication | `RB_INTEGRATION_PUBLISH_ENABLED=false`, `RB_INTEGRATION_TRANSPORT=disabled`, `RB_INTEGRATION_KAFKA_TOPIC` unset by default | **Not ready; keep disabled** |
| RB ingestion | `RB_TIMETABLER_SYNC_INGESTION_ENABLED=false` when absent | **Not ready; keep disabled** |
| Deployment automation | Baseline workflow only invoked `~/deploy-staging.sh`. The Timetabler v2 candidate adds protected-environment value transport, atomic `.env` attestation, PM2 singleton start/stop/probe, and failure rollback | Candidate behavior passed local fake-supervisor QA, but the exact staging run/config/process attestation is still required and publication stays off |
| Catalogue state | Live Staff and Location pages show no data after RB reset | Explicit staging acceptance failure |

## Compatibility and bootstrap state that blocks safe publication

The deployed baseline producer and older RB consumer could not be connected merely by
setting environment variables. The Timetabler v2 candidate and deployed RB receiver
now align on the minimum provider contract, but activation remains unsafe until the
fixed-watermark bootstrap/reconciliation and operational gates pass:

1. Exactly one bounded canonical outer transaction object/Kafka record per finalized
   Timetabler change set, with `transaction_count`, `complete=true`, one hash and
   `member_only` Staff/Location/activity final-state members.
2. One positive globally commit-ordered `source_sequence` per source scope/change set,
   a single source-scope Kafka key, whole-transaction retry/dead-letter/replay, and
   aggregate versions inside each member.
3. Deterministic absolute calendar/occurrence facts with aware local/UTC boundaries,
   IANA timezone, week references and explicit DST behavior.
4. Bearer-authenticated `/api/integration/resource-booking/v1/snapshot` pagination at a
   fixed transaction watermark. Latest aggregate states are regrouped by originating
   source sequence into the same composite outer contract, never multiple hashes for
   one sequence.
5. Kafka TLS/SASL/idempotence configuration, preflight/health/liveness, publisher
   heartbeat, protected-environment value transport, and a default-off/fail-closed PM2
   singleton deployment gate.

Resource Booking receiver `a3dcecd0c50daff88bcd030dd5dc039b09d98feb` is now deployed
for the checked-in v2 fixture, whose exact SHA-256 is
`bf4ae32e685af60415f768fd567fab0423086d5c67ec41dc262142a51e2f2d92`.
Its ingestion, repair and reverse flags remain false. This resolves the code-shape
blocker but not bootstrap acceptance: no configuration flag may bypass the required
fixed-watermark snapshot reconciliation and mirror-only verification.

An executable probe using the strict Resource Booking parser from exact commit
`fc9ecf670d8a1a542cba8b3efe7d1b057779285d` rejected the baseline representative Timetabler
envelope with:

```text
EXPECTED_REJECTION ContractError: Unknown Timetabler envelope fields: actor_id, aggregate, change_set_id, committed_at, committed_state, event_version, ordering_key, origin, request_id
```

The architecture requires transformation to remain Resource Booking/adapter-owned.
Therefore the preferred correction is to make the RB adapter accept and transform an
approved TT canonical transaction contract, while Timetabler adds only the minimal
transaction completeness/source watermark and authenticated generic snapshot support
needed to publish its own source facts. Do not make Timetabler emit RB persistence
schemas.

## Required contract and code corrections before staging activation

1. Record the deployed Resource Booking receiver's contract-test evidence and jointly approve the checked-in v2 provider/consumer fixtures for Staff,
   Location, activity/allocation replacement, tombstone, bulk/change-set completeness,
   timezone/DST, payload limits, ordering, replay and acknowledgement/error classes.
2. Run Resource Booking consumer tests proving the one-change-set composite, global
   source watermark, Kafka partition/offset behavior, fixed-watermark snapshot and
   replay-after-watermark semantics against the real provider output.
3. Run the deployed RB-owned transformation layer and checked-in shared fixtures
   against the real Timetabler producer/snapshot output before applying any records.
4. Deploy the Timetabler candidate with publication/approval false and record the exact
   SHA/run, migration `0102`, redacted protected-environment attestation, zero publisher
   instances, API health and pending source watermark.
5. Prove broker topics, partitions, ACL/TLS/SASL identities, retention, maximum message
   size, DLQ, `read_committed`, consumer group and manual offset behavior in staging.

## Secret-safe staging configuration attestation

Operators must record the following values or present/absent/hashed attestations in an
immutable run artifact. Never paste passwords, tokens, private keys or full connection
strings.

### Timetabler

- exact deployed Git SHA and migrations through `api.0102` applied;
- `RB_INTEGRATION_CAPTURE_ENABLED=true`;
- `RB_INTEGRATION_PUBLISH_ENABLED=false` until the activation step;
- approved/attested `RB_INTEGRATION_DEPLOYMENT_ID`, shared `RB_INTEGRATION_SOURCE_SCOPE=default`, and schema/contract version;
- approved `RB_INTEGRATION_TRANSPORT=kafka` and topic only after contract tests pass;
- Kafka bootstrap endpoint present, TLS/SASL mode and producer principal name, with
  secret values redacted;
- publisher service definition/version, execution account, enabled/running state,
  restart policy, log destination and lag/dead-letter probe;
- outbox counts by state, oldest age, maximum ID, minimum/maximum aggregate versions,
  dead letters and quarantine count.

### Resource Booking

- exact deployed Git SHA and Phase 1 schema applied;
- `RB_TIMETABLER_SYNC_PHASE=phase1`;
- ingestion, repair and reverse allocation all false before bootstrap;
- source scope exactly matching the approved TT deployment/source scope;
- command topic, DLQ, consumer group and broker identity matching the TT/broker record;
- HTTPS Timetabler source base URL and snapshot path; service token presence/rotation
  identifier only, never the token;
- consumer service installed but stopped, with restart/drain/probe definitions;
- preflight result, inbox/checkpoint/DLQ/quarantine counts and reset baseline timestamp.

### Broker

- exact cluster/environment identifier and broker version;
- approved topic and DLQ existence, partition count/keying, replication, retention and
  maximum message bytes;
- TT producer principal can produce/describe only the required topic;
- RB consumer principal can read/describe that topic and its consumer group and can
  write the approved DLQ where required;
- a non-production canary proves TLS/SASL, headers, payload size, `read_committed`,
  manual offset/rebalance and restart behavior without using production facts.

## Coordinated fixed-watermark bootstrap and catch-up

Run this only after the contract/code corrections and configuration attestations above
pass against the exact deployed builds.

1. **Freeze the change window.** Name TT, RB, QA and operations operators; record start
   time, SHAs, broker positions and rollback owner. Keep TT capture on, TT publication
   off, RB ingestion/repair/reverse delivery off, and shared RB resources unavailable.
2. **Preflight both stores read-only.** Verify TT migration/outbox integrity and current
   Staff/Location/activity counts. Run RB Phase 1 preflight against the reset database.
   Resolve identity, type, code, term/calendar/timezone and unsafe allocation findings.
3. **Seed complete TT source facts into the outbox.** In bounded supervised transactions,
   run `seed_resource_booking_outbox` for `staff`, then `location`, then `activity`,
   advancing and recording every `next_after_id`. Capture remains on so concurrent
   mutations receive later source versions. This step mutates only the TT integration
   outbox and requires an approved change ticket.
4. **Capture one fixed high-water mark.** After all seed cursors finish, obtain the
   approved snapshot watermark from the same source-order mechanism used by live
   events. Record row/event counts and the exact maximum source version/change set.
5. **Dry-run the complete RB snapshot.** With ingestion and repair still off, read all
   pages at the fixed watermark and require stable watermark/cursors, valid hashes,
   complete change sets and expected Staff/Location/activity totals. No page may be
   filtered for the acceptance result.
6. **Apply the bootstrap through the canonical RB service.** Temporarily enable only the
   supervised RB repair gate, execute the fixed-watermark snapshot apply with an
   identified operator, then disable repair again. Do not direct-load tables. Every
   imported Staff/Location must have `source_system=timetabler` and
   `integration_mode=mirror_only`.
7. **Reconcile before live ingestion.** Run a new unfiltered dry-run against the same
   high-water mark. Require zero unexplained missing/stale/orphaned facts, zero identity
   or timezone divergence, zero mirror-only violations and healthy projections.
8. **Prepare replay after the mark.** Prove the Kafka offset/source-version boundary for
   the first transaction after the snapshot high-water. Earlier seed events must be
   safely recognized as already applied; later events must remain ordered and none may
   be skipped. Record the broker offsets and TT outbox IDs/change sets used for proof.
9. **Start canary delivery.** Start the RB consumer with ingestion explicitly enabled
   for the approved source scope, then start/enable the TT publisher for a bounded
   canary. Confirm publisher acknowledgement, RB inbox durable outcome, manual offset,
   mirror application and projection separately. A transport acknowledgement alone is
   not success.
10. **Verify live catch-up and user behavior.** Require zero lag/gaps/DLQ/quarantine,
    repeat unfiltered reconciliation to zero unexplained drift, compare source and RB
    Staff/Location counts, verify both catalogues display data, and prove an RB booking
    against a Timetabler mirror is rejected with `resource_sync_mode_mirror_only`.
11. **Soak and approve.** Run the agreed load/message-size, restart, rebalance, rollback
    and DR exercises. Attach immutable logs/results and obtain named RB, TT, QA, product
    and operations GO approvals before Phase 1 acceptance or Phase 2 production work.

## Rollback during bootstrap or canary

On any contract, hash, ordering, count, mirror-only, lag, DLQ, reconciliation or health
failure: stop TT publication first, stop RB ingestion, keep reverse delivery and repair
off, preserve TT outbox and RB inbox/checkpoints, record the last good source watermark
and broker offsets, and do not delete/reset evidence. Diagnose, replay/reconcile through
the canonical services, and repeat the fixed-watermark procedure. Shared writes remain
paused or Phase 1 mirror-only throughout.
