## Context

Timetabler currently mutates scheduling state in several independent code paths and then performs Redis, WebSocket, or Kafka work directly. The principal engine consumer in `kafka_consumer/tt_response.py` uses Confluent Kafka defaults (auto commit and auto offset store), rejects responses older than the configured 30-second request timeout, mutates activity/resource relations without an encompassing transaction, performs Redis/WebSocket/Kafka effects before processing is durably acknowledged, and catches exceptions inside `schedule()` without propagating a processing failure. Its `swap()` response path is another direct allocation writer.

The repository also contains writers outside that handler:

- `api/views/admin/schedule_request.py::manual_constraint_break()` applies a final constraint-break schedule locally and duplicates variant/resource-map logic.
- `api/views/admin/booking_schedule.py` schedules Staff or Location independently according to a request `type`, so a complete booking can span separate calls.
- `api/views/admin/unschedule.py`, `api/views/admin/resources_update_current.py`, `api/views/admin/booking_swap.py`, and `api/views/admin/activity_delete.py` directly replace or remove allocation state.
- `api/views/admin/resources_update_requirement.py` can replace Staff and Location relations for an already-scheduled activity as a side effect of requirement edits.
- `api/views/admin/import_table.py::import_tt_activity()` can set scheduled fields, while `import_tt_booking()` directly inserts a scheduled booking and `TtActivityLocation` relation. Staff and Location imports are also direct lifecycle writers.
- `api/management/commands/20260705-manual-schedule-activity.py` and `import-manual-schedule-activity.py` directly schedule activities.
- Variant/JTA endpoints and variant merging in engine/manual/swap paths can change weeks, identities, or delete activities.
- The routed `api/views/admin/simulate_schedule_response.py` contains another full direct scheduling/variant writer, and legacy `kafka_consumer/schedule_response.py` contains a second engine response writer. Deployment usage must be proven; any obsolete path must be disabled/removed rather than left as a bypass.
- `api/views/admin/activity_template_delete.py`, `module_delete.py`, and `academic_term_delete.py` cascade-delete scheduled `TtActivity` rows through model foreign keys, while `week_pattern_delete.py` can set a scheduled activity's `week_pattern` to null and change occurrence meaning. They require pre-delete final-state capture/tombstones and atomic resource-map repair, or an explicit prohibition when scheduled allocations are affected.

`api/utils.py::recalculate_resource_map()` already opens `transaction.atomic()`. When called inside the new outer transaction, that block becomes a savepoint on the same connection and sees the outer transaction's uncommitted relation changes; this behavior is intentionally retained and tested.

The Scheduling Engine returns one complete result and does not query Timetabler SQL while BE applies it. No reviewed workflow requires an intermediate Timetabler commit. Therefore the outbound request and response wait stay outside a database transaction, the response is parsed and validated before a transaction begins, and final application can be one atomic database unit.

## Goals / Non-Goals

**Goals:**

- Commit a durable canonical outbound record in the same transaction as every authoritative Staff, Location, activity, or allocation mutation.
- Represent final replacement state, including the previous affected resource/week scope and tombstones, so moves and merges remove stale Resource Booking mirrors.
- Make engine response application atomic, idempotent, retryable, and safe under Kafka redelivery.
- Converge final allocation writers on a shared change-set application service so atomicity and outbox rules are difficult to bypass.
- Keep network calls and non-database external effects outside database transactions.
- Give operators retry, replay, dead-letter, lag, correlation, reconciliation, rollout, and rollback controls.
- Minimize Timetabler-specific knowledge of Resource Booking; Timetabler emits deterministic calendar/absolute occurrence source facts, while Resource Booking transformation, identity mapping, validation/application, downstream retries, and mirror reconciliation orchestration belong to the Resource Booking-owned adapter.

**Non-Goals:**

- Resource Booking-to-Timetabler booking commands or two-way synchronization; those are Phase 2.
- A shared reservation/conflict authority; Resource Booking mirrors remain mirror-only in this phase and conflicting Resource Booking writes against them are prohibited operationally.
- Periodic bulk synchronization as the primary integration mechanism. Reconciliation is a repair mechanism only.
- Resource Booking schema-specific payloads, recurrence expansion, or direct HTTP/adapter calls from a Timetabler database transaction.
- Treating `confirmed_schedule` as a Resource Booking acknowledgement or dependable engine handshake.

## Hard Phase Dependency

Phase 2 implementation MUST NOT begin until all Phase 1 implementation tasks are complete; the full Phase 1 unit, integration, regression, failure-injection, concurrency, and performance suites have passed for Timetabler and Resource Booking/adapter; reconciliation has demonstrated the agreed convergence criteria; and an approved completion report records versions, environments, results, known limitations, and rollback readiness. Phase 1 planning/review may proceed in parallel with Phase 2 planning, but Phase 2 production-code work is blocked.

## Decisions

### 1. Keep the integration boundary canonical and Timetabler-owned

Timetabler emits a versioned canonical envelope, not a Resource Booking payload. One finalized Timetabler change set becomes one outer transport object and one Kafka record. The Resource Booking-owned adapter transforms canonical Staff, Location, activity, and allocation members, validates/applies the supplied absolute timezone-aware occurrence facts, maps identities, calls Resource Booking, retries downstream work, suppresses duplicates, and reconciles Resource Booking state.

The canonical envelope contains at least:

- outer transaction/event ID, schema/event/source type, source scope, one positive globally commit-ordered source sequence, the source-scope ordering key, member count, `complete=true`, aware commit time, and one member-array hash;
- bounded `member_only` canonical aggregate events carrying stable Timetabler aggregate identity, monotonic aggregate version, change-set index/count, and deployment/term identity where applicable;
- `correlation_id`, `causation_id`, request/engine request ID when applicable, actor/audit identity, and `origin`;
- a final-state snapshot or stable final-state references sufficient for the adapter to construct the mirror;
- prior and final affected Staff IDs, Location IDs, week IDs, scheduled slot/time/duration, activity/variant identities, and status;
- tombstones for deleted/merged variants, activities, Staff, and Locations; and
- canonical timezone/term/week references and deterministic local-offset/UTC occurrence boundaries needed by the adapter without Timetabler knowing Resource Booking schemas.

Staff/Location IDs are stable source identities. If deployment-level uniqueness is required, the configured integration deployment/tenant ID namespaces them. Exact event transport, schema registry, and retention are resolved as contract tasks before implementation.

### 2. Add transactional outbox and applied-response receipt models

Planned models under `api/models/` and migrations under `api/migrations/`:

- `IntegrationOutbox`: immutable member envelope/payload, aggregate key/version, source scope/global change-set sequence/index/count, origin/correlation/causation/request metadata, state (`pending`, `retry`, `published`, `dead_letter`), attempts/errors/timestamps, and replay audit metadata. Unique constraints prevent duplicate aggregate versions and member indices.
- `IntegrationTransportCursor`: transactionally allocates one scope-global sequence when the first member of a change set is captured; the cursor lock remains held until commit so allocation order is commit order and sibling members reuse that sequence.
- `IntegrationPublisherState`: non-secret publisher heartbeat/status and last published sequence for preflight/liveness.
- `AppliedEngineResponse`: unique engine `request_id`, response hash/protocol version, change-set ID, applied timestamp, and outcome. A duplicate matching response is a no-op; a repeated request ID with a different hash is rejected and alerted rather than applied.

Outbox member rows and their finalized member count are inserted by the same `transaction.atomic()` block as the authoritative mutation. The exact composite envelope is size/hash validated before that block can commit. Publisher status transitions affect every member together in separate transactions. A downstream acknowledgement is never required to commit Timetabler state.

### 3. Apply a canonical final schedule change set

A new application-layer service (planned under `api/services/scheduling/`) accepts a fully parsed and validated change set containing successful final assignments, no-slot items, previous affected state, requested origin, correlation metadata, and tombstones. It provides primitives for apply, replace, unschedule/delete, merge, and resource replacement. Existing endpoints/consumers become orchestration adapters around that service.

The service runs this deterministic unit:

1. Open `transaction.atomic()` only after all external input is parsed and validated.
2. Lock affected activities, variant families, Staff, Locations, weeks/resource-map rows, and response receipt keys in deterministic model-and-ID order. Configure bounded lock and statement timeouts.
3. Insert/check the applied-response receipt when the source is an engine response.
4. Delete and insert `TtActivityStaff` and `TtActivityLocation` replacement relations.
5. Update scheduled fields, duration/week relations where requested, and unschedule/delete state.
6. Process variant/JTA merge, week transfer, deletion, and tombstones.
7. Recalculate database resource maps. The current nested atomic block is expected to act as a savepoint and see same-connection uncommitted state.
8. Insert one or more canonical TT-to-RB outbox rows containing the final committed replacement state.
9. Commit. Any exception escapes the atomic scope and rolls back all preceding database work.

No broad exception handler may remain inside the transaction or convert a failed apply into apparent success. Validation errors are handled before the transaction; processing exceptions are caught outside it for logging/retry and must remain visible to the caller/consumer.

### 4. Preserve the safe engine-response sequence

The required sequence is:

`engine request -> engine response -> parse/validate -> transaction.atomic { deterministic locks; receipt; relation replacement; scheduled fields; variants/weeks/deletions; DB resource maps; TT->RB outbox } -> COMMIT -> post-commit Redis/WebSocket/confirmed_schedule/legacy or microservice Kafka triggers -> store/commit consumed response offset`

The outbound request and engine wait are never inside the database transaction. `transaction.on_commit()` may trigger immediate best-effort effects, but any delivery that must survive a process crash uses a durable outbox. There is no synchronous Resource Booking or adapter call inside the transaction.

`confirmed_schedule` remains post-commit and feature-gated for compatibility only. The current engine integration has no dependable handler for it; it is neither an engine success guarantee nor a Resource Booking acknowledgement.

### 5. Use manual Kafka acknowledgement and idempotent redelivery

The response consumer is configured with `enable.auto.commit=false` and `enable.auto.offset.store=false`. It stores/commits an offset only after a successful database/outbox commit and after durable post-commit delivery records have been created. Best-effort UI refresh may fail without withholding the offset when a durable retry record exists.

Crash behavior is explicit:

- Before database commit: no state/outbox exists; message redelivery re-attempts the whole change.
- After database commit and before offset commit: redelivery finds the matching applied-response receipt and performs no second mutation; missing post-commit effects are re-triggered from durable records.
- After offset commit: database/outbox state already exists and publisher recovery is independent of the response topic.

The current age-only 30-second rejection is replaced by receipt/request-state logic. A response that is merely redelivered after rollback is not discarded because its original request is old. Unknown, mismatched, already-terminal, or protocol-invalid responses are quarantined/alerted according to policy; a valid un-applied response remains eligible. The exact first-arrival expiry policy is a decision task and cannot defeat rollback/redelivery safety.

### 6. Keep post-commit effects recoverable

Redis refresh, WebSocket notification, `confirmed_schedule`, legacy/microservice Kafka, HTTP calls, and RB/adapter calls never occur before the authoritative commit. Immediate triggers run after commit. Effects whose loss changes another durable system use an outbox/delivery record; UI-only effects may be best effort if the UI can refresh canonical state.

Redis or downstream failure after commit does not roll back or reapply Timetabler state. Retrying the effect uses the committed change-set ID and remains idempotent.

### 7. Deliver and reconcile complete source transactions from durable canonical state

A dedicated outbox publisher claims complete finalized change sets with safe worker locking, emits one composite Kafka record under the single source-scope partition key, preserves the global source sequence, uses exponential backoff with jitter, and transitions every member together. A failed/dead-letter sequence blocks later delivery. Operator replay resets the whole immutable transaction for audited retry under its same ID/sequence; it never creates a later replacement that can be overtaken.

The adapter owns comparison against Resource Booking. Timetabler exposes a bearer-authenticated bounded snapshot at one fixed global transaction watermark. Latest aggregate members at the fence are regrouped by their originating source sequence into the same outer transport contract, at most one hash/envelope per sequence, strictly ordered and never above the fence. Snapshot mode may include only the originating members still latest at the fence and uses a deterministic snapshot event ID to avoid collision with a different live member set; complete live delivery continues to use the source transaction ID as its event ID. Timetabler does not query or model Resource Booking tables. Reconciliation remains repair, not the primary transaction flow.

Pre-existing Staff and Location rows that predate outbox capture are repaired through a dedicated guarded management command, not a scheduling workflow or fabricated activity mutation. It takes fixed maximum-ID boundaries, locks resources in deterministic batches, compares the canonical current state with the latest finalized member, and appends only missing or stale replacements. Each emitted batch has one change-set ID, next per-aggregate versions, one globally commit-ordered source sequence, and finalization inside the same database transaction. A failed batch leaves no member, version, or sequence behind; a rerun skips identical state. Deleted source rows are not scanned, so an existing tombstone remains authoritative.

If Resource Booking projection semantics change after source version 1 has already been applied, recovery does not delete mirrors or weaken same-version hash protection. A distinct operator-approved projection-revision command re-emits every current Staff/Location state under the accepted snapshot event type at its next aggregate version. Its stable operation key is stored in `request_id` with origin `resource_projection_revision`; finalized outbox members form the per-aggregate completion ledger. Row/advisory locks serialize concurrent retries, completed aggregates are skipped even after partial multi-batch success, and a failed batch rolls back. The operation changes no Staff/Location row, never emits activities, and remains unavailable unless publication, approval, and reverse delivery are all false. A normal bootstrap afterward is a canonical-hash no-op.

### 8. Roll out with mirror-only enforcement

Rollout order:

1. Deploy additive schema and metrics with publisher disabled.
2. Deploy shared atomic writer/engine consumer changes with outbox capture enabled but external RB publishing disabled; verify no regression and bounded transaction timing.
3. Deploy and contract-test the Resource Booking-owned adapter.
4. While publication, activation approval, and reverse delivery are all disabled, explicitly bootstrap pre-existing Staff and Location canonical states through the guarded resource-only command. Seed activity/allocation state only from authoritative existing events or a separately approved source-state procedure; never fabricate activities to populate Resource Booking. Take a fixed snapshot watermark and reconcile.
5. Mark Timetabler-sourced Resource Booking records/resources mirror-only and prohibit Resource Booking from creating conflicting bookings against them.
6. Enable publisher per environment/tenant, monitor outbox lag/error/dead-letter counts and mirror divergence, and complete the Phase 1 acceptance suite.
7. Produce and approve the Phase 1 completion report before Phase 2 implementation starts.

Rollback is additive and preserves evidence: disable outbound publication, keep capturing outbox rows if the atomic writer remains deployed, stop adapter application, and retain schema/outbox/receipts for replay. Do not purge or renumber events. If reverting the atomic engine consumer, first stop response consumption and prove no ambiguous in-flight responses. Mirrored Resource Booking data remains source-tagged and read-only until a controlled adapter rollback/reconciliation completes.

## Operational and Configuration Design

Planned settings in `backend/settings.py` and `.env-sample` include integration enablement, deployment/tenant identity, publisher transport endpoint/topic, schema version, batch size, retry/backoff/dead-letter thresholds, lock/statement timeouts, consumer manual-commit/max-poll settings, reconciliation cadence, and feature flags for post-commit `confirmed_schedule`. Secrets remain in the deployment secret store.

The staging GitHub Actions workflow is the deployment enforcement point rather than an operator-only runbook. It transports the non-secret Phase 1 values and secret presence from the protected `staging` environment, stops any earlier publisher before the base deploy, atomically writes the service environment, forces phase `phase1`, schema `2`, reverse delivery false, `confirmed_schedule` false and publisher concurrency one, verifies migrations through `0102`, and either leaves the PM2 publisher stopped or starts exactly one instance after enabled preflight and liveness pass. Explicit, mutually exclusive workflow-dispatch inputs may invoke either the idempotent Staff/Location bootstrap or a stable-key projection revision only after the gate attests publication, activation approval, and reverse delivery are all false; ordinary push deployment invokes neither. Both the outer workflow and the supervised gate attempt failure rollback to publication/approval false without printing broker credentials or snapshot tokens. Empty or unattested deployment/source identities cannot be activated.

Metrics/logging/tracing include outbox pending/retry/dead-letter totals, oldest-event age, publish latency, attempt count, apply transaction duration, lock/deadlock/timeout counts, duplicate response count, request-to-apply latency, post-commit effect failures, consumer rebalances/offset lag, and reconciliation missing/stale/orphan totals. Logs carry event/change-set/request/correlation IDs without sensitive payload leakage. Alerts cover sustained lag, dead letters, repeated version gaps, receipt hash mismatches, transaction timing budget breaches, and divergence.

## Timing and Concurrency Budget

The engine currently retains temporary reservations for approximately 30 seconds after producing a result. Representative single and worst-case bulk results must measure response parse, lock wait, atomic apply, resource-map recalculation, commit, Redis refresh, and offset handling. If the agreed percentile budget cannot finish the necessary finalization inside that reservation window, the cross-service plan must add engine reservation TTL extension/renewal rather than accepting a conflict window. Kafka `max.poll.interval.ms`, batch limits, consumer pause/resume, and heartbeat behavior must be sized from the same measurements.

## Coordinated E2E Passed; Load Gate Remains

The execution contract is `docs/resource_booking/phase1-e2e-load-test-plan.md`. Its 13 E2E cases passed in coordinated run `phase1-e2e-20260807-01`, covering Staff/Location lifecycle, bulk final engine response, advice-only Pre-Schedule, accepted final drag/drop, native booking, reschedule, unschedule/cancel/delete/variant merge, manual constraint-break, duplicate/delayed delivery, a Phase 1 mirror-only concurrency race, controlled publisher/consumer restart, safe application/configuration rollback with durable resume, cleanup, and final unfiltered reconciliation. Its eight load cases remain not started: the accepted E2E-final baseline of 1,810 resources/62 activities at watermark 117 (with 1,794/46/19 retained only as fixed-watermark recovery provenance), steady `P` for 30 minutes (`P=2` by default), `5P` for five minutes, the largest real approved bulk transaction, maximum contract recurrence, the exact 2 MiB boundary, controlled restart under load, and 60 minutes of stability.

Default thresholds are end-to-end p95 ≤ 5 seconds/p99 ≤ 15 seconds, lag drain within five minutes, Resource Booking apply p95 ≤ 2 seconds/p99 ≤ 5 seconds, Timetabler database lock-wait p95 ≤ 500 ms with no unhandled deadlock, booking API p95 impact ≤ 20% with error-rate increase below 0.1 percentage point, and no sustained CPU or memory above 80%. Accepted conflicts, partial apply, unexplained drift, mirror-only violations, gaps, unresolved quarantine/dead letter, and final backlog are all zero-tolerance failures. Reverse delivery remains false.

Any future load execution is restricted to an explicitly approved non-production environment and supported APIs with collision-resistant run IDs; it must not touch production or an unrelated database and must not use direct SQL or table fixtures. Phase 2 and reverse delivery stay disabled.

The Timetabler load executor is a default-disabled, independently dispatched
profile runner. A manifest-free read-only attestation supplies the exact
service/configuration/database/home fingerprints used to approve one immutable
cross-service manifest. The run order is ENTRY, LOAD-01, PROVISION, LOAD-02–06,
ARM-FENCE, LOAD-07–08, CLEANUP, FINAL. The manifest contains a fixed predecessor graph, never future
evidence hashes; each dispatch securely verifies the actual predecessor
mode-0444 cross-user exchange artifact and digest. It requires exact service commits, default
threshold approval (`P=2`, `5P=10` unless a separate versioned telemetry
decision exists), seven named roles, an approved window, and exact source
fences. It attributes every new
source sequence to a run request, verifies canonical member/transaction hashes,
waits for complete ordered publication, samples source/publisher/host state,
and hard-aborts on correctness or resource invariants. Provision uses supported
APIs only and creates equal 32-way lifecycle/allocation shards plus 1,364
activities after LOAD-01. Rate profiles prove accepted start-time fidelity, not
queued-future counts; queued calls are cancelled on abort. LOAD-04 requires an
approved real maximum activity/member count. LOAD-05 cannot run until the
maximum recurrence contract of 32 activities × 12 weeks = 384 occurrences is
bound to its approval record. LOAD-06 calibrates the
exact under/over canonical envelopes before mutation, requires the admissible
case within 4 KiB below 2 MiB, and proves over-boundary rollback with no
source/domain/outbox partial state. LOAD-07 restarts only the exact Timetabler
publisher behind a separately written short-lived shared fence artifact and
attests 0–5 second skew. The guarded `arm-fence` action first verifies the
sealed LOAD-06 predecessor and unchanged source fence, then creates/reuses only
the canonical uid-1073 mode-0444 fence with a cryptographic nonce and real
approval reference; it performs no load, restart, database mutation, or domain
mutation. Resource Booking owns its independently coordinated consumer action.
Cleanup uses thirteen retry-safe, pre-calibrated supported-API
transactions rather than one potentially oversized delete.

Canonical load transaction timestamps always include an explicit timezone:
aware database values retain their offset, while naive values are localized
with the configured Timetabler `TIME_ZONE` before serialization. A single
guarded `correct-provision-evidence` compatibility action may repair the
already sealed offset-less PROVISION evidence through the shared manifest's
strict `evidence_corrections.provision` declaration. It verifies and preserves
the original immutable artifact, exact source/publisher fence, and complete
run-owned inventory, writes only the declared `provision-corrected.json`,
`provision-corrected-v2.json`, `provision-corrected-v3.json`,
`provision-corrected-v4.json`, `provision-corrected-v5.json`,
`provision-corrected-v6.json`, or `provision-corrected-v7.json` immutable
artifact, and has
no API authentication, restart, domain/database/outbox mutation, load, DR, or
reverse-delivery capability. LOAD-02 accepts that file only as the declared
`PROVISION` predecessor with the supplied exact digest.

LOAD-02 recovery from an interrupted run is bounded to the open interval
between its immutable entry and final fences. Timetabler must reconstruct an
exact deterministic request prefix from finalized, published durable outbox
transactions, reject every gap/unrelated request/incomplete transaction, and
submit only the remaining deterministic iterations. Final evidence retains
all 3,600 source transactions, while active API latency, lock, host, and rate
measurements truthfully exclude the historical prefix.

The fixed recovery has a distinct non-mutating `checkpoint-load-02` action.
It binds the exact run, current commits, corrected PROVISION predecessor, open
LOAD-02 fence, and strict durable prefix proof into immutable shared
`load-02-checkpoint.json`. The artifact carries every prefix source
transaction plus canonical digest and explicit negative capability assertions;
it never authenticates to an API or writes/restarts either service.

Corrected PROVISION evidence may be regenerated after LOAD-02 begins only when
the identical strict prefix proof succeeds. Its PROVISION transaction inventory
stays historical; current source progress is disclosed separately and cannot
be mistaken for PROVISION membership.

The first LOAD-02 execution demonstrated that client HTTP completion is not a
Scheduling Engine completion receipt: `schedule-request` transactionally
enqueues a `PostCommitDelivery`, the delivery UUID becomes the engine/Kafka
request ID, and the engine result may arrive after a later per-shard API call.
Immediate unschedule can therefore be a valid no-op while the activity is still
unscheduled. The failed run is not resumable from client-request/outbox counts.
Before any cleanup, a fixed-run read-only containment audit must correlate the
API-log aggregate, post-commit correlation and generated UUID, Kafka response,
applied receipt or quarantine, finalized/published change set, and run-owned
activity state across two stable reads. Missing or contradictory correlation is
reported as ambiguity, never relabelled as a source transaction. Any unresolved
delivery/response/receipt chain keeps late-mutation risk true; quarantine keeps
operator-replay risk explicit. The audit has no authentication, evidence-write,
domain/database/outbox/Kafka/process/load/cleanup/DR/Phase-2/reverse capability.

Any subsequent LOAD-02 execution is engine-aware and non-resumable. Timed load
contains only recurring term-26 schedule and unschedule transactions. A
single-worker shard may issue its dependent unschedule only
after the preceding `schedule-request` has one published post-commit delivery,
one valid terminal engine response, one matching atomic receipt, and exactly
one finalized/published source transaction. Source attribution uses the
engine delivery UUID for engine-applied transactions and the client request ID
for synchronous API transactions. Cross-shard completion order may differ from
HTTP submission order, but the complete transaction/request bijection and the
exact stage fence remain mandatory. Any partial execution is audited and
cleaned as an abandoned run; it is never replayed from an inferred prefix.
Each transaction also seals term dates, recurrence cardinality, exact
Staff/Location preset cardinality, and the assigned/released Timetabler state.
Resource Booking acceptance requires an audited occurrence for every recurrence
member with exactly one Staff and one Location allocation, allocated on schedule
and released on unschedule. This evidence change does not alter the
Timetabler--Scheduling Engine contract.

The separately approved Jan--June 2026 recovery is an existing-source-state
projection backfill, not an engine schedule or a correction of the observed
45-vs-46 term-count mismatch. A read-only plan binds term 26, IDs 672--716,
the source/database/configuration fence, current scheduled states, truthful
typed allocations (including Location-only activities 686 and 690), absolute
occurrences, next aggregate versions, and the exact-size complete transport
preview. Execution revalidates that plan while locking the source rows and
global cursor, then uses the ordinary activity outbox capture service to append
and finalize one deterministic 45-member transaction. It performs no activity,
allocation, week, Staff, Location, engine, RB, reverse-delivery, Phase 2, load,
or DR action. Finalization remains below the configured message limit or the
entire outbox/version/cursor write rolls back. Retry accepts only the identical
sealed change set and final fence; publisher verification is required before
handoff to the independently owned RB receipt/reconciliation gate.

Both services consume one canonical manifest unchanged. Entry requires named
operator, abort authority, Timetabler owner, Resource Booking owner, QA owner,
product owner, and operations owner records with role/reference/approval time,
plus an explicit ISO-8601 execution window. A runner fails closed outside that
window or when the complete profile observation cannot finish within it. The
manifest fixes deterministic workload/observation windows for all profiles,
including PROVISION `5400+300` within a 120-minute workflow timeout, LOAD-02 `1800+300`, LOAD-03 `300+300`, bounded LOAD-04/05/06
`300+300`, LOAD-07 `300+300` around its shared epoch/hash, and LOAD-08
`3600+300` seconds.

Each canonical profile result is retained as a Timetabler-owned mode-0444 file
at the traversal-safe deterministic staging path
`/var/tmp/mayvins-timetabler-phase1-load/<run-id>/<profile-lower>.json`. The
fixed root/run directories are uid 1073 mode 0755 and files are uid 1073 mode
0444 for the attested Resource Booking reader uid 1069. The manifest and
read-only attestation bind the path hash, UIDs, modes, and prohibition on
credentials. It contains the
complete transaction list and raw source/host/publisher measurements, while
Actions emits only its SHA-256 and bounded locator metadata. Resource Booking
may read that exact file on the attested shared host but receives no authority
to touch Timetabler data/processes. Timetabler never receives RB credentials or
controls an RB process. Cleanup/final preserve the file for coordinated
verification and require a separate RB-owned unfiltered reconciliation result.

Disaster recovery is explicitly outside this gate: database backup/PITR restore, Kafka cluster rebuild/restore, host/AZ/region/site failover, infrastructure/DNS/secrets-vault rebuild, and multi-region/business-continuity exercises are neither required nor authorized. Controlled process restart, safe application/release/configuration rollback, checkpoint recovery, durable resume, and reconciliation remain mandatory operational tests and are not DR.

## Risks / Trade-offs

- A shared change-set service touches many mature write paths. Staged migration, path-by-path characterization tests, and feature flags reduce but do not eliminate regression risk.
- Larger atomic engine batches may increase lock duration and deadlock pressure. Deterministic ordering, timeouts, bounded batches based on measurements, and concurrency tests are mandatory.
- Publisher success does not prove Resource Booking application unless the chosen contract includes a durable adapter acknowledgement. Versioned reconciliation remains necessary.
- Capturing previous and final affected state increases outbox volume but is required for correct replacement and tombstones.
- Disabling Kafka auto commit exposes older consumer assumptions. Receipt idempotency and crash-window tests are prerequisites to rollout.
- The current 30-second age filter cannot remain the sole expiry control; changing it requires explicit quarantine and observability for truly late/unknown responses.
- Atomic application is engine-protocol compatible only if all safeguards described above are implemented and tested; there is no zero-risk claim.

## Acceptance Criteria

- Every Staff and Location create/update/archive/delete path commits exactly one appropriately versioned canonical lifecycle change with the mutation or commits neither.
- Every final allocation writer, including normal/bulk/legacy/simulated engine paths, final drag/drop, manual constraint break, booking, unschedule/delete, current/requirement resource replacement, variants/JTA, imports, cascade/pattern deletes, and scheduling management commands, uses the atomic/outbox guarantees or is formally blocked/decommissioned.
- Pre-schedule processing emits no allocation event; accepting the final drop emits exactly one committed replacement change set.
- Moved/changed/deleted allocations remove stale old occurrences through replacement scope/tombstones rather than only adding new state.
- Failure injection after every database stage produces total rollback, no outbox, and no external effects.
- Duplicate/redelivered engine responses apply once, and the three required crash windows recover without duplicate state.
- Post-commit Redis/WebSocket/downstream failures retry or refresh without rolling back or double-applying database state.
- Concurrency/deadlock, large-batch, 30-second engine reservation, and Kafka max-poll tests meet recorded budgets.
- Replay, ordering/versioning, reconciliation, dashboards, alerts, and operator runbooks pass acceptance exercises.
- Resource Booking mirrors are demonstrably mirror-only and reconciled, and the approved cross-service Phase 1 completion report blocks/unblocks Phase 2 explicitly.
