Skip to content

ADR-021: Messaging after MassTransit v8

This ADR is for planning only: no work starts until Chris answers the open decisions in §13. It answers #3986.

SyRF runs MassTransit 8.4.0 across five deployables, and open-source v8 maintenance ends after 2026.

Two owner decisions are already recorded:

  • 2026-10-03: licensing v9 is rejected.
  • 2026-10-04: SyRF moves to Wolverine, in one coordinated switch-over with no side-by-side running. Extensive isolated testing has to give enough confidence before the production switch.

Until the switch, SyRF runs a bridge: the last open-source release, 8.5.11, under an explicit risk window.

Status

In-Review, 2026-10-04. Two decisions are recorded:

  • v9 rejected (2026-10-03);
  • Wolverine with a single switch-over chosen (2026-10-04).

The decisions still open are listed in §13. Evidence markers:

  • V: VERIFIED by reading the code on main at 85e6facf7.
  • C: CORRECTED, where the earlier text or the review appendix was wrong.
  • NV: NOT VERIFIED, with the reason.
  • W: a web fact, with a link and the date checked.

1. Context

1.1 External facts (W, checked 2026-10-03/04)

Fact Source
v9 is commercial (Massient, Inc., 9.2.3). Rejected by owner decision 2026-10-03. nuget, massient.com
v8 stays Apache-2.0 ("MT v8 remains open-source", Chris Patterson). 8.5.11 was published 2026-09-30. It depends on RabbitMQ.Client 7.2.2, MongoDB.Driver 3.12.0, Quartz [3.22, 4.0) and EF Core Relational 10.0.0, and its nuspec has a net10.0 group. X post; NuGet API
The end of v8 maintenance ("fixes through 2026") is stated only by third parties. No primary Massient wording was found, and Massient's licence says only that v8 is unsupported under it. Jovanovic, licence
.NET 10 LTS support ends 2028-11-14. .NET 11 (STS) is at RC1, GA expected November 2026. .NET policy
Wolverine 6.45.0 (2026-10-02): MIT, net9/net10. Packages for RabbitMQ, SqlServer and EntityFrameworkCore. There is no official MongoDB store; the community Wolverine.MongoDB 1.1.0 has a single owner. Wolverine uses Microsoft DI. JasperFx support plans cost $3k–15k a year and are optional. nuget, sagas, support
Wolverine recurring schedules: opts.Schedules.ScheduleRecurring takes cron with a time zone. ExclusiveNodeWithParallelism(n) runs a listener on one node, n in parallel, with failover. recurring, exclusive node
Wolverine scheduling (ScheduleAsync) is durable only with database persistence. No API to cancel a scheduled message is documented. message bus
Wolverine testing: tracked sessions (TrackActivity, InvokeMessageAndWaitAsync), StubAllExternalTransports, StubWolverineMessageHandling. testing
Rebus 8.9.5 (MIT). Corrections against the previous version of this ADR (C):
• Rebus.RabbitMq 10.1.1 (2026-01-19) ships net8.0/net9.0/netstandard2.0, which run on net10, so the missing net10 build is not a blocker.
• Rebus.MongoDb 9.0.0 (2024-12-18) depends on MongoDB.Driver ≥ 3.0.0, which is compatible with SyRF's 3.10.0.
• Request/reply comes from the Rebus.Async add-on (10.0.0, 2023-11-15).
NuGet API (Rebus)
Brighter (Paramore.Brighter) 10.7.0 (2026-07-29): net8/9/10. Has a MongoDB outbox/inbox (driver 3.10.0), Quartz/Hangfire schedulers and RPC. No sagas. nuget
CAP (DotNetCore.CAP) 10.0.2 (2026-08-01): net8/9/10. Has MongoDB storage (driver 3.9.0) and delayed publish. No sagas and no request/reply. nuget
NServiceBus 10.2.9 (net10, proprietary) has official MongoDB persistence. It costs from $10 per production endpoint per day; Community is free up to 3 endpoints and 10k msg/day. pricing
OpenTransit (re-checked 2026-10-03): last push 2026-08-12 (an upstream merge). No releases, no NuGet packages and no named maintainers. GitHub

Interop footnote: Wolverine can talk to MassTransit over RabbitMQ (UseMassTransitInterop), but with limits: one DefaultIncomingMessage<T> per listener, and warned-of reply-queue "hiccups" (interop). The single switch-over makes interop irrelevant, so it is not a scoring criterion.

1.2 What SyRF uses: inventory (V unless marked)

ID Item Count / shape Evidence
I1 Bus bootstrap and naming One helper builds every in-cluster bus: AddConsumers(entryAssembly), the publish scheduler and ConfigureEndpoints. MapCommandQueues maps each ICommand to a queue; three bulk-PDF commands share one queue. MassTransitHelpers.cs:30-62,174-202; IntegrationExtensions.cs:20-56
I2 Consumers 31 IConsumer + 4 IJobConsumer: PM 28 + 4, API 2, the PDF agent 1. One Fault<T> consumer, about 14 definitions with ConcurrentMessageLimit, and 2 per-pod temporary fan-out endpoints in the API. PM.Endpoint/Consumers/*.cs; PM Program.cs:484-582; ActivityClaimRevokedConsumer.cs:24,106; ProjectStatisticsChangedConsumer.cs:24,122; SearchUploadSavedFaultConsumer.cs:14
I3 Request/response 10 clients: API ×8, PDF agent ×1, PM ×1. BulkPdfNotifierAuthorityController.cs:17-23; StudyController.cs:40; SearchController.cs:249,270; PdfAgentMassTransitRegistration.cs:25; PM Program.cs:581
I4 Saga (MongoDB) SearchImportJobStateMachine (325 lines): 5 events (one is Fault<…>), a 60-minute timeout and partitioner(1), stored in pmSearchImportJobState. Production (read-only aggregate, 2026-10-03) holds 4 Completed, 1 Error, 65 Uploading and 572 Parsing, mostly orphaned. SearchImportJobStateMachine.cs:15-84,282,298-320; SearchImportJobState.cs:13-38; MongoMassTransitExtensions.cs:12-27
I5 Job Service sagas (SQL Server, EF) JobSaga, JobTypeSaga and JobAttemptSaga live in the Quartz service. Four jobs set JobOptions (timeouts 5 min to 2 h, retries, SetConcurrentJobLimit(3)). Jobs are submitted by the API and the Lambda. QuartzServiceCollectionExtensions.cs:131-147; Migrations/20250324010206_InitialCreate.cs; ReferenceFileParseJobConsumer.cs:68-90; RoBProcessingJobConsumer.cs:53-67; S3FileReceivedFunction.cs:252
I6 Scheduling (Quartz, SQL Server) 9 SchedulePublish + 4 CancelScheduledPublish, with tokens stored in MongoDB. Only the idle timer has a generation guard: the suspended-session and liveness timers rely on the broker cancel alone. 5 recurring schedules (FEAT-024). UseScheduledRedelivery on 5 definitions. The PDF agent's RetryLater uses the shared Quartz. NotificationHub.cs:281,742,902,1551,1571,1590,1598; ReviewController.cs:357,1548; SlotReservation.cs:212-232; MarkSessionIdleConsumer.cs:60-64; PM Services/*Schedule.cs
I7 Outbox No MassTransit transactional outbox, only UseInMemoryOutbox. SyRF's own outboxes (MongoDB/DynamoDB) do not depend on the transport. In the persistence plan's IDomainEventOutbox, only IDomainEventSink touches the bus, so it is compatible. MassTransitHelpers.cs:128,146,158; the October 2026 persistence and messaging plan (IDomainEventOutbox; temporary planning doc)
I8 Retry No global policy. 15 per-consumer retry, redelivery or outbox sites, plus JobOptions.SetRetry (#3984). ReviewSessionConsumerRetry.cs:14-31
I9 Topology Exchanges are Namespace:Type, with interface-hierarchy exchange-to-exchange bindings. Queues are <Assembly − .Messages>.<Type> and <Assembly − .Endpoint>.<Type>; 5 use the default formatter. MassTransit infrastructure names (quartz, Job*, MassTransit.*). The B1 permission regexes are built from these names. broker-security-contract.md:189-215,235-262
I10 Contracts 35 files, 46 interfaces, sent through 68 interface-typed anonymous initialisers (a MassTransit-only feature). PM.Messages → PM.Core. Serialiser: System.Text.Json. PM.Messages.csproj:9; grep counts
I11 S3-notifier Lambda Creates a bus per S3 record. C: the vhost comes from each object's metadata, so any cached sender must be keyed by vhost. The reconciler Lambda caches one bus. Production runs older notifier code (#3669). S3FileReceivedFunction.cs:56-60,202-227,507-527; BulkPdfScheduledReconciler.cs:1139-1191
I12 PDF agent (ARRNC) One consumer and a claim client, with a Paused mode (no receive endpoint). It holds no database credentials. PdfAgentMassTransitRegistration.cs:14-40
I13 Tests 60 files import MassTransit; 14 use ITestHarness; 5 use Testcontainers RabbitMQ. The e2e/ stack runs RabbitMQ and SQL Server. StateMachineTestFixture.cs:28-50; BulkStudyUpdateAtomicBrokerTests.cs:51,223-256
I14 Broker TLS (B1) RabbitMqTls makes certificate checks strict on 5671 (syrf#3878). The live B1 PRs are held until main can be deployed to production (Chris, 2026-10-02). RabbitMqTls.cs:21-74
I15 Packages and DI 13 csproj files reference MassTransit, with no central package management. SharedKernel's MassTransit reference is unused (V). DI is Lamar 15.0.1. 8.x already runs on net10.0 in production. SyrfHelpers.cs:19; Directory.Build.props:8

NV:

  • Production SQL counts for Quartz triggers and job sagas: the read was refused by the session's permission classifier (D9).
  • Daily message volume.
  • Whether open PRs touch these files: gh pr list was refused.

2. Scope

ID Problem Impact Evidence Issue
S1 v8 maintenance ends after 2026 and v9 is rejected Unpatched messaging from 2027 on the Internet-facing broker path (Lambda, ARRNC) §1.1 #3986
S2 SyRF is on 8.4.0, not 8.5.11 It misses free fixes csproj ×13 #3986
S3 Many MassTransit-only features are in use (Job Service, cancel tokens, interface initialisers, Fault<T>) The switch touches every programme I2–I10 #3986
S4 Contracts depend on the domain Contract changes cost more PM.Messages.csproj:9 #3988

3. Options

Effort is in engineer-weeks and is a PROPOSAL: S is under 1, M 1–3, L 3–8, XL over 8. "Built in" means the library provides the feature, so SyRF does not own that code.

Inventory need Wolverine (chosen) Rebus Brighter CAP NServiceBus In-house (RabbitMQ.Client) OpenTransit
I4 saga + 60-min timeout built in (EF/SQL); SyRF chooses Mongo state (D3b) built in (Mongo) none none built in (Mongo) build as v8, but nothing shipped
I5 job timeout/retry/limit 3 handlers + ExclusiveNodeWithParallelism(3) + error policies handlers; no cluster-wide limit handlers; no limit handlers; no limit handlers build as v8
I6 durable delay built in (SQL store) built in (Mongo timeouts) Quartz/Hangfire built in built in build as v8
I6 recurring (5) cron built in none via Quartz none none build as v8
I6 cancel tokens none documented → generation guards none via Quartz none none build as v8
I3 request/reply (10) built in Rebus.Async add-on RPC none built in (callbacks) build as v8
I7/I8 inbox, outbox, retry built in built in built in built in built in build as v8
I13 test harness tracked sessions basic basic basic good build as v8
Store for the above SQL Server (already run) MongoDB (could retire SQL Server) MongoDB + Quartz SQL MongoDB MongoDB n/a SQL + Mongo
Licence / steward MIT; JasperFx (also maintains Lamar); releases 2026-09-30 and 2026-10-02 MIT; one main maintainer MIT; community MIT; community proprietary; about $11k–15k a year (PROPOSAL) SyRF forever none yet
Effort (PROPOSAL) ~20–28 wk total (§8) ~22–30 wk +saga build +saga and RPC build ~18–26 wk + licence 30+ wk 0, but nothing exists

3.1 Why Wolverine

It covers the most of SyRF's inventory with built-in features, so SyRF owns the least code.

  • Scheduling: cron recurring schedules (the 5 FEAT-024 schedules) and durable delayed scheduling.
  • Jobs: ExclusiveNodeWithParallelism for RoB's cluster-wide limit of 3.
  • Reliability: inbox, outbox and retry policies.
  • State: EF/SQL Server sagas.
  • Messaging patterns: request/reply.
  • Testing: tracked sessions.
  • Stewardship: an active, company-backed steward (JasperFx, which also maintains Lamar, SyRF's DI).
  • It runs on the SQL Server SyRF already operates for Quartz and the job sagas, so no new kind of store is needed.

3.2 Why the others lost

  • Rebus has one real advantage: everything lives in MongoDB, so SQL Server could be retired. It lost because:
  • it has no recurring cron (I6), so SyRF would build a leader-elected scheduler;
  • it has no cluster-wide concurrency limit (I5);
  • request/reply needs a 2023 add-on (I3);
  • the Mongo store is unchanged since 2024-12;
  • it has a single main maintainer.

Each gap is SyRF-owned code, and together they outweigh retiring SQL Server. - Brighter and CAP have no sagas, and CAP has no request/reply. - NServiceBus is a recurring proprietary licence, the same objection as v9. - In-house would make SyRF own everything. - OpenTransit has shipped nothing; it is re-checked under trigger T6.

4. Decision

  1. Bridge (P1): upgrade 8.4.0 → 8.5.11 under Directory.Packages.props, and drop the unused SharedKernel reference.
  2. Risk window (§5): run 8.5.11 unsupported from 2027-01-01 until the production switch. The target is the switch in Q3 2027 (PROPOSAL); the hard limit is 2028-11-14.
  3. Target: Wolverine, single coordinated switch-over of all five deployables (§7 S0–S5).
  4. Saga (D3b, recommended): SyRF-owned state in Mongo, using the existing repository, Audit.Version concurrency and Wolverine handlers, plus a durable scheduled timeout.
  5. Job Service: handlers with SyRF job records. Timeouts become cancellation tokens; retries come from Wolverine error policies; the cluster-wide limit is ExclusiveNodeWithParallelism(3).
  6. Scheduling: a Wolverine SQL Server message store on the existing instance (new database). Generation guards replace broker cancels. The PDF agent stays credential-free: its RetryLater becomes a command to PM.
  7. Host switch on main, not a long-lived branch (D1-new, recommended). A startup setting Messaging:Host = MassTransit | Wolverine defaults to MassTransit, and both stacks live on main until S5.
  8. Why: main is very active, with five programmes adding consumers. A long-lived branch would drift on all ~35 consumers and the contracts, and the drift would only surface at merge, right before the switch.
  9. Cost: both packages ship in images until S5. Business logic is extracted into host-neutral handler classes, with thin MassTransit consumers and thin Wolverine handlers calling them.
  10. Guard: all five deployables must use the same value. It is one environment-level value in cluster-gitops; the Lambda and the agent get it through their own config. A test refuses a mixed configuration in any chart render, and §7 S4 checks it.
  11. Independent work: contract decoupling (#3988), STJ round-trip tests (#3984), and the persistence-plan outbox. The outbox is compatible: only IDomainEventSink is retargeted in S1.

5. Risk window for unsupported 8.5.11

5.1 What actually ends

  • Nothing stops working. v8 has no licence key or expiry, and its packages stay on NuGet under Apache-2.0 (W).
  • What ends is upstream fixes: no CVE patches and no fixes for future .NET, driver or broker changes. The "through 2026" end comes only from third parties (W). 8.5.11 shipped on 2026-09-30, so re-check NuGet monthly.

5.2 Mitigations

# Mitigation Owner
R1 Pin MassTransit 8.5.11 and its transitive dependencies (RabbitMQ.Client 7.2.2, MongoDB.Driver ≥3.12, Quartz 3.22+, EF Core 10) centrally. Any major bump needs review against this ADR. P1
R2 Monthly check of NuGet, GitHub advisories and Dependabot for MassTransit, RabbitMQ.Client, Quartz and MongoDB.Driver, with the result recorded on #3986. maintainer
R3 Finish B1 (strict 5671, per-principal permissions, close public 5672) before 2027-03-31 (PROPOSAL). auth session
R4 Fork-and-patch fallback: a vendored 8.5.11 source build, used only to apply a security fix (Apache-2.0 allows it). on trigger

5.3 Triggers

# Trigger Response
T1 A CVE with CVSS ≥ 7.0 (PROPOSAL) in MassTransit or a transitive dependency, with no fixed 8.x R4 within 14 days (PROPOSAL)
T2 .NET 11/12 needed before 2028-11-14 and incompatible Stay on .NET 10 LTS; pull the switch forward
T3 RabbitMQ.Client 8.x or MongoDB.Driver 4.x forced (security, Atlas, broker) and 8.5.11 cannot use it Pull the switch forward; R4 meanwhile
T4 Quartz 4.x or EF Core 11 incompatible with MassTransit.Quartz or MassTransit.EntityFrameworkCore Hold the versions; pull forward
T5 A RabbitMQ server upgrade breaks the 8.x client Hold the broker version; pull forward
T6 OpenTransit ships a maintained net10 release Note it on #3986; the decision stands unless Chris re-opens it

6. MVP boundary and flags

  • MVP: P1 + S0. That is 8.5.11 on every deployable, proven in staging, plus the characterisation baseline and the generation guards, which ship to production early and on their own.
  • Out of scope here, and where each is tracked:
  • global retry policy: #3984;
  • contract decoupling: #3988;
  • B1 live steps: the auth session;
  • liveness probes: review item 14.
  • Flags:
  • P1 is not flagged; roll back by image tag.
  • S0's generation guards are not flagged: they are additive, stale-delivery no-ops, and they are tested.
  • The migration is flagged through the startup host switch (§4 item 7), not a runtime flag. The default is MassTransit, so it is changed per environment at the switch, and rollback means switching back.

7. Programme: single switch-over

Every PR meets these criteria:

# Condition → result Verification
C1 Any messaging package version → resolved from Directory.Packages.props only CI package listing (one project at a time)
C2 A PR touches a handler → its test file is updated, and the S0 suite passes on the default host unit/Testcontainers
C3 Topology differs from the S0 snapshot → a diff is attached and the B1 regexes are regenerated snapshot + B1 fixture
C4 Docs → this ADR and .claude/rules are updated in the same PR docs
C5 Production → never promoted by a PR. Production is a separate Chris-approved step (S4). governance

Effort is a PROPOSAL. E2E runs only in the hermetic e2e/ stack; staging is for human testing.

P1: Bridge to 8.5.11 (S–M, 1–2 wk)

  • Files: the 13 MassTransit csproj files; a new Directory.Packages.props; SyRF.SharedKernel.csproj.
  • Approach: an EF model check on JobServiceSagaOverrideDbContext.
  • Rollback: revert the image tag.
# Condition → result Verification
P1.1 Build → every MassTransit package resolves to 8.5.11; no 8.4.0 is left transitively CI package listing
P1.2 Job Service EF model → no diff, or a reviewed migration plus a rollback script EF tooling + Testcontainers SQL
P1.3 Saga, scheduler and broker suites → pass Testcontainers
P1.4 Hermetic stack: import, bulk update, presence idle/suspend → each completes E2E
P1.5 Staging → an import reaches Completed; no _error growth for 24 h (PROPOSAL) staging check (human)
P1.6 Strict TLS → an untrusted chain is rejected fixture

S0: Characterisation baseline on MassTransit (L, 3–4 wk)

  • Files:
  • a new SyRF.Messaging.Characterisation.Tests, written against a host-neutral driver that can run unchanged against either host;
  • generation guards for RemoveSuspendedSession and CheckConnectionLiveness, matching SlotReservation.cs:218-224 and MarkSessionIdleConsumer.cs:60-64.
  • Rollback: revert. The tests are additive, and the guards make stale deliveries no-ops.
# Condition → result Verification
S0.1 Topology and queue inventory (exchanges, queues, bindings, the message types per queue) → snapshot committed, and it fails on drift Testcontainers RabbitMQ
S0.2 Each of the 46 contracts → round-trips, with golden JSON contract
S0.3 Saga: every path (started, saved, parsed, completed, each fault) plus the 60-min timeout → golden terminal states and the SearchImportJobState BSON shape (CSUUID) Testcontainers
S0.4 Each of the 4 jobs → timeout, retry count, ignored exceptions and concurrency limit pinned Testcontainers
S0.5 Every scheduled message (9 sites) and recurring schedule (5) → fires once, at the expected time Testcontainers SQL
S0.6 Each request/reply client (10) → reply and timeout behaviour pinned Testcontainers
S0.7 Bulk-PDF claim/progress/finalize → serialised, single writer Testcontainers
S0.8 Lambda: records for two vhosts in one invocation → each reaches its own vhost; malformed-record isolation pinned LocalStack + Testcontainers
S0.9 Agent Paused mode → connects with no receive endpoint Testcontainers
S0.10 A stale suspended/liveness timer after a reconnect → a no-op without a cancel (ships to production early, Chris-approved) unit + Testcontainers
S0.11 Interface inheritance: every consumer of a base or shared interface → listed with its exchange-to-exchange binding, and a test asserts it receives (Wolverine routes by concrete type) Testcontainers
S0.12 The API's per-pod temporary endpoints → a test pins fan-out to every pod and auto-delete on shutdown Testcontainers

S1: Build Wolverine behind the host switch on main (XL, 10–14 wk)

  • Scope: all five deployables.
  • Thin Wolverine handlers call host-neutral logic.
  • Concrete record contracts implement the 46 interfaces; MassTransit can still publish them as the interface.
  • Explicit routes keep queue names where possible.
  • Fault<T> paths become explicit failure events.
  • The 15 retry sites become Wolverine error policies.
  • Persistence: a SQL Server message store, with API and PM connected; the saga as SyRF Mongo state (D3b); jobs as handlers with SyRF job records and ExclusiveNodeWithParallelism(3); the 5 recurring schedules on ScheduleRecurring.
  • Client paths: the Lambda is send-only, with a sender cache per vhost; the agent's RetryLater is routed via PM; strict TLS is ported from RabbitMqTls.
  • B1: permission regexes regenerated from the Wolverine topology, including reply queues (pinned names; automatic reply queues need queue-declare permission).
  • IDomainEventSink: retargeted.
  • Rollback: none needed. The default host stays MassTransit.
# Condition → result Verification
S1.1 Messaging:Host=Wolverine on each deployable → it starts under Lamar and registers every handler, route and schedule unit + Testcontainers
S1.2 Mixed host values across charts → the render test fails chart test
S1.3 Default host → the S0 suite still green (no regression on main) CI
S1.4 Wolverine topology → a reviewed diff from S0.1, with the B1 regexes regenerated (reply queues included) and the auth session's sign-off snapshot + review
S1.5 The Lambda package on Wolverine → builds and runs in LocalStack LocalStack

S2: Confidence gates in isolated environments only (L, 3–4 wk)

All gates run in the hermetic e2e/ stack and Testcontainers, never against staging.

# Condition → result Verification
S2.1 The full S0 suite on Wolverine → green, unchanged Testcontainers
S2.2 Full E2E on Wolverine → green, no new failures against the known main baseline E2E
S2.3 Soak on Bramble: 24 h, ≥ 2× the observed peak daily volume (volume NV, measure first) → every _error queue = 0; p95 handling latency ≤ 1.2× the MassTransit baseline from the same run (PROPOSAL) soak harness (Bramble)
S2.4 Broker restart mid-traffic → no message lost; consumers reconnect within 60 s (PROPOSAL) chaos test
S2.5 A pod killed mid-job (RoB, bulk update) → the job resumes or retries per S0.4; ADR-020 rollback semantics hold chaos test
S2.6 SQL Server outage of 5 min (PROPOSAL) → scheduled messages fire late but exactly once, and recurring schedules resume chaos test
S2.7 Lambda cold start → no worse than 8.x + 20% (PROPOSAL) LocalStack
S2.8 Strict TLS on 5671 from every client → an untrusted chain is rejected fixture
S2.9 B1 permission fixture (positive and negative sets, reply queues) → passes on the Wolverine topology broker fixture
S2.10 Cut-over rehearsal on a production-shaped snapshot (the staging copy, or anonymised data; no production reads without Chris) → the drain runbook and the rollback runbook each complete end to end, timed: drain ≤ 3 h, rollback ≤ 1 h (PROPOSAL) rehearsal record
S2.11 Saga-state migration script on the snapshot → counts per state match; CSUUID ids are kept; the orphan prune runs as a dry run with an export script + Testcontainers

S3: Staging switch-over and bake (M, 2–3 wk elapsed)

Staging may be down or reseeded.

  • Bake: 2 weeks (PROPOSAL) of normal tester use. Every scheduled and recurring path must be seen firing: daily statistics, the fold sweep, presence timers and the saga timeout.
  • Rollback drill: run the rollback runbook on staging, then switch forward again.
# Condition → result Verification
S3.1 Staging switched by the S2.10 runbook → finishes within the rehearsed time staging check (human)
S3.2 Bake period → no _error growth; every path in the bake list observed firing staging check + logs
S3.3 Rollback drill → MassTransit restored within ≤ 1 h; then forward again staging check

S4: Production switch-over (M, ~1 wk including hypercare)

A separate Chris-approved step with a go/no-go checklist:

  • S2 and S3 green;
  • B1 status known;
  • 3669 resolved;

  • D9 counts in hand.

Maintenance window (D4):

  1. Stop producers: API admission, the Lambda trigger and agent intake.
  2. Drain every queue.
  3. Wait out running jobs (≤ 2 h) and the saga timeout (60 min).
  4. Run the approved orphan prune with export (D5).
  5. Migrate live saga rows.
  6. Re-create active session timers and register recurring schedules on the new store.
  7. Deploy in coordinated order: PM, then Quartz/SQL store, then API, then the Lambda (ACK, with the CodeSha256 check), then the ARRNC agent.
  8. Verify, then reopen.

Rollback (time-boxed, PROPOSAL 4 h from reopen; roll forward only after that):

  1. Stop producers.
  2. Drain the Wolverine queues.
  3. Redeploy all previous images and the Lambda version.
  4. Restore the saga export.
  5. Quartz and the job-saga tables are untouched until S5.
# Condition → result Verification
S4.1 Before the window → _error counts, saga-state counts and job-saga counts recorded Chris-run read-only check
S4.2 Drain step → every queue reads 0, and no running Job* saga is left runbook record
S4.3 After the switch → all five deployables report Wolverine; first import, bulk update, RoB, presence and recurring each succeed post-switch checklist
S4.4 First 24 h → no _error growth (PROPOSAL); a rollback trigger means S4 rollback governance

S5: Remove MassTransit and retire Quartz (S–M, 1–2 wk, after ≥ 4 weeks stable, PROPOSAL)

# Condition → result Verification
S5.1 No MassTransit package and no host switch left in the solution CI package listing
S5.2 SyRF.Quartz and the job-saga tables → retired after a backup; the B1 regexes have no MassTransit alternatives governance + fixture
S5.3 Full hermetic E2E → green E2E

8. Order, critical path and timeline

P1 ─► S0 ─► S1 ─► S2 ─► S3 ─► S4 (Chris go) ─► S5
      #3984 / #3988 / persistence-plan sink: parallel, separate PRs
  • Order: S0's generation guards can ship before the rest of S0. S1 can start on contracts and handler extraction while S0 finishes, because they touch different files. S2 needs all of S1.
  • Effort (PROPOSAL): P1 1–2 + S0 3–4 + S1 10–14 + S2 3–4 + S3 2–3 + S4 1 + S5 1–2, which is about 21–30 engineer-weeks. Without side-by-side running, testing replaces incremental proof, so this is higher than the earlier phased estimate.

Timeline (PROPOSAL):

Window Work
Oct–Nov 2026 P1
Nov 2026 – Jan 2027 S0
Jan–May 2027 S1
May–Jun 2027 S2
Jul 2027 S3
Aug 2027 S4
Sep–Oct 2027 S5

The same people also carry the auth cutover, FEAT-024, Bulk PDF production, AF2 and the notifications stack. The switch happens after December 2026, so the §5 risk window is required.

9. Risks

Risk Mitigation
A single release spans 5 deployables, including the off-cluster ARRNC agent and an ACK-managed Lambda One environment-level host value; S1.2 mixed-config test; coordinated deploy order; agent and Lambda steps in the rehearsed runbook
Rollback is heavier than per-endpoint flags (drain, then redeploy everything, then restore state) Rehearsed and timed in S2.10 and S3.3; 4 h time-box; Quartz and job tables kept until S5
Main drift Host switch on main instead of a long-lived branch; S1.3 keeps the default host green
Behaviour gaps found only in production S0 golden suite on both hosts; soak and chaos (S2); 2-week staging bake
No scheduled-message cancel in Wolverine S0.10 generation guards
Interface-inheritance routing or per-pod fan-out lost S0.11, S0.12
B1 regexes stale, or reply queues denied S1.4, S2.9; auth session sign-off
Saga or job state lost Drain, export, dry run (S2.11), D9 counts
A CVE during the bridge period R1–R4, T1
Wolverine maintainer concentration MIT; optional JasperFx support
Both stacks in images until S5 (size, accidental use) S5 removal; C1 pinning

10. Coordination

  • Open PRs: NV (gh pr list was refused). Run gh pr list --state open --json number,title,files per phase.
  • Auth session: owns B1 (TLS, principals, regexes) and signs off S1.4 and S2.9. S4 follows the auth production cutover, and B1 lands first (R3).
  • Persistence plan: only IDomainEventSink touches the bus. S1 retargets it; do not adopt the MassTransit Mongo outbox (#3984).
  • FEAT-024: 5 recurring schedules and 7 consumers; S0.5 and S3.2 cover them.
  • Bulk PDF (#3669): the Lambda, the agent and the shared queue; S4 needs #3669 resolved.
  • ADR-020: job consumers; S2.5 reuses its crash suite.
  • Notifications stack (#3932, #3938–#3947, #3965): new consumers put their logic in host-neutral handlers from S1 onwards.

11. Consequences

  • SyRF runs an unmaintained 8.5.11 through most of 2027, under named mitigations and triggers.
  • About 21–30 engineer-weeks (PROPOSAL), plus one planned production maintenance window.
  • Gains:
  • an MIT stack with no licence administration;
  • SyRF-owned saga and job state;
  • built-in cron, concurrency and test tooling;
  • no Quartz service after S5;
  • narrower B1 regexes.
  • Changes:
  • SQL Server stays long term (Wolverine's store), and API and PM connect to it;
  • contracts become concrete records;
  • tests move from ITestHarness to tracked sessions.

12. Option 1 record

Licensing v9 would have been 1–2 weeks of effort with no topology or state change. Its price is unpublished (licence). Rejected by owner decision 2026-10-03.

13. Decisions for Chris

Recorded: v9 rejected (2026-10-03); Wolverine with a single switch-over (2026-10-04). Open:

ID Decision Recommendation
D0 Approve P1 (8.5.11 + central package management) now Approve
D1 Long-lived branch or host switch on main Host switch on main (§4 item 7)
D2 Accept the §5 risk window and triggers T1–T6 Accept: target Q3 2027, hard limit 2028-11-14
D3b Saga target SyRF-owned Mongo state, not the EF/SQL saga
D4 Production maintenance window and comms PROPOSAL: a weekday 05:00–09:00 UTC, users notified a week ahead, with banner text agreed in S3
D5 Prune the 637 orphaned saga rows (export first, dry run on the snapshot) Approve for S4
D6 Soak thresholds (S2.3) PROPOSAL: 24 h, ≥ 2× peak volume, _error = 0, p95 ≤ 1.2× baseline
D7 Staging bake length (S3) 2 weeks, plus every scheduled and recurring path observed
D8 Sequencing against the auth cutover and B1 B1 before 2027-03-31; S4 after the auth production cutover; S1.4 regexes signed off by the auth session
D9 Supply read-only production counts of Quartz triggers and job sagas (refused to this session) Needed before S2.10 and S4
D10 Keep SQL Server long term Yes: Wolverine's store needs it (Rebus was the only all-Mongo route, and it lost, §3.2)
D11 Timing: S0 starts Nov 2026 after P1 Approve

14. References

  • Review: docs/planning/architecture-review-2026-10.md items 11–14 (PR #3961); synthesis §5.1.
  • The October 2026 persistence and messaging plan (docs/planning/, temporary): IDomainEventOutbox/IDomainEventSink.
  • The application-authority broker security contract (docs/planning/application-authority-transition/broker-security-contract.md, temporary).
  • ADR-015, ADR-017, ADR-020.
  • Issues: #3984, #3986, #3988, #3669.