ADR-021: Messaging after MassTransit v8¶
This ADR is for planning only: no work starts until Chris answers the open decisions in §13. It answers #3986.
SyRF runs MassTransit 8.4.0 across five deployables, and open-source v8 maintenance ends after 2026.
Two owner decisions are already recorded:
- 2026-10-03: licensing v9 is rejected.
- 2026-10-04: SyRF moves to Wolverine, in one coordinated switch-over with no side-by-side running. Extensive isolated testing has to give enough confidence before the production switch.
Until the switch, SyRF runs a bridge: the last open-source release, 8.5.11, under an explicit risk window.
Status¶
In-Review, 2026-10-04. Two decisions are recorded:
- v9 rejected (2026-10-03);
- Wolverine with a single switch-over chosen (2026-10-04).
The decisions still open are listed in §13. Evidence markers:
- V: VERIFIED by reading the code on
mainat85e6facf7. - C: CORRECTED, where the earlier text or the review appendix was wrong.
- NV: NOT VERIFIED, with the reason.
- W: a web fact, with a link and the date checked.
1. Context¶
1.1 External facts (W, checked 2026-10-03/04)¶
| Fact | Source |
|---|---|
| v9 is commercial (Massient, Inc., 9.2.3). Rejected by owner decision 2026-10-03. | nuget, massient.com |
| v8 stays Apache-2.0 ("MT v8 remains open-source", Chris Patterson). 8.5.11 was published 2026-09-30. It depends on RabbitMQ.Client 7.2.2, MongoDB.Driver 3.12.0, Quartz [3.22, 4.0) and EF Core Relational 10.0.0, and its nuspec has a net10.0 group. | X post; NuGet API |
| The end of v8 maintenance ("fixes through 2026") is stated only by third parties. No primary Massient wording was found, and Massient's licence says only that v8 is unsupported under it. | Jovanovic, licence |
| .NET 10 LTS support ends 2028-11-14. .NET 11 (STS) is at RC1, GA expected November 2026. | .NET policy |
Wolverine 6.45.0 (2026-10-02): MIT, net9/net10. Packages for RabbitMQ, SqlServer and EntityFrameworkCore. There is no official MongoDB store; the community Wolverine.MongoDB 1.1.0 has a single owner. Wolverine uses Microsoft DI. JasperFx support plans cost $3k–15k a year and are optional. |
nuget, sagas, support |
Wolverine recurring schedules: opts.Schedules.ScheduleRecurring takes cron with a time zone. ExclusiveNodeWithParallelism(n) runs a listener on one node, n in parallel, with failover. |
recurring, exclusive node |
Wolverine scheduling (ScheduleAsync) is durable only with database persistence. No API to cancel a scheduled message is documented. |
message bus |
Wolverine testing: tracked sessions (TrackActivity, InvokeMessageAndWaitAsync), StubAllExternalTransports, StubWolverineMessageHandling. |
testing |
| Rebus 8.9.5 (MIT). Corrections against the previous version of this ADR (C): • Rebus.RabbitMq 10.1.1 (2026-01-19) ships net8.0/net9.0/netstandard2.0, which run on net10, so the missing net10 build is not a blocker. • Rebus.MongoDb 9.0.0 (2024-12-18) depends on MongoDB.Driver ≥ 3.0.0, which is compatible with SyRF's 3.10.0. • Request/reply comes from the Rebus.Async add-on (10.0.0, 2023-11-15). |
NuGet API (Rebus) |
| Brighter (Paramore.Brighter) 10.7.0 (2026-07-29): net8/9/10. Has a MongoDB outbox/inbox (driver 3.10.0), Quartz/Hangfire schedulers and RPC. No sagas. | nuget |
| CAP (DotNetCore.CAP) 10.0.2 (2026-08-01): net8/9/10. Has MongoDB storage (driver 3.9.0) and delayed publish. No sagas and no request/reply. | nuget |
| NServiceBus 10.2.9 (net10, proprietary) has official MongoDB persistence. It costs from $10 per production endpoint per day; Community is free up to 3 endpoints and 10k msg/day. | pricing |
| OpenTransit (re-checked 2026-10-03): last push 2026-08-12 (an upstream merge). No releases, no NuGet packages and no named maintainers. | GitHub |
Interop footnote: Wolverine can talk to MassTransit over RabbitMQ (UseMassTransitInterop), but with
limits: one DefaultIncomingMessage<T> per listener, and warned-of reply-queue "hiccups"
(interop). The
single switch-over makes interop irrelevant, so it is not a scoring criterion.
1.2 What SyRF uses: inventory (V unless marked)¶
| ID | Item | Count / shape | Evidence |
|---|---|---|---|
| I1 | Bus bootstrap and naming | One helper builds every in-cluster bus: AddConsumers(entryAssembly), the publish scheduler and ConfigureEndpoints. MapCommandQueues maps each ICommand to a queue; three bulk-PDF commands share one queue. |
MassTransitHelpers.cs:30-62,174-202; IntegrationExtensions.cs:20-56 |
| I2 | Consumers | 31 IConsumer + 4 IJobConsumer: PM 28 + 4, API 2, the PDF agent 1. One Fault<T> consumer, about 14 definitions with ConcurrentMessageLimit, and 2 per-pod temporary fan-out endpoints in the API. |
PM.Endpoint/Consumers/*.cs; PM Program.cs:484-582; ActivityClaimRevokedConsumer.cs:24,106; ProjectStatisticsChangedConsumer.cs:24,122; SearchUploadSavedFaultConsumer.cs:14 |
| I3 | Request/response | 10 clients: API ×8, PDF agent ×1, PM ×1. | BulkPdfNotifierAuthorityController.cs:17-23; StudyController.cs:40; SearchController.cs:249,270; PdfAgentMassTransitRegistration.cs:25; PM Program.cs:581 |
| I4 | Saga (MongoDB) | SearchImportJobStateMachine (325 lines): 5 events (one is Fault<…>), a 60-minute timeout and partitioner(1), stored in pmSearchImportJobState. Production (read-only aggregate, 2026-10-03) holds 4 Completed, 1 Error, 65 Uploading and 572 Parsing, mostly orphaned. |
SearchImportJobStateMachine.cs:15-84,282,298-320; SearchImportJobState.cs:13-38; MongoMassTransitExtensions.cs:12-27 |
| I5 | Job Service sagas (SQL Server, EF) | JobSaga, JobTypeSaga and JobAttemptSaga live in the Quartz service. Four jobs set JobOptions (timeouts 5 min to 2 h, retries, SetConcurrentJobLimit(3)). Jobs are submitted by the API and the Lambda. |
QuartzServiceCollectionExtensions.cs:131-147; Migrations/20250324010206_InitialCreate.cs; ReferenceFileParseJobConsumer.cs:68-90; RoBProcessingJobConsumer.cs:53-67; S3FileReceivedFunction.cs:252 |
| I6 | Scheduling (Quartz, SQL Server) | 9 SchedulePublish + 4 CancelScheduledPublish, with tokens stored in MongoDB. Only the idle timer has a generation guard: the suspended-session and liveness timers rely on the broker cancel alone. 5 recurring schedules (FEAT-024). UseScheduledRedelivery on 5 definitions. The PDF agent's RetryLater uses the shared Quartz. |
NotificationHub.cs:281,742,902,1551,1571,1590,1598; ReviewController.cs:357,1548; SlotReservation.cs:212-232; MarkSessionIdleConsumer.cs:60-64; PM Services/*Schedule.cs |
| I7 | Outbox | No MassTransit transactional outbox, only UseInMemoryOutbox. SyRF's own outboxes (MongoDB/DynamoDB) do not depend on the transport. In the persistence plan's IDomainEventOutbox, only IDomainEventSink touches the bus, so it is compatible. |
MassTransitHelpers.cs:128,146,158; the October 2026 persistence and messaging plan (IDomainEventOutbox; temporary planning doc) |
| I8 | Retry | No global policy. 15 per-consumer retry, redelivery or outbox sites, plus JobOptions.SetRetry (#3984). |
ReviewSessionConsumerRetry.cs:14-31 |
| I9 | Topology | Exchanges are Namespace:Type, with interface-hierarchy exchange-to-exchange bindings. Queues are <Assembly − .Messages>.<Type> and <Assembly − .Endpoint>.<Type>; 5 use the default formatter. MassTransit infrastructure names (quartz, Job*, MassTransit.*). The B1 permission regexes are built from these names. |
broker-security-contract.md:189-215,235-262 |
| I10 | Contracts | 35 files, 46 interfaces, sent through 68 interface-typed anonymous initialisers (a MassTransit-only feature). PM.Messages → PM.Core. Serialiser: System.Text.Json. |
PM.Messages.csproj:9; grep counts |
| I11 | S3-notifier Lambda | Creates a bus per S3 record. C: the vhost comes from each object's metadata, so any cached sender must be keyed by vhost. The reconciler Lambda caches one bus. Production runs older notifier code (#3669). | S3FileReceivedFunction.cs:56-60,202-227,507-527; BulkPdfScheduledReconciler.cs:1139-1191 |
| I12 | PDF agent (ARRNC) | One consumer and a claim client, with a Paused mode (no receive endpoint). It holds no database credentials. | PdfAgentMassTransitRegistration.cs:14-40 |
| I13 | Tests | 60 files import MassTransit; 14 use ITestHarness; 5 use Testcontainers RabbitMQ. The e2e/ stack runs RabbitMQ and SQL Server. |
StateMachineTestFixture.cs:28-50; BulkStudyUpdateAtomicBrokerTests.cs:51,223-256 |
| I14 | Broker TLS (B1) | RabbitMqTls makes certificate checks strict on 5671 (syrf#3878). The live B1 PRs are held until main can be deployed to production (Chris, 2026-10-02). |
RabbitMqTls.cs:21-74 |
| I15 | Packages and DI | 13 csproj files reference MassTransit, with no central package management. SharedKernel's MassTransit reference is unused (V). DI is Lamar 15.0.1. 8.x already runs on net10.0 in production. | SyrfHelpers.cs:19; Directory.Build.props:8 |
NV:
- Production SQL counts for Quartz triggers and job sagas: the read was refused by the session's permission classifier (D9).
- Daily message volume.
- Whether open PRs touch these files:
gh pr listwas refused.
2. Scope¶
| ID | Problem | Impact | Evidence | Issue |
|---|---|---|---|---|
| S1 | v8 maintenance ends after 2026 and v9 is rejected | Unpatched messaging from 2027 on the Internet-facing broker path (Lambda, ARRNC) | §1.1 | #3986 |
| S2 | SyRF is on 8.4.0, not 8.5.11 | It misses free fixes | csproj ×13 | #3986 |
| S3 | Many MassTransit-only features are in use (Job Service, cancel tokens, interface initialisers, Fault<T>) |
The switch touches every programme | I2–I10 | #3986 |
| S4 | Contracts depend on the domain | Contract changes cost more | PM.Messages.csproj:9 |
#3988 |
3. Options¶
Effort is in engineer-weeks and is a PROPOSAL: S is under 1, M 1–3, L 3–8, XL over 8. "Built in" means the library provides the feature, so SyRF does not own that code.
| Inventory need | Wolverine (chosen) | Rebus | Brighter | CAP | NServiceBus | In-house (RabbitMQ.Client) | OpenTransit |
|---|---|---|---|---|---|---|---|
| I4 saga + 60-min timeout | built in (EF/SQL); SyRF chooses Mongo state (D3b) | built in (Mongo) | none | none | built in (Mongo) | build | as v8, but nothing shipped |
| I5 job timeout/retry/limit 3 | handlers + ExclusiveNodeWithParallelism(3) + error policies |
handlers; no cluster-wide limit | handlers; no limit | handlers; no limit | handlers | build | as v8 |
| I6 durable delay | built in (SQL store) | built in (Mongo timeouts) | Quartz/Hangfire | built in | built in | build | as v8 |
| I6 recurring (5) | cron built in | none | via Quartz | none | none | build | as v8 |
| I6 cancel tokens | none documented → generation guards | none | via Quartz | none | none | build | as v8 |
| I3 request/reply (10) | built in | Rebus.Async add-on | RPC | none | built in (callbacks) | build | as v8 |
| I7/I8 inbox, outbox, retry | built in | built in | built in | built in | built in | build | as v8 |
| I13 test harness | tracked sessions | basic | basic | basic | good | build | as v8 |
| Store for the above | SQL Server (already run) | MongoDB (could retire SQL Server) | MongoDB + Quartz SQL | MongoDB | MongoDB | n/a | SQL + Mongo |
| Licence / steward | MIT; JasperFx (also maintains Lamar); releases 2026-09-30 and 2026-10-02 | MIT; one main maintainer | MIT; community | MIT; community | proprietary; about $11k–15k a year (PROPOSAL) | SyRF forever | none yet |
| Effort (PROPOSAL) | ~20–28 wk total (§8) | ~22–30 wk | +saga build | +saga and RPC build | ~18–26 wk + licence | 30+ wk | 0, but nothing exists |
3.1 Why Wolverine¶
It covers the most of SyRF's inventory with built-in features, so SyRF owns the least code.
- Scheduling: cron recurring schedules (the 5 FEAT-024 schedules) and durable delayed scheduling.
- Jobs:
ExclusiveNodeWithParallelismfor RoB's cluster-wide limit of 3. - Reliability: inbox, outbox and retry policies.
- State: EF/SQL Server sagas.
- Messaging patterns: request/reply.
- Testing: tracked sessions.
- Stewardship: an active, company-backed steward (JasperFx, which also maintains Lamar, SyRF's DI).
- It runs on the SQL Server SyRF already operates for Quartz and the job sagas, so no new kind of store is needed.
3.2 Why the others lost¶
- Rebus has one real advantage: everything lives in MongoDB, so SQL Server could be retired. It lost because:
- it has no recurring cron (I6), so SyRF would build a leader-elected scheduler;
- it has no cluster-wide concurrency limit (I5);
- request/reply needs a 2023 add-on (I3);
- the Mongo store is unchanged since 2024-12;
- it has a single main maintainer.
Each gap is SyRF-owned code, and together they outweigh retiring SQL Server. - Brighter and CAP have no sagas, and CAP has no request/reply. - NServiceBus is a recurring proprietary licence, the same objection as v9. - In-house would make SyRF own everything. - OpenTransit has shipped nothing; it is re-checked under trigger T6.
4. Decision¶
- Bridge (P1): upgrade 8.4.0 → 8.5.11 under
Directory.Packages.props, and drop the unused SharedKernel reference. - Risk window (§5): run 8.5.11 unsupported from 2027-01-01 until the production switch. The target is the switch in Q3 2027 (PROPOSAL); the hard limit is 2028-11-14.
- Target: Wolverine, single coordinated switch-over of all five deployables (§7 S0–S5).
- Saga (D3b, recommended): SyRF-owned state in Mongo, using the existing repository,
Audit.Versionconcurrency and Wolverine handlers, plus a durable scheduled timeout. - Job Service: handlers with SyRF job records. Timeouts become cancellation tokens; retries come
from Wolverine error policies; the cluster-wide limit is
ExclusiveNodeWithParallelism(3). - Scheduling: a Wolverine SQL Server message store on the existing instance (new database). Generation guards replace broker cancels. The PDF agent stays credential-free: its RetryLater becomes a command to PM.
- Host switch on main, not a long-lived branch (D1-new, recommended). A startup setting
Messaging:Host = MassTransit | Wolverinedefaults to MassTransit, and both stacks live on main until S5. - Why: main is very active, with five programmes adding consumers. A long-lived branch would drift on all ~35 consumers and the contracts, and the drift would only surface at merge, right before the switch.
- Cost: both packages ship in images until S5. Business logic is extracted into host-neutral handler classes, with thin MassTransit consumers and thin Wolverine handlers calling them.
- Guard: all five deployables must use the same value. It is one environment-level value in cluster-gitops; the Lambda and the agent get it through their own config. A test refuses a mixed configuration in any chart render, and §7 S4 checks it.
- Independent work: contract decoupling (#3988), STJ round-trip tests (#3984), and the
persistence-plan outbox. The outbox is compatible: only
IDomainEventSinkis retargeted in S1.
5. Risk window for unsupported 8.5.11¶
5.1 What actually ends¶
- Nothing stops working. v8 has no licence key or expiry, and its packages stay on NuGet under Apache-2.0 (W).
- What ends is upstream fixes: no CVE patches and no fixes for future .NET, driver or broker changes. The "through 2026" end comes only from third parties (W). 8.5.11 shipped on 2026-09-30, so re-check NuGet monthly.
5.2 Mitigations¶
| # | Mitigation | Owner |
|---|---|---|
| R1 | Pin MassTransit 8.5.11 and its transitive dependencies (RabbitMQ.Client 7.2.2, MongoDB.Driver ≥3.12, Quartz 3.22+, EF Core 10) centrally. Any major bump needs review against this ADR. | P1 |
| R2 | Monthly check of NuGet, GitHub advisories and Dependabot for MassTransit, RabbitMQ.Client, Quartz and MongoDB.Driver, with the result recorded on #3986. | maintainer |
| R3 | Finish B1 (strict 5671, per-principal permissions, close public 5672) before 2027-03-31 (PROPOSAL). | auth session |
| R4 | Fork-and-patch fallback: a vendored 8.5.11 source build, used only to apply a security fix (Apache-2.0 allows it). | on trigger |
5.3 Triggers¶
| # | Trigger | Response |
|---|---|---|
| T1 | A CVE with CVSS ≥ 7.0 (PROPOSAL) in MassTransit or a transitive dependency, with no fixed 8.x | R4 within 14 days (PROPOSAL) |
| T2 | .NET 11/12 needed before 2028-11-14 and incompatible | Stay on .NET 10 LTS; pull the switch forward |
| T3 | RabbitMQ.Client 8.x or MongoDB.Driver 4.x forced (security, Atlas, broker) and 8.5.11 cannot use it | Pull the switch forward; R4 meanwhile |
| T4 | Quartz 4.x or EF Core 11 incompatible with MassTransit.Quartz or MassTransit.EntityFrameworkCore | Hold the versions; pull forward |
| T5 | A RabbitMQ server upgrade breaks the 8.x client | Hold the broker version; pull forward |
| T6 | OpenTransit ships a maintained net10 release | Note it on #3986; the decision stands unless Chris re-opens it |
6. MVP boundary and flags¶
- MVP: P1 + S0. That is 8.5.11 on every deployable, proven in staging, plus the characterisation baseline and the generation guards, which ship to production early and on their own.
- Out of scope here, and where each is tracked:
- global retry policy: #3984;
- contract decoupling: #3988;
- B1 live steps: the auth session;
- liveness probes: review item 14.
- Flags:
- P1 is not flagged; roll back by image tag.
- S0's generation guards are not flagged: they are additive, stale-delivery no-ops, and they are tested.
- The migration is flagged through the startup host switch (§4 item 7), not a runtime flag. The default is MassTransit, so it is changed per environment at the switch, and rollback means switching back.
7. Programme: single switch-over¶
Every PR meets these criteria:
| # | Condition → result | Verification |
|---|---|---|
| C1 | Any messaging package version → resolved from Directory.Packages.props only |
CI package listing (one project at a time) |
| C2 | A PR touches a handler → its test file is updated, and the S0 suite passes on the default host | unit/Testcontainers |
| C3 | Topology differs from the S0 snapshot → a diff is attached and the B1 regexes are regenerated | snapshot + B1 fixture |
| C4 | Docs → this ADR and .claude/rules are updated in the same PR |
docs |
| C5 | Production → never promoted by a PR. Production is a separate Chris-approved step (S4). | governance |
Effort is a PROPOSAL. E2E runs only in the hermetic e2e/ stack; staging is for human testing.
P1: Bridge to 8.5.11 (S–M, 1–2 wk)¶
- Files: the 13 MassTransit csproj files; a new
Directory.Packages.props;SyRF.SharedKernel.csproj. - Approach: an EF model check on
JobServiceSagaOverrideDbContext. - Rollback: revert the image tag.
| # | Condition → result | Verification |
|---|---|---|
| P1.1 | Build → every MassTransit package resolves to 8.5.11; no 8.4.0 is left transitively | CI package listing |
| P1.2 | Job Service EF model → no diff, or a reviewed migration plus a rollback script | EF tooling + Testcontainers SQL |
| P1.3 | Saga, scheduler and broker suites → pass | Testcontainers |
| P1.4 | Hermetic stack: import, bulk update, presence idle/suspend → each completes | E2E |
| P1.5 | Staging → an import reaches Completed; no _error growth for 24 h (PROPOSAL) |
staging check (human) |
| P1.6 | Strict TLS → an untrusted chain is rejected | fixture |
S0: Characterisation baseline on MassTransit (L, 3–4 wk)¶
- Files:
- a new
SyRF.Messaging.Characterisation.Tests, written against a host-neutral driver that can run unchanged against either host; - generation guards for
RemoveSuspendedSessionandCheckConnectionLiveness, matchingSlotReservation.cs:218-224andMarkSessionIdleConsumer.cs:60-64. - Rollback: revert. The tests are additive, and the guards make stale deliveries no-ops.
| # | Condition → result | Verification |
|---|---|---|
| S0.1 | Topology and queue inventory (exchanges, queues, bindings, the message types per queue) → snapshot committed, and it fails on drift | Testcontainers RabbitMQ |
| S0.2 | Each of the 46 contracts → round-trips, with golden JSON | contract |
| S0.3 | Saga: every path (started, saved, parsed, completed, each fault) plus the 60-min timeout → golden terminal states and the SearchImportJobState BSON shape (CSUUID) |
Testcontainers |
| S0.4 | Each of the 4 jobs → timeout, retry count, ignored exceptions and concurrency limit pinned | Testcontainers |
| S0.5 | Every scheduled message (9 sites) and recurring schedule (5) → fires once, at the expected time | Testcontainers SQL |
| S0.6 | Each request/reply client (10) → reply and timeout behaviour pinned | Testcontainers |
| S0.7 | Bulk-PDF claim/progress/finalize → serialised, single writer | Testcontainers |
| S0.8 | Lambda: records for two vhosts in one invocation → each reaches its own vhost; malformed-record isolation pinned | LocalStack + Testcontainers |
| S0.9 | Agent Paused mode → connects with no receive endpoint | Testcontainers |
| S0.10 | A stale suspended/liveness timer after a reconnect → a no-op without a cancel (ships to production early, Chris-approved) | unit + Testcontainers |
| S0.11 | Interface inheritance: every consumer of a base or shared interface → listed with its exchange-to-exchange binding, and a test asserts it receives (Wolverine routes by concrete type) | Testcontainers |
| S0.12 | The API's per-pod temporary endpoints → a test pins fan-out to every pod and auto-delete on shutdown | Testcontainers |
S1: Build Wolverine behind the host switch on main (XL, 10–14 wk)¶
- Scope: all five deployables.
- Thin Wolverine handlers call host-neutral logic.
- Concrete
recordcontracts implement the 46 interfaces; MassTransit can still publish them as the interface. - Explicit routes keep queue names where possible.
Fault<T>paths become explicit failure events.- The 15 retry sites become Wolverine error policies.
- Persistence: a SQL Server message store, with API and PM connected; the saga as SyRF Mongo state
(D3b); jobs as handlers with SyRF job records and
ExclusiveNodeWithParallelism(3); the 5 recurring schedules onScheduleRecurring. - Client paths: the Lambda is send-only, with a sender cache per vhost; the agent's RetryLater is
routed via PM; strict TLS is ported from
RabbitMqTls. - B1: permission regexes regenerated from the Wolverine topology, including reply queues (pinned names; automatic reply queues need queue-declare permission).
IDomainEventSink: retargeted.- Rollback: none needed. The default host stays MassTransit.
| # | Condition → result | Verification |
|---|---|---|
| S1.1 | Messaging:Host=Wolverine on each deployable → it starts under Lamar and registers every handler, route and schedule |
unit + Testcontainers |
| S1.2 | Mixed host values across charts → the render test fails | chart test |
| S1.3 | Default host → the S0 suite still green (no regression on main) | CI |
| S1.4 | Wolverine topology → a reviewed diff from S0.1, with the B1 regexes regenerated (reply queues included) and the auth session's sign-off | snapshot + review |
| S1.5 | The Lambda package on Wolverine → builds and runs in LocalStack | LocalStack |
S2: Confidence gates in isolated environments only (L, 3–4 wk)¶
All gates run in the hermetic e2e/ stack and Testcontainers, never against staging.
| # | Condition → result | Verification |
|---|---|---|
| S2.1 | The full S0 suite on Wolverine → green, unchanged |
Testcontainers |
| S2.2 | Full E2E on Wolverine → green, no new failures against the known main baseline |
E2E |
| S2.3 | Soak on Bramble: 24 h, ≥ 2× the observed peak daily volume (volume NV, measure first) → every _error queue = 0; p95 handling latency ≤ 1.2× the MassTransit baseline from the same run (PROPOSAL) |
soak harness (Bramble) |
| S2.4 | Broker restart mid-traffic → no message lost; consumers reconnect within 60 s (PROPOSAL) | chaos test |
| S2.5 | A pod killed mid-job (RoB, bulk update) → the job resumes or retries per S0.4; ADR-020 rollback semantics hold | chaos test |
| S2.6 | SQL Server outage of 5 min (PROPOSAL) → scheduled messages fire late but exactly once, and recurring schedules resume | chaos test |
| S2.7 | Lambda cold start → no worse than 8.x + 20% (PROPOSAL) | LocalStack |
| S2.8 | Strict TLS on 5671 from every client → an untrusted chain is rejected | fixture |
| S2.9 | B1 permission fixture (positive and negative sets, reply queues) → passes on the Wolverine topology | broker fixture |
| S2.10 | Cut-over rehearsal on a production-shaped snapshot (the staging copy, or anonymised data; no production reads without Chris) → the drain runbook and the rollback runbook each complete end to end, timed: drain ≤ 3 h, rollback ≤ 1 h (PROPOSAL) | rehearsal record |
| S2.11 | Saga-state migration script on the snapshot → counts per state match; CSUUID ids are kept; the orphan prune runs as a dry run with an export | script + Testcontainers |
S3: Staging switch-over and bake (M, 2–3 wk elapsed)¶
Staging may be down or reseeded.
- Bake: 2 weeks (PROPOSAL) of normal tester use. Every scheduled and recurring path must be seen firing: daily statistics, the fold sweep, presence timers and the saga timeout.
- Rollback drill: run the rollback runbook on staging, then switch forward again.
| # | Condition → result | Verification |
|---|---|---|
| S3.1 | Staging switched by the S2.10 runbook → finishes within the rehearsed time | staging check (human) |
| S3.2 | Bake period → no _error growth; every path in the bake list observed firing |
staging check + logs |
| S3.3 | Rollback drill → MassTransit restored within ≤ 1 h; then forward again | staging check |
S4: Production switch-over (M, ~1 wk including hypercare)¶
A separate Chris-approved step with a go/no-go checklist:
- S2 and S3 green;
- B1 status known;
-
3669 resolved;¶
- D9 counts in hand.
Maintenance window (D4):
- Stop producers: API admission, the Lambda trigger and agent intake.
- Drain every queue.
- Wait out running jobs (≤ 2 h) and the saga timeout (60 min).
- Run the approved orphan prune with export (D5).
- Migrate live saga rows.
- Re-create active session timers and register recurring schedules on the new store.
- Deploy in coordinated order: PM, then Quartz/SQL store, then API, then the Lambda (ACK, with the
CodeSha256check), then the ARRNC agent. - Verify, then reopen.
Rollback (time-boxed, PROPOSAL 4 h from reopen; roll forward only after that):
- Stop producers.
- Drain the Wolverine queues.
- Redeploy all previous images and the Lambda version.
- Restore the saga export.
- Quartz and the job-saga tables are untouched until S5.
| # | Condition → result | Verification |
|---|---|---|
| S4.1 | Before the window → _error counts, saga-state counts and job-saga counts recorded |
Chris-run read-only check |
| S4.2 | Drain step → every queue reads 0, and no running Job* saga is left |
runbook record |
| S4.3 | After the switch → all five deployables report Wolverine; first import, bulk update, RoB, presence and recurring each succeed |
post-switch checklist |
| S4.4 | First 24 h → no _error growth (PROPOSAL); a rollback trigger means S4 rollback |
governance |
S5: Remove MassTransit and retire Quartz (S–M, 1–2 wk, after ≥ 4 weeks stable, PROPOSAL)¶
| # | Condition → result | Verification |
|---|---|---|
| S5.1 | No MassTransit package and no host switch left in the solution | CI package listing |
| S5.2 | SyRF.Quartz and the job-saga tables → retired after a backup; the B1 regexes have no MassTransit alternatives |
governance + fixture |
| S5.3 | Full hermetic E2E → green | E2E |
8. Order, critical path and timeline¶
P1 ─► S0 ─► S1 ─► S2 ─► S3 ─► S4 (Chris go) ─► S5
#3984 / #3988 / persistence-plan sink: parallel, separate PRs
- Order: S0's generation guards can ship before the rest of S0. S1 can start on contracts and handler extraction while S0 finishes, because they touch different files. S2 needs all of S1.
- Effort (PROPOSAL): P1 1–2 + S0 3–4 + S1 10–14 + S2 3–4 + S3 2–3 + S4 1 + S5 1–2, which is about 21–30 engineer-weeks. Without side-by-side running, testing replaces incremental proof, so this is higher than the earlier phased estimate.
Timeline (PROPOSAL):
| Window | Work |
|---|---|
| Oct–Nov 2026 | P1 |
| Nov 2026 – Jan 2027 | S0 |
| Jan–May 2027 | S1 |
| May–Jun 2027 | S2 |
| Jul 2027 | S3 |
| Aug 2027 | S4 |
| Sep–Oct 2027 | S5 |
The same people also carry the auth cutover, FEAT-024, Bulk PDF production, AF2 and the notifications stack. The switch happens after December 2026, so the §5 risk window is required.
9. Risks¶
| Risk | Mitigation |
|---|---|
| A single release spans 5 deployables, including the off-cluster ARRNC agent and an ACK-managed Lambda | One environment-level host value; S1.2 mixed-config test; coordinated deploy order; agent and Lambda steps in the rehearsed runbook |
| Rollback is heavier than per-endpoint flags (drain, then redeploy everything, then restore state) | Rehearsed and timed in S2.10 and S3.3; 4 h time-box; Quartz and job tables kept until S5 |
| Main drift | Host switch on main instead of a long-lived branch; S1.3 keeps the default host green |
| Behaviour gaps found only in production | S0 golden suite on both hosts; soak and chaos (S2); 2-week staging bake |
| No scheduled-message cancel in Wolverine | S0.10 generation guards |
| Interface-inheritance routing or per-pod fan-out lost | S0.11, S0.12 |
| B1 regexes stale, or reply queues denied | S1.4, S2.9; auth session sign-off |
| Saga or job state lost | Drain, export, dry run (S2.11), D9 counts |
| A CVE during the bridge period | R1–R4, T1 |
| Wolverine maintainer concentration | MIT; optional JasperFx support |
| Both stacks in images until S5 (size, accidental use) | S5 removal; C1 pinning |
10. Coordination¶
- Open PRs: NV (
gh pr listwas refused). Rungh pr list --state open --json number,title,filesper phase. - Auth session: owns B1 (TLS, principals, regexes) and signs off S1.4 and S2.9. S4 follows the auth production cutover, and B1 lands first (R3).
- Persistence plan: only
IDomainEventSinktouches the bus. S1 retargets it; do not adopt the MassTransit Mongo outbox (#3984). - FEAT-024: 5 recurring schedules and 7 consumers; S0.5 and S3.2 cover them.
- Bulk PDF (#3669): the Lambda, the agent and the shared queue; S4 needs #3669 resolved.
- ADR-020: job consumers; S2.5 reuses its crash suite.
- Notifications stack (#3932, #3938–#3947, #3965): new consumers put their logic in host-neutral handlers from S1 onwards.
11. Consequences¶
- SyRF runs an unmaintained 8.5.11 through most of 2027, under named mitigations and triggers.
- About 21–30 engineer-weeks (PROPOSAL), plus one planned production maintenance window.
- Gains:
- an MIT stack with no licence administration;
- SyRF-owned saga and job state;
- built-in cron, concurrency and test tooling;
- no Quartz service after S5;
- narrower B1 regexes.
- Changes:
- SQL Server stays long term (Wolverine's store), and API and PM connect to it;
- contracts become concrete records;
- tests move from
ITestHarnessto tracked sessions.
12. Option 1 record¶
Licensing v9 would have been 1–2 weeks of effort with no topology or state change. Its price is unpublished (licence). Rejected by owner decision 2026-10-03.
13. Decisions for Chris¶
Recorded: v9 rejected (2026-10-03); Wolverine with a single switch-over (2026-10-04). Open:
| ID | Decision | Recommendation |
|---|---|---|
| D0 | Approve P1 (8.5.11 + central package management) now | Approve |
| D1 | Long-lived branch or host switch on main | Host switch on main (§4 item 7) |
| D2 | Accept the §5 risk window and triggers T1–T6 | Accept: target Q3 2027, hard limit 2028-11-14 |
| D3b | Saga target | SyRF-owned Mongo state, not the EF/SQL saga |
| D4 | Production maintenance window and comms | PROPOSAL: a weekday 05:00–09:00 UTC, users notified a week ahead, with banner text agreed in S3 |
| D5 | Prune the 637 orphaned saga rows (export first, dry run on the snapshot) | Approve for S4 |
| D6 | Soak thresholds (S2.3) | PROPOSAL: 24 h, ≥ 2× peak volume, _error = 0, p95 ≤ 1.2× baseline |
| D7 | Staging bake length (S3) | 2 weeks, plus every scheduled and recurring path observed |
| D8 | Sequencing against the auth cutover and B1 | B1 before 2027-03-31; S4 after the auth production cutover; S1.4 regexes signed off by the auth session |
| D9 | Supply read-only production counts of Quartz triggers and job sagas (refused to this session) | Needed before S2.10 and S4 |
| D10 | Keep SQL Server long term | Yes: Wolverine's store needs it (Rebus was the only all-Mongo route, and it lost, §3.2) |
| D11 | Timing: S0 starts Nov 2026 after P1 | Approve |
14. References¶
- Review:
docs/planning/architecture-review-2026-10.mditems 11–14 (PR #3961); synthesis §5.1. - The October 2026 persistence and messaging plan (
docs/planning/, temporary):IDomainEventOutbox/IDomainEventSink. - The application-authority broker security contract (
docs/planning/application-authority-transition/broker-security-contract.md, temporary). - ADR-015, ADR-017, ADR-020.
- Issues: #3984, #3986, #3988, #3669.