Screening source-write development benchmark¶
ProjectScreeningWriteBenchmark measures screening source saves on a real MongoDB replica set in three
arms, selected per cell by the ScreeningWriteArm enum:
| Arm (artifact metric) | What one submission executes |
|---|---|
SourceOnly (*_source_only) |
Statistics writes off. With ReviewEligibilityPolicy off: the capacity-save overload with tracking off, a version-filtered FindOneAndReplace (find + findAndModify). With it on: the controller's flag-off branch, an uncached Project read per attempt and TrySaveActivityReviewAsync (its eligibility transaction and Project token write). |
Transactional (*_materialized) |
Unchanged from earlier artifacts: the six-argument capacity save inside the statistics snapshot transaction, with the durable-mode read, envelope and receipt pre-check and the real screening writer and coordinator. ReviewEligibilityPolicy off only. |
Fold (*_fold) |
ProjectScreeningFoldSave exactly as ReviewController.TrySaveScreeningOnFoldPathAsync drives it with materializedProjectStatisticsFold on and the project in fold mode, over a target that mirrors ReviewScreeningFoldTarget: private load with the persisted inclusion statuses, durable-mode and control reads, duplicate checks, deadline, free fold-only retries and the three-attempt budget, then SaveScreeningOnFoldPathAsync (or TrySaveActivityReviewAsync with eligibility on). The real ProjectStatisticsFoldWorker drains after every batch, outside the timed window. |
Every cell runs with ReviewEligibilityPolicy off and on. With it off, all three arms run; with it on,
SourceOnly and Fold run (*_eligibility_source_only, *_eligibility_fold). The transactional arm is
not run with eligibility on: the gate does not compare against it, and running it would need a new
driver path rather than the unchanged historical one.
Like-for-like. With tracking off, the fold arm's write is the plain version-filtered
FindOneAndReplace; the SourceOnly arm's capacity overload with tracking off issues the same
command with the same filter, minus statistics, which is why it is the gate's baseline. The
controller's own flag-off branch with tracking off uses TrySaveExistingAsync (ReplaceOne, one
update) instead; both are one write round trip. The controller's GetOrDefaultAsync before
TrySaveScreeningAsync (cached process-wide for two seconds) is outside every arm: SourceOnly and
Transactional time their own source load, which the fold arm replaces with its private load, as the
design's command budget counts it. Each request owns a separate repository and cache, reloads after an
optimistic conflict and keeps the API's three-attempt limit. Reviewer membership and authorization
(except the eligibility stage check), event dispatch and HTTP are outside the fixture.
The minimum corpus, PS-WRITE-01, contains one screening project and ten studies. Each matrix cell
uses 1, 2, 5 or 10 reviewers submitting to the same study or separate studies. Later rounds alternate
inclusion/exclusion decisions by the same reviewers (the reviewer-correction pattern). With
eligibility on, every reviewer is an active member with the stage Review grant. The artifact's corpus
counts describe maximum matrix dimensions, not a claim that every cell has ten active reviewers. This
is deliberately not a large-project PS-DS-02 benchmark. Scale values other than 1 are refused.
Reproduce¶
From the PR worktree root:
SYRF_STATS_DATASET=PS-WRITE-01 SYRF_STATS_ITERATIONS=100 SYRF_STATS_WARMUP=10 \
SYRF_STATS_RESULTS_DIR=/tmp/syrf-screening-write-results \
dotnet test src/libs/project-management/SyRF.ProjectManagement.Mongo.Data.Tests/SyRF.ProjectManagement.Mongo.Data.Tests.csproj \
--filter FullyQualifiedName~ProjectScreeningWriteBenchmark
One run measures every arm in every cell, with ReviewEligibilityPolicy off and on; there is no
per-arm variable. Command counting is opt-in. Add SYRF_STATS_COUNT_ROUNDTRIPS=1 to the command above for the
command breakdown; leave it unset for a no-subscriber control timing. Without the dataset variable,
ordinary CI skips this benchmark. The existing replica-set fixture owns
and removes its mongo:8.0 container. Existing shared-runner resource limits apply when
SYRF_TEST_JOB_KEY is set; the artifact records that policy and the runtime/server environment.
No deployed database, credentials or live flags are needed. Run it only on an idle host: the
programme's threshold is a one-minute load average per CPU of at most 0.5 before the run starts. The
benchmark's own start-up (the test host and a fresh MongoDB container) lifts the load of every run, so the load
during a run is a measurement, not a hygiene check.
This command is a development run. Gate (b) acceptance is a different recipe: five timing runs of 300
iterations and one counted run, aggregated by scripts/aggregate-statistics-write-benchmark.py; see
the acceptance run. Each run's own gate (b)
figures are in the artifact note <cell>_fold_gate_b (and <cell>_eligibility_fold_gate_b), and are one
run's figures, not the gate; the historical transactional gate stays in <cell>_p95_gate. Set
SYRF_STATS_SOURCE_COMMIT=$(git rev-parse HEAD) to record the commit in the artifact's source_commit
note (otherwise it reads unrecorded).
What the harness now records¶
Following the write-overhead diagnosis
(screening-write-overhead-diagnosis.md, #3255), three things
changed in what this harness reports. The artifact schema is extended, never renamed, so an existing
reader of Durations, Notes or the *_outcomes keys is unaffected.
- Command counting is opt-in via
SYRF_STATS_COUNT_ROUNDTRIPS. ACommandStartedEventsubscriber forces the driver to materialise every command document, and the materialised arm issues many times more commands than the source-only arm, so the recorder's own cost is asymmetric between the two arms being compared. Running once without it gives a no-subscriber control timing. When it is off,commandsandcommit_attemptsreadnot_countedrather than zero, and noRoundTripsrecord is emitted. - Driver handshake and monitoring commands are not counted (
hello,isMaster,buildInfo,ping,saslStart,saslContinue,authenticate,getnonce,endSessions), the same list the sharedMongoCommandRecorderhas always applied. A pooled connection opened inside a timed window sends them (connections_openedstill reports it), and counting them gave 4.01 instead of 4 commands per fold submission in the first 2026-10-05 acceptance run (below).ScreeningWriteCommandCountTestspins the list and that application commands (find,findAndModify,killCursors, ...) are still counted. - Commands are broken down by driver command name, in the artifact's existing
RoundTripsfield, so a reduction is attributable to a specific command class rather than only to a batch total. The per-name counts are windowed exactly as the total is, so they sum tocommands. - Per-attempt command counts are recorded in the
*_commands_per_attemptnote, for the single-reviewer cells only. With two or more reviewers the submissions run concurrently against one shared counter, so a per-attempt delta would include another reviewer's commands; the artifact says so rather than publishing a figure it cannot support.
The duplicate pre-check's assertion also moved outside the timed window.
ResolveByReceiptAsync itself stays inside it, because production runs the same pre-check before
opening a transaction.
Evidence and interpretation¶
The existing BenchmarkResults format records p50/p95 submission duration and repository-attempt
duration. Submission time includes source reload, envelope/receipt reads, classification and retries.
It excludes seeding, assertion queries, HTTP middleware, assignment, event dispatch and background
work. The source-only baseline performs no statistics admission reads. Candidate/baseline ordering
alternates by reviewer count. Each variant has its own warmup.
Every batch checks exact persisted decisions against successful submissions, and enabled batches
also check exact authoritative/projection parity, operation receipt/delta presence, committed projection
revision advancement and the coalesced outbox's client revision. Exhausted retries are reported as
failures to save. A repeated decision can save source without a classified delta; these saves are
counted separately and correctly have no statistics receipt/delta. There is no silent exception-to-success
conversion. Driver command monitoring counts all commands during submissions, commit attempts and
UnknownTransactionCommitResult failures; it excludes seeding and post-batch assertion queries.
There is no injected unknown-commit failure in this timing run.
The historical p95 overhead gate (*_p95_gate, the transactional arm against source-only) is strictly
less than 10%. Its result is emitted for each matrix cell, never asserted as a CI correctness test on a
shared development host. Gate (b), for the fold arm, is defined
below. All-submission latency includes exhausted
requests, so a lower p95 with a worse success rate cannot establish a usable performance improvement.
The artifact includes saved, conflict, exhausted and no-delta counts so that tradeoff remains visible.
Maximum serialized summary bytes are also recorded.
This slice keeps active-reviewer tracking off and enables only ProjectScreening maintenance. It does not establish reviewer-family maintenance overhead, reservation/assignment throughput, annotation write cost, non-capacity API latency, large-project capacity, live rollout consistency or production acceptance. Those required programme measurements remain ordered follow-up evidence. Existing deterministic replica-set tests separately prove the non-capacity transaction conflict response and rollback.
Development result: 5 September 2026¶
The 100-iteration artifact records 10 warmup batches per variant. All exact correctness assertions passed. Every p95 overhead gate failed. These measurements leave the performance acceptance gate open.
| Target | Reviewers | Baseline p95 ms | Materialized p95 ms | Overhead | Baseline saved | Materialized saved |
|---|---|---|---|---|---|---|
| Same study | 1 | 8.57 | 64.64 | 654% | 100/100 | 100/100 |
| Same study | 2 | 11.24 | 47.56 | 323% | 200/200 | 117/200 |
| Same study | 5 | 20.35 | 44.77 | 120% | 300/500 | 141/500 |
| Same study | 10 | 28.90 | 44.55 | 54% | 329/1000 | 146/1000 |
| Different studies | 1 | 6.19 | 48.65 | 686% | 100/100 | 100/100 |
| Different studies | 2 | 6.14 | 94.55 | 1441% | 200/200 | 200/200 |
| Different studies | 5 | 7.71 | 80.96 | 950% | 500/500 | 350/500 |
| Different studies | 10 | 21.89 | 134.89 | 516% | 1000/1000 | 598/1000 |
The different-study result exposes contention on shared project statistics despite independent source
records: 1,406 optimistic conflicts and 402 exhausted submissions in its ten-reviewer candidate cell.
No unknown-commit failure was observed. Summary documents remained at most 1,982 bytes in this
small corpus. The artifact's no_delta_saves field counts successful source saves without a statistics
transition, so it includes every source-only baseline save; it is not a count of unchanged source data.
Required continuation #3255 owns the bounded write-path performance and contention improvement before claiming activation readiness. Increasing retries alone would trade exhausted requests for more load and latency; it does not establish the required overhead threshold. Re-run this matrix after the chosen change and add the required reviewer-family and active-tracking variants as their writers become available. The present slice does not silently expand maintenance or alter the existing three-attempt source contract to make its numbers pass.
Dependency integration measurement: 13 September 2026¶
After merging main and reviewer-maintenance prerequisite #3297 at 0803a4ec5, the opt-in smoke run
(3 measured batches, 1 warmup) passed. A separate 30-batch artifact
then ran all eight cells with five warmup batches per variant against disposable MongoDB 8.0.28
on .NET 10.0.0. Each cell still has one project and ten studies, and only ProjectScreening is enabled.
The measured total is 1,080 submissions per variant; warmup is excluded. Every source/projection,
receipt/delta and outbox correctness assertion passed. No unknown-commit failure was observed.
All eight overhead gates failed again. Thirty batches are a modest local measurement, not a large-project, current-controller, staging or production acceptance result. The historical 100-iteration artifact above remains unchanged. This run records that dependency revision; it does not establish a comparable host-controlled trend against September 5.
| Target | Reviewers | Baseline p95 ms | Materialized p95 ms | Overhead | Baseline saved | Materialized saved |
|---|---|---|---|---|---|---|
| Same study | 1 | 11.10 | 84.44 | 660.79% | 30/30 | 30/30 |
| Same study | 2 | 24.68 | 57.83 | 134.28% | 60/60 | 36/60 |
| Same study | 5 | 26.49 | 52.90 | 99.70% | 90/150 | 46/150 |
| Same study | 10 | 30.73 | 60.30 | 96.22% | 97/300 | 44/300 |
| Different studies | 1 | 8.40 | 62.83 | 647.97% | 30/30 | 30/30 |
| Different studies | 2 | 8.46 | 93.91 | 1010.16% | 60/60 | 60/60 |
| Different studies | 5 | 11.36 | 96.04 | 745.78% | 150/150 | 105/150 |
| Different studies | 10 | 14.60 | 86.96 | 495.69% | 300/300 | 176/300 |
Baseline exhausted 263 of 1,080 submissions after 1,005 total optimistic conflicts; materialized exhausted 553 after 1,883 conflicts. Exhaustion is a failed save, never counted as success. During measured submissions the driver recorded 4,649 baseline commands versus 28,104 materialized commands, including 527 materialized commit attempts. Those command counts exclude assertion queries and seeding. The higher command count and contention are observed diagnostic evidence; this slice introduces no unmeasured partitioning or write-protocol changes to chase the threshold.
Reproduce this sample with the command above using SYRF_STATS_ITERATIONS=30 and
SYRF_STATS_WARMUP=5. Focused integration validation additionally passed 67 tests covering durable
mode saves, optimistic contention and reviewer materialization; with no dataset variable the benchmark
was skipped as intended. Performance follow-up remains #3255, and rollout stays gated by the failed
write-overhead criterion and the separate live acceptance requirements.
The later integration of #3297 at 266a8dac3 includes Sonar fixes, legacy null-list recovery and a real Project-document admission write for reviewer maintenance. The recorded artifact predates those fixes. Reviewer maintenance remains disabled in this benchmark; measuring its additional write and contention requires a separately enabled workload before rollout acceptance. No fresh performance measurement is claimed by this integration.
Paired run 2026-09-15 (loaded host)¶
main at 0fe1716f6 and #3475 at 6546bb25d were run back to back, same host, same minute,
SYRF_STATS_ITERATIONS=30 SYRF_STATS_WARMUP=5 SYRF_STATS_COUNT_ROUNDTRIPS=1, MongoDB 8.0.28 on
.NET 10.0.0. Artifacts:
main arm,
#3475 arm.
Both the pre-#3475 main arm and the post-#3475 candidate arm ran on the same loaded host; neither
was an idle-host fixture. Load average was 20-45 with roughly 130 CI containers running throughout
both arms. This measurement therefore fails the hygiene rule this programme set for itself — "run
unchanged main and the candidate back to back on the same idle host" — in one specific respect:
the p95 columns below are indicative only and no ratio derived from them may be quoted. The
command, conflict, exhaustion and saved counts are deterministic work counts rather than timings and
are the reliable part of this run.
| Cell | Commands (main → #3475) | Conflicts | Exhausted | Saved | Materialised p95 ms | Source-only p95 ms |
|---|---|---|---|---|---|---|
| same_study 1 | 900 → 630 | 0 → 0 | 0 → 0 | 30 → 30 | 118.72 → 56.53 | 58.55 → 12.62 |
| same_study 2 | 1390 → 1090 | 82 → 80 | 25 → 25 | 35 → 35 | 95.67 → 53.62 | 13.29 → 17.18 |
| same_study 5 | 2812 → 2553 | 343 → 343 | 111 → 111 | 39 → 39 | 84.33 → 56.52 | 15.73 → 26.41 |
| same_study 10 | 5196 → 4974 | 779 → 770 | 251 → 246 | 49 → 54 | 58.99 → 71.41 | 41.99 → 30.72 |
| different_studies 1 | 900 → 630 | 0 → 0 | 0 → 0 | 30 → 30 | 47.91 → 42.95 | 6.44 → 8.68 |
| different_studies 2 | 2612 → 1860 | 58 → 60 | 0 → 0 | 60 → 60 | 103.98 → 75.32 | 38.86 → 5.90 |
| different_studies 5 | 5090 → 3779 | 193 → 193 | 44 → 43 | 106 → 107 | 103.06 → 93.50 | 7.31 → 5.85 |
| different_studies 10 | 9296 → 7142 | 434 → 428 | 125 → 118 | 175 → 182 | 94.07 → 98.10 | 9.94 → 10.15 |
What this establishes¶
- The command reductions are real and match the unit measurement. The uncontended cells fall
900 -> 630 over 30 submissions, i.e. 30 -> 21 commands per save, exactly the figure
ProjectScreeningWriteRoundTripBudgetTestsasserts. The candidate artifact's per-command-name breakdown for that cell isfind 390, update 120, insert 60, findAndModify 30, commitTransaction 30— 13 finds, 4 updates, 2 inserts, 1 findAndModify and 1 commit per save — and its per-attempt note records 16 commands inside the repository save, leaving 5 before the transaction. - The contended cells drop by more in absolute terms, because a conflicting attempt also does less wasted work: 2612 -> 1860 at two different-study reviewers, 5090 -> 3779 at five, 9296 -> 7142 at ten.
What this does not establish¶
- A4's contention effect is not demonstrated. Conflicts and exhaustion are essentially unchanged
(58->60 / 0->0; 193->193 / 44->43; 434->428 / 125->118). Ordering the contended writes first was
expected to cut the work a loser does, not the number of losers, and that is what the counts
show — a loser still loses. The small movements in
saved(175 -> 182 at ten different-study reviewers, 49 -> 54 at ten same-study) are within the noise of a loaded host and are not claimed as an improvement. - No latency claim. Materialised p95 is lower in six of the eight cells and higher in two, but the
source-only p95 of the same cell moved by as much as 6.6x between arms (
different_studies 2: 38.86 ms on main against 5.90 ms on the candidate) purely from host load. That is why the artifacts' own*_p95_gateoverhead percentages are not reproduced here: under this load they measure the scheduler, not the change. - Every one of the eight <10% overhead gates still fails, in both arms. Nothing here moves the performance acceptance gate.
At the time, an idle-host re-run was owed before any rollout claim; it was run on 2026-09-30 (see Idle-host run 2026-09-30 below) and the gate still fails. The defensible statement about this change remains the one the command counts support: nine fewer round trips per conflict-free save, and proportionally less wasted work per losing attempt.
Idle-host run 2026-09-30¶
main at 65ae1099b, MongoDB 8.0.28 on .NET 10.0.0, PS-WRITE-01 at SYRF_STATS_ITERATIONS=100
SYRF_STATS_WARMUP=10, exactly the Reproduce command above. Three sequential runs from one scratch
worktree, one build:
- (a) timing, no command counting: artifact.
- (b) command count,
SYRF_STATS_COUNT_ROUNDTRIPS=1: artifact. - © repeat of (a), for run-to-run variance: artifact.
Reviewer maintenance is not an arm and there is no env toggle for it. The harness fixes
ActiveReviewerTrackingEnabled = false and toggles only MaterializedProjectStatisticsWrites between
the two arms, so no run of this benchmark measures enabled reviewer maintenance. STATUS asks for it;
that measurement still needs a separately enabled workload and is not claimed here.
Host conditions¶
The machine is also the CI runner host (48 CPUs). The doc's idle threshold is load per CPU at most 0.5.
| Run | Load average at start | Peak 1-minute load during/after | Peak load per CPU | Docker containers running |
|---|---|---|---|---|
| (a) timing | 4.42 | 7.88 | 0.16 | 56-58 |
| (b) counting | 7.53 | 10.75 | 0.22 | 57-58 |
| © repeat | 7.78 | 14.85 | 0.31 | 57 |
Load stayed below the threshold in every snapshot, so no waiting was needed. About 57 CI containers were present throughout but mostly idle; the load-average figures are the evidence, not a claim of an empty host. The 15 September runs above had load 20-45.
Results (run a)¶
Overhead is materialized p95 over source-only p95 minus one. The gate is strictly under 10%.
| Target | Reviewers | Source-only p50 ms | Source-only p95 ms | Materialized p50 ms | Materialized p95 ms | Overhead | Gate |
|---|---|---|---|---|---|---|---|
| Same study | 1 | 4.08 | 7.76 | 37.51 | 61.21 | 689% | FAIL |
| Same study | 2 | 5.75 | 11.12 | 31.40 | 43.85 | 294% | FAIL |
| Same study | 5 | 10.99 | 19.07 | 25.63 | 43.95 | 131% | FAIL |
| Same study | 10 | 18.87 | 30.20 | 28.80 | 44.24 | 46% | FAIL |
| Different studies | 1 | 4.79 | 5.82 | 35.40 | 39.31 | 575% | FAIL |
| Different studies | 2 | 4.90 | 5.94 | 42.82 | 81.20 | 1268% | FAIL |
| Different studies | 5 | 4.70 | 5.92 | 43.89 | 81.84 | 1283% | FAIL |
| Different studies | 10 | 5.15 | 13.88 | 43.54 | 84.29 | 507% | FAIL |
Run-to-run variance (overhead per cell in runs a, c and the counting run b, whose subscriber adds its own asymmetric cost): same study 689/611/596, 294/293/289, 131/128/116, 46/51/53; different studies 575/662/766, 1268/1426/1438, 1283/1366/1167, 507/1088/570. The verdict is stable; the individual percentages are not, especially the source-only p95 in the low-millisecond cells (a 1 ms shift in a 5-6 ms baseline moves the ratio by 15-20%). Do not quote a single percentage to more than order of magnitude.
Outcomes and command counts¶
Saved, conflicts and exhaustion are from run a; commands are from the counting run b (100 batches, so submissions per cell are 100 times the reviewer count).
| Target | Reviewers | Source-only saved | Materialized saved | Source-only conflicts / exhausted | Materialized conflicts / exhausted | Source-only commands (per submission) | Materialized commands (per submission) |
|---|---|---|---|---|---|---|---|
| Same study | 1 | 100/100 | 100/100 | 0 / 0 | 0 / 0 | 200 (2.0) | 2500 (25.0) |
| Same study | 2 | 200/200 | 116/200 | 100 / 0 | 268 / 84 | 700 (3.5) | 4143 (20.7) |
| Same study | 5 | 302/500 | 142/500 | 896 / 198 | 1125 / 358 | 3300 (6.6) | 9172 (18.3) |
| Same study | 10 | 310/1000 | 151/1000 | 2376 / 690 | 2612 / 849 | 7724 (7.7) | 17669 (17.7) |
| Different studies | 1 | 100/100 | 100/100 | 0 / 0 | 0 / 0 | 200 (2.0) | 2500 (25.0) |
| Different studies | 2 | 200/200 | 200/200 | 0 / 0 | 198 / 0 | 400 (2.0) | 7600 (38.0) |
| Different studies | 5 | 500/500 | 350/500 | 0 / 0 | 646 / 150 | 1000 (2.0) | 15513 (31.0) |
| Different studies | 10 | 1000/1000 | 601/1000 | 0 / 0 | 1401 / 399 | 2000 (2.0) | 28782 (28.8) |
No unknown-commit failure occurred and every parity, receipt and outbox assertion passed. Exhaustion is a failed save; a lower materialized p95 in a contended cell does not offset its lower success rate.
Verdict: FAIL, on an idle host¶
All eight cells fail the under-10% p95 gate, with overhead from 46% (ten reviewers, same study) to 1,283% (five reviewers, different studies). The host met the doc's idle definition, so unlike the 15 September paired run these timings are acceptance-grade evidence, and they say the gate is not met. Only the single-reviewer cells are free of contention; they fail by 575-689%, so the cost is structural per-save work and not contention alone.
Different-study cells regress the most in relative terms: source-only saves never conflict there (0 conflicts at every reviewer count), while the materialized arm exhausts 150 of 500 submissions at five reviewers and 399 of 1,000 at ten because every save contends on the shared per-project statistics rows, which drops materialized success to 60% of submissions at ten reviewers.
Cost drivers¶
Uncontended (one reviewer), the materialized save issues 25 commands against 2 (find +
findAndModify) source-only. Names are from the counting run's breakdown; the collection order comes
from ProjectScreeningWriteRoundTripBudgetTests output at the same commit.
| Command | Source-only per save | Materialized per save | Added |
|---|---|---|---|
find |
1 | 16 | +15 |
update |
0 | 4 | +4 |
insert |
0 | 2 | +2 |
aggregate |
0 | 1 | +1 |
commitTransaction |
0 | 1 | +1 |
findAndModify |
1 | 1 | 0 |
This is 25 commands per save, up from the 21 recorded for #3475 on 15 September: the four added
commands are the bounded ledger admission probes (aggregate and find on
pmProjectStatisticsRevision) and the two pmProjectStatisticsReceiptRetirement reads (receipt
retirement guards), all merged after that measurement. In the table these four are the aggregate
row's +1 and three of the find row's +15 (one revision probe plus the two retirement reads);
the rest of the table was already present in the 21-command count. Under contention the materialized arm also
issues one abortTransaction per lost attempt (for example 1,396 aborts at ten different-study
reviewers), each preceded by the full read set again.
Top three added command types per save: find (+15), update (+4), insert (+2). The
15 extra reads dominate: the traced order shows pmProjectStatisticsGlobalControl read 3 times
(pre-transaction, envelope authority capture, in-transaction), pmProjectStatisticsControl,
pmProjectStatisticsSourceOperationReceipt, pmProjectStatisticsReceiptRetirement and
pmProjectStatisticsPublicationGuard each read twice (once for the envelope/pre-check, once again inside
the transaction), plus two pmProjectStatisticsRevision reads.
Optimisation candidates (not implemented)¶
Each is grounded in the traced sequence above and needs its own correctness review, since the duplicate reads exist to close races between the pre-check and the transaction.
- Read
pmProjectStatisticsGlobalControlonce per submission, not three times. Hold the snapshot captured inBeginOperationAsyncand re-validate it inside the transaction with the write's own version guard (or one read whose result the transaction reuses) instead of three separate finds. - Fold the pre-transaction duplicate pre-check into the transaction.
ResolveByReceiptAsyncand receipt-retirement preflight read the receipt and retirement collections, and the transaction reads both again. A single in-transaction read, with the duplicate result returned from the transaction, removes roughly four finds per save at the cost of opening a transaction for a duplicate. - Combine the Control and PublicationGuard checks into the same read as the receipt admission
(one query against a single guard document per project, or one
$lookup) instead of one find per guard family. - Merge the two
updatewrites on Control and PublicationGuard, and the read-then-updateon theNotificationOutbox, into a single upsert (the outbox already uses a coalesced slot). The 4 updates and 2 inserts are the writes the transaction has to keep, so the saving here is the paired read, not the write. - Address contention separately from round trips. In the different-study cells every save writes
the same
pmProjectStatisticsCurrentand Control rows, so conflicts and exhaustion (399 of 1,000 at ten reviewers) do not fall with command count. Options: shard the statistics row by study bucket, or accumulate deltas in the already-written revision/ledger rows and fold them asynchronously. This is the design question behind #3255 and #3251, not a local tweak.
Even a large round-trip reduction is unlikely to reach under 10% while a save must run a multi-document
transaction against a source path that is a single findAndModify; the gate as written may need Chris's
explicit decision, as STATUS already notes.
Limits¶
- One project, ten studies, capacity-save path only. Not the current controller path, not HTTP, not staging or production.
- Reviewer maintenance and active-reviewer tracking are off; their overhead remains unmeasured.
- The source-only p95 in low-millisecond cells is noisy, so percentages differ by up to 2x between identical runs (see above). The pass/fail verdict does not.
Fold arm run 2026-10-02 (slice 6, gate b)¶
Async point-fold slice 6 (design, section 9): the
three-arm matrix with ReviewEligibilityPolicy off and on, at 75498052f on #3949 (stacked on the
slice-6a P0 branch, with slices 3, 4 and 5 merged), MongoDB 8.0.28 on .NET 10.0.0, PS-WRITE-01 at
SYRF_STATS_ITERATIONS=100 SYRF_STATS_WARMUP=10, exactly the Reproduce command above, one build:
- (a) timing, no command counting: artifact.
- (b) command count,
SYRF_STATS_COUNT_ROUNDTRIPS=1: artifact. - © repeat of (a), for run-to-run variance: artifact.
Every run passed every per-batch check: exact persisted decisions; for the transactional arm, parity, receipts, deltas and the outbox as before; for the fold arm, acceptance 3 after every batch, warmup included (overlay before the drain = authoritative = the legacy FullStats ProjectScreening section; pending entries = the batch's entry-carrying saves; the real worker drains; stored row = authoritative; one receipt per consumed entry; no pending set left).
Host conditions: not acceptance-grade timing¶
The machine is the CI runner host (48 CPUs); the threshold is a one-minute load per CPU of at most 0.5 (load 24). From 22:22 to 00:00 UTC the host was polled every 90 seconds and each run was started the moment load fell to 24 or below, but CI bursts arrived within one to three minutes of every start:
| Run | Load at start | 1-minute load sampled each minute during the run | Peak per CPU | Load after | Containers |
|---|---|---|---|---|---|
| (a) timing | 23.85 | 37.03, 30.67, 39.95, 38.39, 28.81, 16.29, 31.19 | 0.89 | 42.74 | 73-78 |
| (b) counting | 27.25 | 39.70, 25.79, 17.58, 18.03, 22.17, 26.57, 31.79 | 0.83 | 30.52 | 78-83 |
| © repeat | 14.77 | 17.82, 20.35, 46.86, 50.60, 41.57, 52.81, 46.69 | 1.10 | 30.02 | 77-82 |
The timings below are therefore not acceptance evidence: load exceeded the threshold for most of every run. They are reported because the verdicts were identical in (a) and © in fifteen of sixteen cells and the failing margins are mostly several milliseconds, not noise-sized. The deterministic results (saved, exhausted, conflict attribution, command counts and the per-batch exactness checks) do not depend on host load and are reliable.
Two earlier attempts at a777e0021 failed after seven minutes with Connection refused because there
are too many open connections: a client per arm left about 13 connections open on the server after
disposal, and the forty arms of the new matrix reached mongod's limit. The harness now shares one
instrumented client across arms (6ff78055c); the whole matrix peaks at 33 open connections.
Results (run a, gate b per cell)¶
The Verdict column applies the gate to loaded-host timings, so it is provisional (see
the provisional verdict).
| Target | Reviewers | Eligibility | Source-only p50 / p95 ms | Fold p50 / p95 ms | Overhead p95 | Latency (<10% or <=2 ms) | Exhausted fold / source | Fold saved / expected | Statistics-caused | Verdict |
|---|---|---|---|---|---|---|---|---|---|---|
| Same study | 1 | off | 7.03 / 8.19 | 7.37 / 8.55 | +4.4% (+0.36 ms) | pass | 0 / 0 | 100 / 100 | 0 | PASS |
| Same study | 1 | on | 7.42 / 12.04 | 13.41 / 16.57 | +37.7% (+4.54 ms) | FAIL | 0 / 0 | 100 / 100 | 0 | FAIL |
| Same study | 2 | off | 7.51 / 12.12 | 11.08 / 20.17 | +66.4% (+8.05 ms) | FAIL | 0 / 0 | 200 / 200 | 0 | FAIL |
| Same study | 2 | on | 16.45 / 27.93 | 15.60 / 38.84 | +39.1% (+10.91 ms) | FAIL | 0 / 0 | 200 / 200 | 0 | FAIL |
| Same study | 5 | off | 16.11 / 23.75 | 26.53 / 42.26 | +77.9% (+18.51 ms) | FAIL | 196 / 200 | 304 / 300 | 0 | FAIL |
| Same study | 5 | on | 26.25 / 35.45 | 33.38 / 46.72 | +31.8% (+11.27 ms) | FAIL | 261 / 277 | 239 / 223 | 0 | FAIL |
| Same study | 10 | off | 18.37 / 28.86 | 34.31 / 49.90 | +72.9% (+21.04 ms) | FAIL | 632 / 677 | 368 / 323 | 0 | FAIL |
| Same study | 10 | on | 31.14 / 43.51 | 44.84 / 53.87 | +23.8% (+10.36 ms) | FAIL | 722 / 769 | 278 / 231 | 0 | FAIL |
| Different studies | 1 | off | 5.14 / 6.07 | 6.95 / 8.01 | +32.0% (+1.94 ms) | pass | 0 / 0 | 100 / 100 | 0 | PASS |
| Different studies | 1 | on | 10.74 / 13.05 | 13.36 / 16.02 | +22.8% (+2.98 ms) | FAIL | 0 / 0 | 100 / 100 | 0 | FAIL |
| Different studies | 2 | off | 5.09 / 6.48 | 7.21 / 9.11 | +40.6% (+2.63 ms) | FAIL | 0 / 0 | 200 / 200 | 0 | FAIL |
| Different studies | 2 | on | 16.33 / 20.98 | 15.78 / 28.17 | +34.3% (+7.19 ms) | FAIL | 0 / 0 | 200 / 200 | 0 | FAIL |
| Different studies | 5 | off | 4.93 / 5.97 | 6.94 / 13.12 | +119.7% (+7.15 ms) | FAIL | 0 / 0 | 500 / 500 | 0 | FAIL |
| Different studies | 5 | on | 17.51 / 34.43 | 31.08 / 44.60 | +29.5% (+10.17 ms) | FAIL | 200 / 199 | 300 / 301 | 0 | FAIL |
| Different studies | 10 | off | 5.07 / 7.64 | 7.04 / 15.48 | +102.7% (+7.84 ms) | FAIL | 0 / 0 | 1000 / 1000 | 0 | FAIL |
| Different studies | 10 | on | 28.95 / 41.42 | 43.45 / 50.14 | +21.1% (+8.72 ms) | FAIL | 688 / 700 | 312 / 300 | 0 | FAIL |
Exhausted and Fold saved are the literal gate (b) equalities. Where they fail, the fold arm
exhausted fewer submissions than source-only (for example 632 against 677 at ten same-Study
reviewers) in every cell but one, where it exhausted one more (200 against 199 at five different-Study
reviewers with eligibility on, where the Project token serializes every save). Contended exhaustion is
timing-dependent and differs run to run in both arms.
No fold-arm conflict was statistics-caused in any cell. No conflict was fold-caused either, because the
worker drains between batches, not during them; the free fold-only retry is covered by
ProjectScreeningFoldEndToEndTests.AFoldBetweenTheLoadAndTheWriteIsAFreeRetry.
Run © agrees: PASS in same_study_1 (eligibility off) and FAIL in every other cell; the one
difference is different_studies_1 (eligibility off), +1.94 ms in (a) and +2.00 ms in ©, against a
+2 ms allowance.
Transactional arm, for reference (run a, eligibility off)¶
| Target | Reviewers | Source-only p95 ms | Transactional p95 ms | Transactional saved | Historical gate |
|---|---|---|---|---|---|
| Same study | 1 | 8.19 | 43.22 | 100/100 | 428% FAIL |
| Same study | 2 | 12.12 | 47.52 | 122/200 | 292% FAIL |
| Same study | 5 | 23.75 | 43.41 | 136/500 | 83% FAIL |
| Same study | 10 | 28.86 | 45.85 | 156/1000 | 59% FAIL |
| Different studies | 1 | 6.07 | 38.79 | 100/100 | 539% FAIL |
| Different studies | 2 | 6.48 | 87.61 | 200/200 | 1252% FAIL |
| Different studies | 5 | 5.97 | 86.10 | 350/500 | 1342% FAIL |
| Different studies | 10 | 7.64 | 83.09 | 600/1000 | 988% FAIL |
Outcomes and conflicts (run a)¶
Conflicts are attributed per attempt from the reload the save already performs. Source: Study means the Study's version moved by a write that was not a fold; Source: Project means the Study did not move, so the eligibility transaction lost on the Project (its admission token or the first-review marker). Fold-caused would be a reload whose version moved only by folds. Statistics-caused counts statistics write rejections in the fold arm, and for the transactional arm a reload where the Study did not move. Unattributed is a third-attempt conflict with no reload in the timed path.
| Target | Reviewers | Eligibility | Arm | Saved | Exhausted | Conflicts | Source: Study | Source: Project | Fold-caused | Statistics-caused | Unattributed |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Same study | 1 | off | SourceOnly | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Same study | 1 | off | Fold | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Same study | 1 | off | Transactional | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Same study | 1 | on | SourceOnly | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Same study | 1 | on | Fold | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Same study | 2 | off | SourceOnly | 200 | 0 | 100 | 100 | 0 | 0 | 0 | 0 |
| Same study | 2 | off | Fold | 200 | 0 | 100 | 100 | 0 | 0 | 0 | 0 |
| Same study | 2 | off | Transactional | 122 | 78 | 256 | 22 | 0 | 0 | 156 | 78 |
| Same study | 2 | on | SourceOnly | 200 | 0 | 132 | 100 | 32 | 0 | 0 | 0 |
| Same study | 2 | on | Fold | 200 | 0 | 144 | 100 | 44 | 0 | 0 | 0 |
| Same study | 5 | off | SourceOnly | 300 | 200 | 900 | 900 | 0 | 0 | 0 | 0 |
| Same study | 5 | off | Fold | 304 | 196 | 896 | 896 | 0 | 0 | 0 | 0 |
| Same study | 5 | off | Transactional | 136 | 364 | 1134 | 134 | 0 | 0 | 636 | 364 |
| Same study | 5 | on | SourceOnly | 223 | 277 | 1028 | 489 | 262 | 0 | 0 | 277 |
| Same study | 5 | on | Fold | 239 | 261 | 985 | 707 | 278 | 0 | 0 | 0 |
| Same study | 10 | off | SourceOnly | 323 | 677 | 2336 | 2336 | 0 | 0 | 0 | 0 |
| Same study | 10 | off | Fold | 368 | 632 | 2215 | 2215 | 0 | 0 | 0 | 0 |
| Same study | 10 | off | Transactional | 156 | 844 | 2600 | 477 | 0 | 0 | 1279 | 844 |
| Same study | 10 | on | SourceOnly | 231 | 769 | 2505 | 1169 | 567 | 0 | 0 | 769 |
| Same study | 10 | on | Fold | 278 | 722 | 2423 | 1863 | 560 | 0 | 0 | 0 |
| Different studies | 1 | off | SourceOnly | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 1 | off | Fold | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 1 | off | Transactional | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 1 | on | SourceOnly | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 1 | on | Fold | 100 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 2 | off | SourceOnly | 200 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 2 | off | Fold | 200 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 2 | off | Transactional | 200 | 0 | 200 | 0 | 0 | 0 | 200 | 0 |
| Different studies | 2 | on | SourceOnly | 200 | 0 | 100 | 0 | 100 | 0 | 0 | 0 |
| Different studies | 2 | on | Fold | 200 | 0 | 100 | 0 | 100 | 0 | 0 | 0 |
| Different studies | 5 | off | SourceOnly | 500 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 5 | off | Fold | 500 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 5 | off | Transactional | 350 | 150 | 650 | 0 | 0 | 0 | 500 | 150 |
| Different studies | 5 | on | SourceOnly | 301 | 199 | 899 | 0 | 700 | 0 | 0 | 199 |
| Different studies | 5 | on | Fold | 300 | 200 | 900 | 0 | 900 | 0 | 0 | 0 |
| Different studies | 10 | off | SourceOnly | 1000 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 10 | off | Fold | 1000 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Different studies | 10 | off | Transactional | 600 | 400 | 1398 | 0 | 0 | 0 | 998 | 400 |
| Different studies | 10 | on | SourceOnly | 300 | 700 | 2400 | 0 | 1700 | 0 | 0 | 700 |
| Different studies | 10 | on | Fold | 312 | 688 | 2379 | 0 | 2379 | 0 | 0 | 0 |
With eligibility on, the source-only arm itself exhausts 199 of 500 submissions at five different-Study reviewers and 700 of 1,000 at ten: every eligibility save writes the Project's admission token, so different Studies contend on one Project document. That is the source-level contention the design records as deferred item 2; the fold arm meets the same contention and adds none.
Command counts (run b)¶
| Target | Reviewers | Eligibility | Source-only per submission | Fold per submission | Transactional per submission | Fold breakdown |
|---|---|---|---|---|---|---|
| Same study | 1 | off | 2.00 | 5.00 | 25.00 | find 400, findAndModify 100 |
| Same study | 1 | on | 7.00 | 10.00 | n/a | commitTransaction 100, find 700, update 200 |
| Same study | 2 | off | 3.50 | 9.50 | 20.44 | find 1600, findAndModify 300 |
| Same study | 2 | on | 10.67 | 17.34 | n/a | abortTransaction 141, commitTransaction 200, find 2587, update 541 |
| Same study | 5 | off | 6.60 | 18.37 | 18.39 | find 7986, findAndModify 1199 |
| Same study | 5 | on | 13.60 | 25.04 | n/a | abortTransaction 988, commitTransaction 234, find 9843, update 1456 |
| Same study | 10 | off | 7.71 | 20.58 | 17.65 | find 17992, findAndModify 2591 |
| Same study | 10 | on | 14.25 | 27.70 | n/a | abortTransaction 2443, commitTransaction 264, find 22027, update 2971 |
| Different studies | 1 | off | 2.00 | 5.00 | 25.00 | find 400, findAndModify 100 |
| Different studies | 1 | on | 7.00 | 10.00 | n/a | commitTransaction 100, find 700, update 200 |
| Different studies | 2 | off | 2.00 | 5.00 | 38.00 | find 800, findAndModify 200 |
| Different studies | 2 | on | 9.50 | 14.00 | n/a | abortTransaction 100, commitTransaction 200, find 2000, update 500 |
| Different studies | 5 | off | 2.00 | 5.00 | 31.10 | find 2000, findAndModify 500 |
| Different studies | 5 | on | 13.20 | 20.80 | n/a | abortTransaction 900, commitTransaction 300, find 7700, update 1500 |
| Different studies | 10 | off | 2.00 | 5.00 | 29.04 | find 4000, findAndModify 1000 |
| Different studies | 10 | on | 14.11 | 22.88 | n/a | abortTransaction 2391, commitTransaction 306, find 17182, update 3003 |
Per submission, uncontended (one reviewer): source-only find + findAndModify = 2; fold = the
private Study find, the durable-mode find and the control find concurrently, then the pre-write
durable-mode re-read find, then findAndModify = 5, in three sequential round trips against two;
transactional = 25. With eligibility on: 7 against 10 (the global and control reads, and the
durable-mode read inside the eligibility transaction). Under contention every fold retry adds the
reload, the durable mode, the control, the receipt and two quarantine reads and the pre-write re-read
before its write (8 commands, 4 sequential round trips against the source-only retry's 2), which is
why the fold arm's contended cells cost as many commands per submission as the transactional arm.
Command budget per save shape (gate b acceptance 1)¶
FoldSaveCommandBudgetTests and FoldSaveEligibilityCommandBudgetTests (API tests, the real
ReviewScreeningFoldTarget, SubmitAnnotationSessionService and AnnotationDeletionFoldTarget) and
ProjectStatisticsFoldCommandBudgetTests (Mongo.Data tests, the real repository) run in ordinary CI on
a real replica set. Every shape asserts no write to any pmProjectStatistics* collection, the same
number of commitTransaction and Project admission-token writes as the same call with statistics and
the fold off, and the count below. A positive control asserts that today's transactional screening save
does trip the statistics-write detector (24 commands, one commit).
The table is the current code (updated 2026-10-04, after the owner decisions recorded in the gate (b) re-baseline); the run (b) figures above were measured before it and keep the older shapes.
| Save shape | Design table, first attempt | As built, first attempt | Before 2026-10-04 | Statistics-off today | Difference from the design |
|---|---|---|---|---|---|
| Capacity-guarded or plain screening | 4 | 4 (2 sequential rounds) | 5 (3 rounds) | 2 | None. The pre-write durable-mode re-read was dropped by owner decision (d), revised |
| Screening charged retry (all MVP families requested) | 6 | 8 for the retry (2 rounds); the lost plain attempt has no tail | 10 (5 rounds), after a 6-command attempt one that ended with the bulk-update-lock probe | 3 (diagnosis, reload, write; 3 rounds) | The screening save carries its annotation half as a second namespace, so the receipt and quarantine reads are per namespace. With one namespace (this benchmark's ProjectScreening-only posture) the retry is 6: reload, mode, control, receipt, quarantine in one concurrent round, then the write |
ReviewEligibilityPolicy screening |
the above + Project + eligibility transaction | 9 | 10 | 7 | None: the fold path adds only its two concurrent control-plane reads; the eligibility transaction no longer re-reads the mode for a fold-path write |
| Annotation session save, untracked | 4 | 4 | 5 | n/a | None |
| Annotation session save, tracked | 4 in the presence transaction | 5 + commit | 5 + commit | same commit | The capacity write inside today's presence transaction keeps its in-transaction mode read; the one commit is today's presence transaction |
| Annotation session delete | 4 | 5 | 6 | n/a | The deletion target's Project read (cached in the host; one find on a miss) |
| Reservation release | 4 | 5 | 5 | 2 | The Project read the entry's context digest is computed from (cached in the host; one find on a miss) |
| Reservation claim | 3, or 4 without the caller's Project | 3 / 4 | 3 / 4 | 1 | None |
| Eligibility admission | today's transaction + global + control | today + 2 | today + 2 | 5 | None |
The capacity-guarded screening shape keeps its diagnosis read after a lost write (it tells "at capacity" from a stale copy), so its charged retry is three sequential rounds rather than two.
Red first: the screening, release and retry assertions were first written at the design table's
counts and failed with the traces above (for example Expected commands to contain 4 item(s) ... but
found 5), then set to the as-built counts with the reason recorded beside each.
Provisional verdict, loaded host (not acceptance evidence)¶
Latency misses gate (b) on these timings; statistics cause no conflicts. This is not the gate (b) verdict. The timings were taken on a loaded host (host conditions), so gate (b) stays open until an idle-host rerun (planned on Bramble). Only the deterministic results below (conflict attribution, saved and exhausted counts, command counts and the per-batch exactness checks) are final.
- Zero statistics-caused conflicts and failures in every cell, both eligibility settings, every run. No fold-arm save failed for a statistics reason.
- Latency misses the limits in fourteen of sixteen cells in run (a) (fifteen in run ©), on a host that was not idle; provisional until the idle-host rerun. The uncontended fold save adds about 2 ms at p50 and p95 (one sequential round trip plus a slightly larger majority-acknowledged write). Contended cells add 7-21 ms at p95, because a fold retry costs four sequential round trips against two.
- Exhaustion equality fails literally in the contended cells: the fold arm exhausted fewer submissions than source-only in all but one of them (one more, 200 against 199).
Cost drivers and the deferred optimisation¶
Drivers 1 and 2 were acted on after the idle-host run: see Gate (b) re-baselined 2026-10-04. The analysis below is as of 2026-10-02.
- The pre-write durable-mode re-read in
SaveScreeningOnFoldPathAsync(and the annotation and deletion writes that share it) is the only extra sequential round trip on an uncontended save. The design (section 2, "Durable reviewer-mode check") accepts that the mode check is not atomic with the write, and invariant 13 makes a mode-transition owner waitFoldPreWriteDeadline+ the writer grace before treating the fleet as transitioned. Whether that wait makes the re-read redundant is a safety decision for the design owner, not a benchmark change. - Retry cost. Every charged retry runs the stored checks (receipt and quarantine, per namespace, plus the drop ranges) and the re-read, which dominates the contended cells.
- Write size. Both writes are a
findAndModifythat returns the document. The fold write also carries the pending set (a larger document both ways) and an explicit majority write concern, which on this replica set matches the server's implicit default, so this is a minor term.
Deferred item 3 (coalescing the global and project
control reads into one $unionWith aggregate, or moving FoldMode onto the Project) was not
applied. Both reads already run concurrently with the private Study load, so coalescing them removes
one command but no sequential round trip; it cannot close a gap of about one round trip, and it would
cross the boundary between the save (which reads the control) and the writer target (which reads the
mode). Driver 1 is the lever that matches the measurement.
Limits¶
- Host load exceeded the threshold during every run, so no timing here is acceptance-grade.
- One project, ten studies; reviewer-tracking off (the capacity filter is not exercised by the timed fold write; the command-budget tests cover the tracked shape). Not HTTP, staging or production.
- The fold worker drains between batches only, so fold-caused conflicts and free retries are not exercised by this benchmark (they are covered by the fold end-to-end tests).
- The fold target mirrors
ReviewScreeningFoldTargetcommand for command (the benchmark project cannot reference the API host); the API command-budget tests run the real target.
Idle-host fold run 2026-10-03 (Bramble)¶
The idle-host rerun the slice-6 verdict was waiting for. Bramble (20 CPUs), main at c45648aeb,
MongoDB 8.0.20 on .NET 10.0.12, PS-WRITE-01 at SYRF_STATS_ITERATIONS=100 SYRF_STATS_WARMUP=10,
exactly the Reproduce command, three runs back to back:
- run 1, timing: artifact.
- run 2, timing repeat: artifact.
- counted,
SYRF_STATS_COUNT_ROUNDTRIPS=1: artifact.
The one-minute load stayed between 2.8 and 5.4 (at most 0.27 per CPU, under the 0.5 threshold), so these timings are acceptance-grade. Every per-batch exactness check passed in every run.
Verdict under the gate as then defined: FAIL. One cell passed in runs 1 and 2
(same_study_1_reviewers), two in the counted run. Every cell had zero statistics-caused and zero
fold-caused conflicts. The failures were of two kinds:
- Latency. The uncontended eligibility-off fold save cost about +1.2-2.2 ms at the mean, most of it
the sequential pre-write durable-mode re-read (one
find, about 1.3 ms here). With eligibility on, the extra sequential read was the durable-mode read inside the eligibility transaction. Under same-study contention the retry dominated: 55-70% of the overhead at 5 and 10 reviewers, because a charged fold retry cost 9 commands in 5 sequential rounds against source-only's 3 in 3. - Exhaustion equality.
fold_exhausted_equals_source_onlyandfold_saved_matchesfailed in every contended cell with five or more reviewers, in both directions, for reasons that were not statistics-caused: free retries never fired (fold_conflicts=0), so both arms raced the same three-attempt budget, and the same arm's exhaustion varied run to run by as much as the arms differed (for example 199 to 157 between runs of one arm).
The contended fold windows also opened new pooled connections (isMaster/saslContinue) while timed,
which explains fold maxima of 116-147 ms against 34-55 ms for source-only.
Gate (b) re-baselined 2026-10-04¶
The design owner decided three changes in one PR (ADR-019 amendment; design decision (d), revised):
- The pre-write durable-mode re-read is dropped. The attempt reads the project's
FoldModeand the durable reviewer mode once, among its concurrent reads. Fold-path eligibility writes skip the mode read inside the eligibility transaction too. - A charged retry is two sequential rounds. The reload runs concurrently with the control-plane reads, the reload's own lock field replaces the separate bulk-update-lock probe, and the unrecorded-drop ranges are read only after an unknown write result. When the reload shows that a fold moved the Study in between, the control plane is read again after the reload, so a fold racing those concurrent reads can never hide a committed id. The command budget table has the new counts.
- The gate is re-baselined, in
ScreeningWriteGateB(pinned byScreeningWriteGateBTestsin the ordinary test lane):
| Clause | Cells | Rule |
|---|---|---|
| Statistics-caused conflicts | every cell, stress cells included | = 0. The run fails after writing its artifact otherwise. |
| Statistics-caused exhaustion | every cell | = 0. An exhausted submission counts only when its final attempt failed on a statistics-caused conflict (for the fold arm, a fold that moved the Study) or a statistics write rejection. Replaces the literal fold exhausted == source_only exhausted and fold saved == submissions - source_only exhausted clauses, which are still reported, as information. |
| Latency | every different-study cell, and the 1-reviewer same-study cells | p95 overhead under 10% or at most +2 ms. Superseded on 2026-10-05 by an absolute budget judged on the median of five runs (below). |
| Latency | same-study cells with 2, 5 and 10 reviewers | Reported as STRESS figures, no PASS/FAIL. Every reviewer edits one Study in lockstep for 100 batches, so these cells measure retry cost under contention that a real project does not sustain, not save cost. |
The harness also warms the connection pool before the first timed window: the shared client keeps a
minimum pool of 64 connections and rounds of concurrent pings open them first
(connection_pool_warmup note), and each arm's _outcomes note reports connections_opened inside its
timed windows. The artifact schema is additive: every earlier key keeps its meaning, and the gate note
appends gate_definition, latency_gated, statistics_caused_exhausted and statistics_clean, with
gate_b_summary tallying the run.
The acceptance rerun¶
Superseded on 2026-10-05: the rerun below was run (see
the 2026-10-05 rerun) and the gate was then redefined; the
current recipe is the acceptance run. As
defined on 2026-10-04: the rerun happens on Bramble after this change merges, not on the CI host. Use the Reproduce
command twice for timing and once with SYRF_STATS_COUNT_ROUNDTRIPS=1, at a load of at most 0.5 per
CPU. Gate (b) passes when, in both timing runs:
gate_b_summaryshowsfailed=0andstatistics_unclean_cells=0;- every latency-gated cell's note shows
verdict=PASS; - the counted run shows the new shapes: 4 commands per uncontended eligibility-off fold submission, 9 with eligibility on, and 6 per charged retry in this ProjectScreening-only posture;
connections_opened=0(or close to it) inside the timed windows of the contended fold arms.
The stress cells' figures are recorded in the results, not judged.
Gate (b) redefined 2026-10-05: absolute budget, median of five runs¶
The 2026-10-05 acceptance rerun (Bramble)¶
The rerun of the acceptance rerun recipe, on Bramble (20 CPUs), main at
5245941e9 (which includes #4011), MongoDB 8.0.20 on .NET 10.0.12, PS-WRITE-01 at
SYRF_STATS_ITERATIONS=100 SYRF_STATS_WARMUP=10, three runs back to back from 00:37 to 01:01 UTC:
- run 1, timing: artifact.
- run 2, timing repeat: artifact.
- counted,
SYRF_STATS_COUNT_ROUNDTRIPS=1: artifact. - The operator's host-load log: the 1-minute load was between 3.06 and 7.38 (at most 0.37 per CPU, under the 0.5 limit) and every run exited 0, so every per-batch exactness check passed.
What the 2026-10-04 changes were for held in every run:
- Zero statistics-caused conflicts and zero statistics-caused exhaustion in every cell, stress cells
included (
statistics_unclean_cells=0in all threegate_b_summarynotes). - The new command shapes in the counted run: 4 commands per uncontended eligibility-off fold submission
(
find3,findAndModify1), 9 with eligibility on, and 6 per charged retry (same Study, two reviewers: 1,400 commands = 200 submissions × 4 + 100 retries × 6). connections_opened=0inside every timed window.
Latency under the 2026-10-04 clause (under 10% or at most +2 ms) did not settle. The per-run tallies passed 4 of 10 gated cells in run 1, 6 in run 2 and 4 in the counted run, and runs 1 and 2 agreed in only 6 of the 10 cells, because the same cell's p95 overhead moved by several milliseconds between runs of one build: at ten reviewers on different Studies, +7.76 ms (48.7%) in run 1 and +1.67 ms (10.0%) in run 2 with eligibility off, +10.66 ms and +25.82 ms with it on. The source-only p95 in the gated cells was about 10-47 ms.
The decision¶
Chris decided on 2026-10-05 (ADR-019 amendment; design decision (k)) to replace the relative limit with an absolute, user-facing budget, and to judge it on enough data that noise cannot flip a verdict:
| Clause | Cells | Rule |
|---|---|---|
| Statistics-caused conflicts | every cell of every run (timing and counted), stress cells included | = 0 (unchanged). A run fails after writing its artifact otherwise. |
| Statistics-caused exhaustion | every cell of every run | = 0 (unchanged; attribution as on 2026-10-04). |
| Latency | every different-study cell and the 1-reviewer same-study cells, eligibility off and on (10 cells) | The median across five timing runs of the cell's p95 overhead (fold p95 minus source-only p95, same cell, same run) is at most 15 ms. Minimum and maximum across runs are reported, not judged. |
| Latency | same-study cells with 2, 5 and 10 reviewers | STRESS: reported, never judged (unchanged). |
| Evidence | At least five timing runs with command counting off, each of at least 300 measured iterations per cell (warm-up 10, as before); one counted run; every artifact from the same dataset, commit, MongoDB server, .NET runtime and host. | |
| Command shapes | the counted run | 4 per uncontended fold save, 9 with eligibility on, 6 per charged retry (unchanged). Judged on the typical shape of each cell: a single-reviewer cell records the commands of every attempt, and attempts that differ from the median attempt are set aside when isolated (at most 2% of the attempts) and reported as a warning with the mean; a consistent change, or outliers above 2%, fail as before. |
| Host | the operator's log | The 1-minute load at each run's <label> BEFORE stamp (taken once the host was idle) is at most 0.5 per CPU, every run (the five timing runs and the counted run) has such a stamp, and every run exited 0. The load sampled while a run is in progress or after it is reported (peak and p95), never judged. |
Why: run-to-run p95 noise of several milliseconds in one cell makes a 2 ms (or 10%) rule unjudgeable; saves take about 10-47 ms at p95 in the gated cells, so up to 15 ms extra is not noticeable to a reviewer, while a save failed by statistics would be, so that clause stays at zero.
In the artifact (the schema stays syrf.project-statistics.baseline/v1; notes are additive):
<cell>_fold_gate_bkeeps every earlier field in its earlier order.gate_definition=2026-10-05;passes_latencyandverdictfollow the 15 ms budget for that one run; the note appendslatency_budget_ms=15.00andsuperseded_2026_10_04_latency_pass(the old relative-or-2 ms clause, as information).gate_b_summaryappendslatency_budget_msandacceptance=median of five timing runs.- A new
source_commitnote carriesSYRF_STATS_SOURCE_COMMIT(unrecordedwhen unset). ScreeningWriteGateBTestspins the budget's boundary (15.00 ms passes, 15.01 ms fails, an exact 15 ms whose binary difference is 15.000000000000002 passes), the stress cells and the statistics clause.
A single run's verdict is not the gate. The gate is the aggregate:
scripts/aggregate-statistics-write-benchmark.py reads the five timing artifacts, the counted artifact and
the host-load log, refuses artifacts that do not share dataset, commit, server, runtime and host, and prints
(and with --json-out/--markdown-out writes) a per-cell table of median, minimum and maximum p95 overhead
in milliseconds and percent, the statistics check of every cell of every run, the command shapes and the
host load, with an overall PASS or FAIL and every reason. Exit code 0 is PASS, 1 is FAIL, 2 means
the artifacts were refused. Its tests (python3 scripts/test-aggregate-statistics-write-benchmark.py, in the
Test statistics evidence scripts workflow) use reduced copies of the 2026-10-05 artifacts.
The acceptance run: five timing runs and one counted run¶
Run on Bramble, never on the CI host, from a clean checkout of main at the commit under test, when
the 1-minute load is at most 0.5 per CPU (10 on Bramble's 20 CPUs); the script waits for that before each
run and records the accepted reading as the run's BEFORE stamp, and a minute-by-minute sampler records the
load throughout as information. One build serves all six runs. At 100 iterations a run took about eight
minutes, so expect about 25 minutes per run and two and a half hours in all. Save this as gate-b-acceptance.sh and run bash gate-b-acceptance.sh from the checkout's root:
#!/usr/bin/env bash
# Gate (b) acceptance run (definition 2026-10-05): five timing runs of 300 iterations and one counted run.
set -euo pipefail
cd "$(git rev-parse --show-toplevel)"
COMMIT=$(git rev-parse HEAD)
OUT=${OUT:-$HOME/scratch/gate-b-$(date -u +%Y%m%dT%H%M%SZ)}
PROJECT=src/libs/project-management/SyRF.ProjectManagement.Mongo.Data.Tests/SyRF.ProjectManagement.Mongo.Data.Tests.csproj
mkdir -p "$OUT"
dotnet build "$PROJECT"
# One line per sample: "<label> <phase> <UTC time> <uptime output> docker_ps=<n>". The aggregator judges the
# 1-minute load on each "<label> BEFORE" line and reads every EXIT= code; the other lines are information.
stamp_line() {
printf '%s %s %s %s docker_ps=%s\n' "$1" "$2" "$(date -u +%FT%TZ)" "$(LC_ALL=C uptime)" "$(docker ps -q | wc -l)"
}
stamp() { stamp_line "$@" >> "$OUT/host-load.log"; }
# Reads a stamp line and succeeds when its 1-minute load is at most 0.5 per CPU.
idle_enough() {
sed -n 's/.*load averages\{0,1\}: \([0-9.]*\).*/\1/p' |
awk -v cpus="$(nproc)" 'BEGIN { ok = 0 } NF { ok = ($1 / cpus <= 0.5) } END { exit !ok }'
}
# Writes "<label> BEFORE" once the host is idle. The stamp is the very reading the wait accepted (it is what the
# aggregator judges), so a run starts only from a stamp at most 0.5 per CPU; neither the build nor the previous
# run's teardown is inside the next run's evidence.
stamp_before_idle() {
local line
until line=$(stamp_line "$1" BEFORE) && idle_enough <<< "$line"; do sleep 30; done
printf '%s\n' "$line" >> "$OUT/host-load.log"
}
( while sleep 60; do stamp sample -; done ) &
SAMPLER=$!
trap 'kill "$SAMPLER" 2>/dev/null || true' EXIT
run() { # $1 = label; $2 = 0 for a timing run, 1 for the counted run
local code=0
stamp_before_idle "$1"
SYRF_STATS_DATASET=PS-WRITE-01 SYRF_STATS_ITERATIONS=300 SYRF_STATS_WARMUP=10 \
SYRF_STATS_COUNT_ROUNDTRIPS="$2" SYRF_STATS_SOURCE_COMMIT="$COMMIT" SYRF_STATS_RESULTS_DIR="$OUT/$1" \
dotnet test "$PROJECT" --no-build --filter FullyQualifiedName~ProjectScreeningWriteBenchmark || code=$?
stamp "$1" "EXIT=$code AFTER"
}
for i in 1 2 3 4 5; do run "run$i" 0; done
run counted 1
python3 scripts/aggregate-statistics-write-benchmark.py "$OUT"/run[1-5]/*.json \
--counted "$OUT"/counted/*.json --host-load "$OUT/host-load.log" --commit "$COMMIT" \
--json-out "$OUT/gate-b-summary.json" --markdown-out "$OUT/gate-b-summary.md"
The script's last command prints the verdict and exits 0 (PASS) or 1 (FAIL); set -e makes the script
exit with it. Gate (b) passes only when that summary says PASS: zero statistics-caused conflicts and
exhaustion in every cell of every run, every gated cell's median p95 overhead at most 15 ms, the typical
command shapes 4 / 9 / 6, the load at every run's BEFORE stamp at most 0.5 per CPU, every run exited 0, and
five 300-iteration timing runs of one commit. The summary also reports the peak and p95 load while the runs
were in progress, and any outlier attempts it set aside; neither fails the gate. Record the result here with
the six artifacts, the host-load log and the summary copied under evidence/source-writes/, as for the earlier
runs (copy the log with a .txt extension: the repository ignores *.log), and on
#3510.
The first five-run acceptance run, 2026-10-05, and the checker's correction¶
The first run of the recipe above (Bramble, 20 CPUs, main at 7f6490a71, MongoDB 8.0.20, .NET 10.0.12, five
timing runs of 300 iterations and one counted run, 02:55 to 04:55 UTC; output kept on Bramble in
~/scratch/gate-b-20261005T025530Z, not committed because the rerun below replaces it) passed everything gate
(b) is about and still returned FAIL, for two measurement artefacts of the checker:
- Latency and statistics. Every latency-gated cell was within the budget at the median across the five runs; the largest was ten reviewers on different Studies with eligibility on, +13.49 ms (runs +12.40 to +22.46 ms) against 15 ms. Zero statistics-caused conflicts and zero statistics-caused exhaustion in every cell of all six runs.
- Host load. The checker judged every sample of the load log against 0.5 per CPU. The only breach was
11.41 on 20 CPUs (0.57 per CPU), the sample taken at 04:35:29, 12 seconds after the counted run's
BEFOREstamp; the run's ownBEFOREstamp, taken 04:35:17 once the host was idle, read 9.70 (0.49 per CPU). The load was the benchmark's start-up (the test host and a fresh MongoDB container), not external load. The rule is now judged on theBEFOREstamps (the host is idle before a run, which is what matters for the timings); the load during and after the runs is reported as a peak and a p95, and never fails. A run without aBEFOREstamp now fails, so a log with no stamps cannot pass vacuously. - Command shapes. In the counted run the cell for one reviewer on different Studies recorded 1,203
commands for 300 fold submissions (4.01 each, against exactly 4), and its per-attempt series
(
different_studies_1_reviewers_fold_commands_per_attempt) had two attempts with 3 and 2 commands instead of - The artifact's
RoundTripsrecord for that cell names them:find900,findAndModify300 and, besides,isMaster1 andsaslContinue2, the handshake and authentication of pooled connections the driver opened inside the timed window (the same cell'sconnections_opened=2); no other counted cell had any. The fold arm sent exactly the 4 commands per submission it should. Two corrections: the benchmark's counter no longer counts driver handshake and monitoring commands (ScreeningWriteCommandCountTests), and the checker judges the typical shape, setting aside isolated outlier attempts (at most 2% of a single-reviewer cell's attempts) with a warning that records them and the mean, while a consistent change such as 5 commands per submission, or outliers above 2%, still fail. The scripts' tests pin both directions.
The gate definition and every threshold are unchanged (15 ms median budget, zero statistics-caused conflicts
and exhaustion, 4 / 9 / 6 commands). The corrected checker judges the unchanged first-run data PASS (exit 0,
with the outlier attempts as a warning), but the owner's decision is to rerun the acceptance run on the
corrected harness and checker; that rerun decides gate (b).
The 2026-10-05 runs judged by the 2026-10-05 rule¶
Not acceptance evidence: they are two timing runs of 100 iterations, not five of 300. Reproduce from the repository root:
EVIDENCE=docs/features/materialized-project-statistics/evidence/source-writes
python3 scripts/aggregate-statistics-write-benchmark.py \
$EVIDENCE/PS-WRITE-01_ScreeningSourceWrites_2026-10-05_bramble-run1.json \
$EVIDENCE/PS-WRITE-01_ScreeningSourceWrites_2026-10-05_bramble-run2.json \
--counted $EVIDENCE/PS-WRITE-01_ScreeningSourceWrites_2026-10-05_bramble-counted.json \
--host-load $EVIDENCE/PS-WRITE-01_ScreeningSourceWrites_2026-10-05_bramble-host-load.txt \
--commit 5245941e9
Verdict: FAIL, exit 1, for two evidence reasons and one latency reason: two timing runs (the rule needs five), 100 iterations each (it needs 300), and ten reviewers on different Studies with eligibility on at a median of 18.24 ms. Statistics are clean in every cell of all three runs, the command shapes are 4 / 9 / 6 and the load before each run peaked at 0.32 per CPU (0.37 during and after the runs). With two runs the median is their mean, so that cell's verdict is decided by one run's +25.82 ms against the other's +10.66 ms; this is the noise the five-run median exists to absorb. Judged one run at a time against 15 ms, run 1 passes all ten gated cells and run 2 passes nine.
| Cell | Gated | Median p95 overhead ms | Min ms | Max ms | Median % | Min % | Max % | Verdict |
|---|---|---|---|---|---|---|---|---|
| Same study, 1 reviewer | yes | 0.32 | -0.49 | 1.12 | 3.2 | -4.0 | 10.3 | PASS |
| Same study, 1 reviewer, eligibility | yes | 2.65 | 0.73 | 4.57 | 16.6 | 4.4 | 28.8 | PASS |
| Same study, 2 reviewers | stress | 5.28 | 5.23 | 5.32 | 31.6 | 31.1 | 32.1 | STRESS |
| Same study, 2 reviewers, eligibility | stress | 6.10 | 5.85 | 6.35 | 24.6 | 24.0 | 25.3 | STRESS |
| Same study, 5 reviewers | stress | 8.41 | 7.78 | 9.04 | 29.4 | 29.0 | 29.9 | STRESS |
| Same study, 5 reviewers, eligibility | stress | 7.40 | 5.46 | 9.35 | 17.1 | 12.7 | 21.4 | STRESS |
| Same study, 10 reviewers | stress | 17.37 | 13.43 | 21.32 | 55.4 | 42.5 | 68.3 | STRESS |
| Same study, 10 reviewers, eligibility | stress | 17.23 | 14.25 | 20.22 | 35.4 | 29.0 | 41.8 | STRESS |
| Different studies, 1 reviewer | yes | -1.34 | -2.56 | -0.12 | -13.3 | -25.3 | -1.3 | PASS |
| Different studies, 1 reviewer, eligibility | yes | 1.35 | 1.16 | 1.54 | 8.4 | 7.3 | 9.4 | PASS |
| Different studies, 2 reviewers | yes | 3.00 | 1.51 | 4.49 | 27.7 | 14.6 | 40.8 | PASS |
| Different studies, 2 reviewers, eligibility | yes | 4.51 | 3.24 | 5.79 | 18.4 | 12.6 | 24.1 | PASS |
| Different studies, 5 reviewers | yes | 2.96 | 1.44 | 4.49 | 19.6 | 9.4 | 29.8 | PASS |
| Different studies, 5 reviewers, eligibility | yes | 7.03 | 6.17 | 7.89 | 17.0 | 15.8 | 18.3 | PASS |
| Different studies, 10 reviewers | yes | 4.71 | 1.67 | 7.76 | 29.3 | 10.0 | 48.7 | PASS |
| Different studies, 10 reviewers, eligibility | yes | 18.24 | 10.66 | 25.82 | 40.0 | 22.5 | 57.5 | FAIL |
Gate (b) therefore stays open until the five-run acceptance run above; production enable stays refused in code until it passes and a production rollout is separately approved.