feat(nwsync): emit blobs in parallel with a bounded worker pool #80

Merged
archvillainette merged 1 commits from feat/79-nwsync-parallel-blob-upload into main 2026-07-31 15:25:34 +00:00
Owner

Closes #79.

Emit is latency-bound, not CPU-bound. Every blob costs two serial round-trips to the zone — a ProbeKey HEAD, then a PutReader PUT — so a hak with a few thousand resources pays a few thousand serialised latencies. Measured on the live sow-assets-manifest backfill: 26 s of CPU across 9.5 minutes of wall clock, on a 4-core host with 5 GB free and peak RSS of 51 MB.

emit now hashes, compresses and stores --jobs N resources at once, default 16 — matching DEPOT_JOBS and the transport's MaxIdleConnsPerHost, so a worker per connection needs no fresh TLS handshake. --jobs 1 is exactly the old behaviour.

Three properties had to survive

Each has a test in internal/nwsync/jobs_test.go:

  • Deterministic manifest bytes. emitterVersion promises a manifest is a function of its artifact, so entries is index-addressed rather than appended to — a worker owns entries[i] alone and the slice comes back in artifact order whatever order uploads finish in. TestEmitProducesTheSameIndexAtEveryJobCount diffs the .nsym and its sidecar between -jobs 1 and -jobs 16.
  • Index still lands last. Any worker's failure aborts before a manifest is written. TestEmitLeavesNoIndexWhenAParallelUploadFails fails every blob PUT with 16 workers in flight and asserts no .nsym appears. Under -race it also covers the shared counters.
  • Identical content still shares one blob. This one bit during development and is the reason to read the diff carefully: serially, the sink's existence check absorbed two resrefs with identical bytes. In parallel both workers probe, both miss, and both upload — TestEmitWritesBlobsAndManifest caught it as "wrote 2 blobs, want 1". Claiming the sha1 in-process restores the dedupe and skips a probe round-trip as well.

Memory

Peak now tracks the resources in flight rather than one resource. The ceiling is N × the 15 MB fileSizeLimit plus its compressed copy — bounded by a constant this package enforces itself, and still not tracking the archive. TestEmitPeakMemoryIsBoundedByJobCount re-runs the #76 regression check at -jobs 8: a hak 8× bigger still costs the same.

This is only cheap because of #78. Before streaming emit, N workers would have meant N whole archives resident.

Not done

Skipping the ProbeKey HEAD on a first-time emit would halve round-trips, but doubles uploaded bytes on a re-run — which is exactly what a backfill is. Noted in #79 so it is not rediscovered; parallelism is the better lever and this PR takes it.

Worker compression still serialises on blobEncoder, which is WithEncoderConcurrency(1) for the memory reason in compressedbuf.go. At 26 s of CPU per hak that is not worth trading memory for, but it is where to look if the numbers ever say otherwise.

🤖 Generated with Claude Code

Closes #79. Emit is latency-bound, not CPU-bound. Every blob costs two serial round-trips to the zone — a `ProbeKey` HEAD, then a `PutReader` PUT — so a hak with a few thousand resources pays a few thousand serialised latencies. Measured on the live sow-assets-manifest backfill: 26 s of CPU across 9.5 minutes of wall clock, on a 4-core host with 5 GB free and peak RSS of 51 MB. `emit` now hashes, compresses and stores `--jobs N` resources at once, default 16 — matching `DEPOT_JOBS` and the transport's `MaxIdleConnsPerHost`, so a worker per connection needs no fresh TLS handshake. `--jobs 1` is exactly the old behaviour. ### Three properties had to survive Each has a test in `internal/nwsync/jobs_test.go`: - **Deterministic manifest bytes.** `emitterVersion` promises a manifest is a function of its artifact, so `entries` is index-addressed rather than appended to — a worker owns `entries[i]` alone and the slice comes back in artifact order whatever order uploads finish in. `TestEmitProducesTheSameIndexAtEveryJobCount` diffs the `.nsym` and its sidecar between `-jobs 1` and `-jobs 16`. - **Index still lands last.** Any worker's failure aborts before a manifest is written. `TestEmitLeavesNoIndexWhenAParallelUploadFails` fails every blob PUT with 16 workers in flight and asserts no `.nsym` appears. Under `-race` it also covers the shared counters. - **Identical content still shares one blob.** This one bit during development and is the reason to read the diff carefully: serially, the sink's existence check absorbed two resrefs with identical bytes. In parallel both workers probe, both miss, and both upload — `TestEmitWritesBlobsAndManifest` caught it as "wrote 2 blobs, want 1". Claiming the sha1 in-process restores the dedupe and skips a probe round-trip as well. ### Memory Peak now tracks the resources in flight rather than one resource. The ceiling is `N` × the 15 MB `fileSizeLimit` plus its compressed copy — bounded by a constant this package enforces itself, and still not tracking the archive. `TestEmitPeakMemoryIsBoundedByJobCount` re-runs the #76 regression check at `-jobs 8`: a hak 8× bigger still costs the same. This is only cheap because of #78. Before streaming emit, N workers would have meant N whole archives resident. ### Not done Skipping the `ProbeKey` HEAD on a first-time emit would halve round-trips, but doubles uploaded bytes on a re-run — which is exactly what a backfill is. Noted in #79 so it is not rediscovered; parallelism is the better lever and this PR takes it. Worker compression still serialises on `blobEncoder`, which is `WithEncoderConcurrency(1)` for the memory reason in `compressedbuf.go`. At 26 s of CPU per hak that is not worth trading memory for, but it is where to look if the numbers ever say otherwise. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
archvillainette added 1 commit 2026-07-31 14:29:10 +00:00
Emit is latency-bound, not CPU-bound. Every blob costs two serial HTTP
round-trips to the zone — an existence probe, then an upload — so a hak
with a few thousand resources pays a few thousand serialised latencies.
A measured backfill spent 26 seconds of CPU across 9.5 minutes of wall
clock, on a host with three of four cores idle and 5 GB free.

Emit now hashes, compresses and stores `--jobs N` resources at once
(default 16, matching DEPOT_JOBS and the transport's idle connections
per host).

Three properties had to survive, and each has a test:

- The manifest's bytes are promised deterministic by emitterVersion, so
  entries is index-addressed rather than appended to: a worker owns
  entries[i] alone and the slice comes back in artifact order whatever
  order the uploads finish in.
- The index is still the publication marker, so any worker's failure
  aborts the run before a manifest is written. Workers drain the rest of
  the channel instead of returning, which keeps the feeder from blocking
  on workers that have gone away.
- Two resrefs holding identical bytes still share one blob. Serially the
  sink's existence check absorbed that; in parallel both workers would
  probe, both miss, and both upload. Claiming the sha1 in-process
  restores the dedupe and skips a probe round-trip as well.

Peak memory is now the resources in flight rather than one resource, so
the ceiling is N times the 15 MB per-resource limit and its compressed
copy — bounded by a constant this package enforces itself, and still
nowhere near tracking the archive.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
archvillainette scheduled this pull request to auto merge when all checks succeed 2026-07-31 15:24:56 +00:00
xtul approved these changes 2026-07-31 15:25:32 +00:00
archvillainette merged commit 1c2acc5530 into main 2026-07-31 15:25:34 +00:00
archvillainette deleted branch feat/79-nwsync-parallel-blob-upload 2026-07-31 15:25:34 +00:00
Sign in to join this conversation.