Compare commits

..
4 Commits
Author SHA1 Message Date
archvillainette fa32dd411f fix(nwsync): stream emit so peak memory tracks the largest resource (#76) (#78)
build-binaries / build-binaries (push) Successful in 2m13s
Closes #76. Part of #54.

`nwsync emit` held roughly 5× the artifact size in RAM, so ovh-main (7 GB, no swap) OOM-killed it on any hak over ~1.4 GB. That blocked the backfill in sow-assets-manifest and would have killed the next release rebuilding a large hak.

## What it does

- `erf.ReadIndex` / `erf.ReadPayload` — parse the header and resource table only, read one payload on demand. `erf.Read` keeps its shape but returns payloads as subslices of the buffer instead of fresh copies, which removes one full copy for the `pipeline` callers too. Callers must not mutate `Resource.Data`; the doc comment says so and no caller does.
- `nwsync.Emit` opens the artifact, hashes it by streaming for the key check (through a section reader, so the file offset stays put), then hashes, compresses and stores one resource at a time. The archive is never resident. Shadowed duplicate resrefs are now never read at all.
- Bounds checks in `ReadIndex` moved to `int64`, so key/resource-list offsets can no longer overflow.

## Not in the issue, but memory-motivated

The zstd blob encoder ran at the default concurrency, which is one encoder per CPU, each holding a window-sized history — about 200 MB of live heap doing nothing on a 24-core runner. `EncodeAll` is single-threaded per call, so concurrency 1 costs nothing. `TestSingleThreadedEncoderMatchesDefault` pins the claim that blobs come out byte-identical.

## Measured

Peak heap during emit, sampled 1 ms:

| hak | before | after |
|-----|--------|-------|
| 8 MB | 155 MB | 21 MB |
| 64 MB | 289 MB | 22 MB |

Flat, as the acceptance asks. `TestEmitPeakMemoryDoesNotScaleWithArtifactSize` fails if the 64 MB fixture costs more than the 8 MB one plus 24 MB of slack.

## Gaps

- The regression check measures Go heap, not RSS, and its largest fixture is 64 MB — a multi-GB run was not done here. A 2.15 GB hak now needs about the same ~22 MB the 64 MB one does, so the 5 GB budget is not close, but that is inference from the flat curve, not a measurement.
- mmap was suggested in the issue and skipped. Payload buffers are still anonymous, but they are one resource each (≤15 MiB), so making them file-backed buys nothing now.

🤖 Generated with [Claude Code](https://claude.com/claude-code)Reviewed-on: #78
Reviewed-by: xtul <mpiasecki720@protonmail.com>
Co-authored-by: vickydotbat <vickydotbat@tutamail.com>
2026-07-31 11:16:23 +00:00
archvillainette 7437653f14 docs: describe the on-disk shape of an emitted NWSync tree (#77)
Documents what `nwsync emit --out DIR` actually produces, so a zone can be checked by hand.

The trap this removes: a blob's filename is the SHA-1 of the resource's *original* bytes, but the file on disk is NWCompressedBuffer framing — a 24-byte `NSYC` header then a zstd frame. `sha1sum <blob>` therefore never matches the name it is sitting under. The doc records the recipe that does match:

```
tail -c +25 <blob> | zstd -dc | sha1sum
```

Also notes the tree layout, that the uncompressed length is a little-endian `uint32` at offset 12, and the rough 4:1 compression ratio on hak content.

Verified two ways: against a real emit of a 250 MB hak (2296 blobs, 59 MB on disk against 249 MB of resources), and against `internal/nwsync/compressedbuf.go`, where the header is written as `[]uint32{blobMagic, blobVersion, algorithmZstd, uint32(len(data)), zstdHeaderVer, zstdDictionary}` — confirming field 3 at offset 12.

Docs only, no behaviour change. Falls out of the investigation in #76; that fix is not in this PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)Reviewed-on: #77
Reviewed-by: xtul <mpiasecki720@protonmail.com>
Co-authored-by: vickydotbat <vickydotbat@tutamail.com>
2026-07-31 11:00:49 +00:00
archvillainette f9051a2c01 nwsync: upload sink, key-addressed CLI, fail-closed publication marker (#73)
build-binaries / build-binaries (push) Successful in 2m18s
Builds the code half of #60, and closes the CLI gap the #53 sweep found: PR #71 shipped #53's original surface rather than the one #56, #62 and #65 settled.

## What lands

**`depot.KeyStore`** — `ProbeKey` / `PutReader` / `GetKey`, addressing the zone by object key rather than by depot sha. NWSync cannot use the sha-addressed path: a blob is named after the sha1 of its *uncompressed* bytes while the body uploaded is the compressed form, and Bunny's `Checksum` header is sha256 of the body. Per #55 this reuses `internal/depot`'s `httpBackend` — same IPv4-pinned transport, same retry, same tri-state probe — and the sha-addressed `Backend` is now rewritten on top of it. No second HTTP client, no per-instance hash-function fields: the caller passes the key and the checksum, which turned out simpler than #55 expected.

**A sink in `internal/nwsync`** — the zone by default, a local tree under `--out DIR` as the conformance path. Blobs upload as they are produced and the index lands last, so the presence of an index is the publication marker. A blob already in the zone is skipped via #55's probe *without* paying for compression (the body is a thunk) — which matters for the backfill, where compression is the expensive part.

**The settled CLI**

```
nwsync emit     [--as NAME] [--out DIR] <artifact-key> <file>
nwsync assemble --group-id N [--tlk-key KEY] [--out DIR] <artifact-key>...
```

Artifact keys are depot keys; an index lives beside its artifact with the extension replaced (#62), derived in exactly one place so `emit` and `assemble` cannot disagree. Flags may now follow positionals — Go's `flag` stops at the first non-flag argument, which cost a run during #59.

**Fail-closed in two places** — an artifact key whose embedded digest does not match the file is refused (publishing an index under the wrong key silently pairs a manifest with the wrong artifact), and `assemble` refuses an artifact with no index rather than publishing a manifest missing a hak.

## Checks

`make check` green. Six new tests run against a Bunny-shaped `httptest` zone that verifies the `Checksum` header the way Bunny does: blobs-then-index ordering, skip-if-present, no index after a failed upload, key/file mismatch, and assemble reading indexes back out of the zone.

Conformance re-run through the new CLI against upstream `nwn_nwsync_write` 2.1.2 over `sow_vfxs_01.hak` (#59's oracle): the manifest is still **byte-identical**.

## Not in this PR

- **Live upload against the real zone.** The nwsync zone and its credential are #61, still open. Everything here is proven against a fake zone only.
- **Consumer wiring** — #65, in the three producer repos.
- **The mid-hak failure *policy*.** The mechanism is here (fail closed, orphan blobs left, re-run resumes); whether a module release may proceed when an emit failed is a human call, still open on #60.

🤖 Generated with [Claude Code](https://claude.com/claude-code)Reviewed-on: #73

Co-authored-by: vickydotbat <vickydotbat@tutamail.com>
2026-07-30 21:56:06 +00:00
archvillainette a131b25e5b feat(nwsync): emit blobs and per-artifact NSYM, assemble merged manifests (#71)
Builds sow-tools#53. Format spec followed is the resolution comment of sow-platform#94, checked line by line against niv/neverwinter.nim at HEAD (`nwsync.nim`, `compressedbuf.nim`, `nwsync/private/libupdate.nim`).

## What lands

`crucible nwsync emit <artifact> --out DIR` — explodes one `.hak`/`.erf`, or one loose file such as the TLK, into NWSync blobs plus a NSYM v3 manifest covering only that artifact, with the same `.json` sidecar upstream writes. Blob path `data/sha1/<h0h1>/<h2h3>/<sha1>`, body in NWCompressedBuffer framing (magic `NSYC`, version 3, algorithm 2, uncompressed size, zstd header version 1, dictionary 0, raw zstd frame). The sha1 that names a blob is over the uncompressed bytes.

`crucible nwsync assemble --order NAMES --entries DIR --out DIR [--group-id N]` — merges the per-artifact manifests into one, reading no bulk data at all. Merge rule is resref shadowing, not concatenation: a resref in more than one artifact resolves to the earliest artifact in `--order`, which is how the game resolves it. `--group-id` stays caller-supplied (1 current, 2 testing; 0 is absent, matching upstream omitting a zero integer meta field).

Rules taken from upstream and not re-invented: `nss`/`ndb`/`gic` always skipped; an unresolvable restype is a hard error, not a skip; a resource over 15 MB fails closed; no `latest` file and no `.origin` file, ever. A `.mod` is refused outright — a persistent world publishes no module contents, so the module contributes no bytes.

## Two deliberate departures

- **Emitter version is its own field, not the build revision.** `emitter_version` is a constant bumped only when emitted bytes change. Keying the refuse-to-merge check on `created_with` would invalidate every published index on every unrelated crucible commit and force a re-emit of the whole 15 GB corpus — the opposite of "nothing downstream ever needs the hak again".
- **`SOURCE_DATE_EPOCH` pins the sidecar timestamp.** The manifest itself was already deterministic; the sidecar's `created` was not, against the determinism rule in `docs/consumer-contract.md`.

## Not in this PR, and why

- **Direct upload.** Only the local `--out` sink exists, which is the conformance path. The upload sink and the mid-hak-failure question are sow-tools#60, and the consumer wiring is #65.
- **The conformance run against upstream.** sow-tools#59 owns getting `nwn_nwsync_write` running and capturing reference output. The format here was read from upstream source rather than from its output, so the byte-for-byte manifest comparison and the after-decompression blob comparison still have to happen — that is what #59 is for. The tests in this PR check the layout against the spec, so a shared misreading would pass them.
- **The `artifacts/haks/sha256/<a>/<b>/<sha256>.nsym` location.** emit writes `<out>/<name>.nsym`; where a publisher puts it is the publisher's business (#65).
- **The acceptance gate** — a real client syncing from an assembled manifest — is unchanged and still open.

## Checks

`make check` green (vet, unit tests, shellcheck, yamllint, workflow contract), `make smoke` green with the new builder, `nix build .#crucible` produces `crucible-nwsync`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)Reviewed-on: #71

Co-authored-by: vickydotbat <vickydotbat@tutamail.com>
2026-07-29 10:50:05 +00:00
21 changed files with 2233 additions and 89 deletions
+1
View File
@@ -19,6 +19,7 @@ Crucible is how the artifact repos turn source into artifacts.
| `crucible-depot` | `crucible depot` | content-addressed depot blob verify/move | | `crucible-depot` | `crucible depot` | content-addressed depot blob verify/move |
| `crucible-hak` | `crucible hak` | ERF/HAK pack/unpack + hak manifests | | `crucible-hak` | `crucible hak` | ERF/HAK pack/unpack + hak manifests |
| `crucible-module` | `crucible module` | build/extract/validate/compare the `.mod` | | `crucible-module` | `crucible module` | build/extract/validate/compare the `.mod` |
| `crucible-nwsync` | `crucible nwsync` | NWSync blob emit + manifest assemble |
| `crucible-topdata` | `crucible topdata` | compile 2da/tlk topdata + packages | | `crucible-topdata` | `crucible topdata` | compile 2da/tlk topdata + packages |
| `crucible-wiki` | `crucible wiki` | render + deploy mechanical wiki pages | | `crucible-wiki` | `crucible wiki` | render + deploy mechanical wiki pages |
+11
View File
@@ -0,0 +1,11 @@
// Command crucible-nwsync is the standalone nwsync builder (equivalent to
// `crucible nwsync`). A single-token binary keeps consumer wrapper scripts simple.
package main
import (
"os"
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/dispatch"
)
func main() { os.Exit(dispatch.RunBuilder("nwsync", os.Args[1:])) }
+60
View File
@@ -34,11 +34,71 @@ aliases.
| `depot` | `verify` | Existence sweep plus sampled download re-hash. | | `depot` | `verify` | Existence sweep plus sampled download re-hash. |
| `depot` | `get` | Fetch one blob with sha re-verify. | | `depot` | `get` | Fetch one blob with sha re-verify. |
| `depot` | `pull` | Incremental verified pull of every referenced blob. | | `depot` | `pull` | Incremental verified pull of every referenced blob. |
| `nwsync` | `emit` | Explode one artifact into NWSync blobs plus its own NSYM manifest. |
| `nwsync` | `assemble` | Merge per-artifact NSYM manifests into one merged manifest. |
`depot status` and `depot get` pick their backend either with `--out DIR`, a `depot status` and `depot get` pick their backend either with `--out DIR`, a
depot tree on disk, or with `--target bunny|cdn`, a remote backend. The two depot tree on disk, or with `--target bunny|cdn`, a remote backend. The two
flags are mutually exclusive. flags are mutually exclusive.
`nwsync emit` runs where an artifact is born (a `.hak`/`.erf`, or a loose file
such as the TLK); `nwsync assemble` runs at module release and reads only the
small per-artifact indexes. Both take **depot keys**: an artifact's index lives
beside the artifact itself with the extension replaced, so `emit` and
`assemble` agree on where it is without being told.
```
nwsync emit [--as NAME] [--out DIR] <artifact-key> <file>
nwsync assemble --group-id N [--tlk-key KEY] [--out DIR] <artifact-key>...
```
Both verbs upload by default; nothing bulky is ever written to the runner's
disk. `--out DIR` writes a local repository tree instead, which is the
conformance path against upstream `nwn_nwsync_write`. The zone comes from
`NWSYNC_STORAGE_ZONE` and `NWSYNC_STORAGE_PASSWORD`, with the host from
`BUNNY_STORAGE_HOST` — NWSync data is a separate zone from the asset depot.
`assemble`'s artifact keys are in `Mod_HakList` order, highest priority first: a
resref in more than one artifact resolves to the earliest one, the way the game
resolves it. `--tlk-key` has its own slot because the TLK shadows nothing.
`--group-id` is per channel — 1 is current, 2 is testing, and 0 leaves the field
out of the sidecar.
`emit` uploads blobs first and the index last, so the presence of an index is
the publication marker: an artifact whose emit died halfway leaves real blobs in
the zone and no index. Blob names are content hashes, so re-running skips
whatever already landed, and `assemble` fails closed on an artifact with no
index rather than publishing a manifest that is missing a hak.
### What an emitted tree looks like
`--out DIR` produces the same tree `emit` would upload, which makes it the way
to check a zone by hand without touching one:
```
<artifact-sha>.nsym binary index
<artifact-sha>.nsym.json the same index, readable
data/sha1/a7/4a/a74aa84a... one blob per resource, two-level fanout
```
A blob's name is the SHA-1 of the resource's **original** bytes, but the file on
disk is not those bytes: each blob is wrapped in NWCompressedBuffer framing, a
24-byte `NSYC` header followed by a zstd frame. Hashing the file directly will
not match its name, which is the obvious
first thing to try and the obvious first thing to be confused by. Strip the
header first:
```
tail -c +25 <blob> | zstd -dc | sha1sum # == the blob's filename
```
The header carries the uncompressed length as a little-endian `uint32` at offset
12, so the decompressed size is checkable without decompressing. Compression is
worth roughly a 4:1
saving on hak content: a 250 MB hak emitted 2296 blobs totalling 59 MB on disk
against 249 MB of resources, as recorded in the sidecar's `on_disk_bytes` and
`total_bytes`.
## Hidden compatibility aliases ## Hidden compatibility aliases
Existing scripts may continue using these names indefinitely. They are accepted Existing scripts may continue using these names indefinitely. They are accepted
+2 -1
View File
@@ -22,12 +22,13 @@
pname = "crucible"; pname = "crucible";
inherit version; inherit version;
src = ./.; src = ./.;
vendorHash = "sha256-hm6mrNAtXv0LidzHUfz4eukTFZouizGtxkZ8gKJFUVI="; vendorHash = "sha256-0I8j7On9YGD2GK9xbj/KkgBrlkMJ6Y6XQv+KCLTgBBU=";
subPackages = [ subPackages = [
"cmd/crucible" "cmd/crucible"
"cmd/crucible-depot" "cmd/crucible-depot"
"cmd/crucible-hak" "cmd/crucible-hak"
"cmd/crucible-module" "cmd/crucible-module"
"cmd/crucible-nwsync"
"cmd/crucible-topdata" "cmd/crucible-topdata"
"cmd/crucible-wiki" "cmd/crucible-wiki"
]; ];
+2
View File
@@ -6,3 +6,5 @@ require (
golang.org/x/text v0.35.0 golang.org/x/text v0.35.0
gopkg.in/yaml.v3 v3.0.1 gopkg.in/yaml.v3 v3.0.1
) )
require github.com/klauspost/compress v1.19.1
+2
View File
@@ -1,3 +1,5 @@
github.com/klauspost/compress v1.19.1 h1:VsB4HPswih7mmZ8WleSFQ75c/Ui1M4trX5oAsJnhSlk=
github.com/klauspost/compress v1.19.1/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
golang.org/x/text v0.35.0 h1:JOVx6vVDFokkpaq1AEptVzLTpDe9KGpj5tR4/X+ybL8= golang.org/x/text v0.35.0 h1:JOVx6vVDFokkpaq1AEptVzLTpDe9KGpj5tR4/X+ybL8=
golang.org/x/text v0.35.0/go.mod h1:khi/HExzZJ2pGnjenulevKNX1W67CUy0AsXcNubPGCA= golang.org/x/text v0.35.0/go.mod h1:khi/HExzZJ2pGnjenulevKNX1W67CUy0AsXcNubPGCA=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM= gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM=
+121
View File
@@ -0,0 +1,121 @@
package depot
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"os"
"strings"
)
// KeyStore is the zone addressed by object key rather than by depot sha. The
// depot names every object after the sha256 of its contents; NWSync does not —
// a blob is named after the sha1 of its *uncompressed* bytes while the body
// uploaded is the compressed form, and a per-artifact index is named after its
// artifact. Both addressing modes want the same transport, retry and probe
// discipline, so the sha-addressed Backend rides on this rather than the other
// way round.
type KeyStore interface {
// ProbeKey returns the existence state of one key. transient=true means a
// retry might change the answer — never read it as "missing, re-upload".
ProbeKey(ctx context.Context, key string) (state ProbeState, transient bool, err error)
// PutReader uploads size bytes read from r to key. checksum is the
// uppercase hex sha256 of those bytes, which Bunny verifies server-side.
PutReader(ctx context.Context, key string, r io.Reader, size int64, checksum string) error
// GetKey fetches the whole object at key. Small objects only — it holds
// the body in memory and does no hash check, because a key is not always
// a content hash.
GetKey(ctx context.Context, key string) ([]byte, error)
}
// NewKeyStore returns a KeyStore for cfg's storage zone. Fails closed on a
// missing host or read key, matching NewBackend.
func NewKeyStore(cfg Config) (KeyStore, error) {
if cfg.StorageHost == "" {
return nil, errors.New("storage backend requires a storage host")
}
if cfg.StorageZone == "" {
return nil, errors.New("storage backend requires a storage zone")
}
if cfg.ReadKey == "" {
return nil, errors.New("storage backend requires a read key")
}
return &httpBackend{name: "bunny", client: newHTTPClient(cfg), cfg: cfg}, nil
}
// keyURL is the storage URL of one object key.
func (b *httpBackend) keyURL(key string) string {
host := b.cfg.StorageHost
// StorageHost is normally a bare host ("storage.bunnycdn.com"); allow a
// full scheme (used by tests against httptest.NewServer) to pass through
// unchanged.
if !strings.Contains(host, "://") {
host = "https://" + host
}
return fmt.Sprintf("%s/%s/%s", strings.TrimSuffix(host, "/"), b.cfg.StorageZone, key)
}
func (b *httpBackend) ProbeKey(ctx context.Context, key string) (ProbeState, bool, error) {
return b.rangeProbe(ctx, b.keyURL(key), map[string]string{"AccessKey": b.cfg.ReadKey})
}
func (b *httpBackend) PutReader(ctx context.Context, key string, r io.Reader, size int64, checksum string) error {
if b.name == "cdn" {
return errors.New("cdn backend is read-only")
}
if b.cfg.WriteKey == "" {
return errors.New("storage backend requires a write key to write")
}
req, err := http.NewRequestWithContext(ctx, http.MethodPut, b.keyURL(key), r)
if err != nil {
return err
}
req.ContentLength = size
req.Header.Set("AccessKey", b.cfg.WriteKey)
// Bunny defines Checksum as sha256 of the body and rejects a mismatch, so
// this is server-side integrity checking, not decoration.
req.Header.Set("Checksum", strings.ToUpper(checksum))
resp, err := b.client.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
_, _ = io.Copy(io.Discard, resp.Body)
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return fmt.Errorf("put %s: unexpected status %d", key, resp.StatusCode)
}
return nil
}
func (b *httpBackend) GetKey(ctx context.Context, key string) ([]byte, error) {
resp, err := b.get(ctx, b.keyURL(key), map[string]string{"AccessKey": b.cfg.ReadKey})
if err != nil {
return nil, err
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
_, _ = io.Copy(io.Discard, resp.Body)
return nil, fmt.Errorf("get %s: unexpected status %d", key, resp.StatusCode)
}
return io.ReadAll(resp.Body)
}
// putFile uploads the file at src to key, streaming it. checksum is the
// uppercase hex sha256 of the file's bytes.
func (b *httpBackend) putFile(ctx context.Context, key, src, checksum string) error {
f, err := os.Open(src)
if err != nil {
return err
}
defer f.Close()
info, err := f.Stat()
if err != nil {
return err
}
return b.PutReader(ctx, key, f, info.Size(), checksum)
}
+4 -36
View File
@@ -10,7 +10,6 @@ import (
"net/http" "net/http"
"os" "os"
"path/filepath" "path/filepath"
"strings"
"time" "time"
) )
@@ -71,16 +70,7 @@ type httpBackend struct {
func (b *httpBackend) Name() string { return b.name } func (b *httpBackend) Name() string { return b.name }
func (b *httpBackend) storageURL(sha string) string { func (b *httpBackend) storageURL(sha string) string { return b.keyURL(BlobKey(sha)) }
host := b.cfg.StorageHost
// StorageHost is normally a bare host ("storage.bunnycdn.com"); allow a
// full scheme (used by tests against httptest.NewServer) to pass through
// unchanged.
if strings.Contains(host, "://") {
return fmt.Sprintf("%s/%s/%s", strings.TrimSuffix(host, "/"), b.cfg.StorageZone, BlobKey(sha))
}
return fmt.Sprintf("https://%s/%s/%s", host, b.cfg.StorageZone, BlobKey(sha))
}
func (b *httpBackend) cdnURL(sha string) string { func (b *httpBackend) cdnURL(sha string) string {
return fmt.Sprintf("%s/%s", b.cfg.CDNBase, BlobKey(sha)) return fmt.Sprintf("%s/%s", b.cfg.CDNBase, BlobKey(sha))
@@ -140,7 +130,8 @@ func (b *httpBackend) Probe(ctx context.Context, sha string) (ProbeState, bool,
return b.rangeProbe(ctx, b.storageURL(sha), map[string]string{"AccessKey": b.cfg.ReadKey}) return b.rangeProbe(ctx, b.storageURL(sha), map[string]string{"AccessKey": b.cfg.ReadKey})
} }
// Put uploads src for sha. cdn is read-only. // Put uploads src for sha. cdn is read-only. A depot object is named after the
// sha256 of its own bytes, so the key's sha doubles as the Checksum header.
func (b *httpBackend) Put(ctx context.Context, sha, src string) error { func (b *httpBackend) Put(ctx context.Context, sha, src string) error {
if b.name == "cdn" { if b.name == "cdn" {
return errors.New("cdn backend is read-only") return errors.New("cdn backend is read-only")
@@ -157,30 +148,7 @@ func (b *httpBackend) Put(ctx context.Context, sha, src string) error {
return nil return nil
} }
f, err := os.Open(src) return b.putFile(ctx, BlobKey(sha), src, sha)
if err != nil {
return err
}
defer f.Close()
req, err := http.NewRequestWithContext(ctx, http.MethodPut, b.storageURL(sha), f)
if err != nil {
return err
}
req.Header.Set("AccessKey", b.cfg.WriteKey)
req.Header.Set("Checksum", strings.ToUpper(sha))
resp, err := b.client.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
_, _ = io.Copy(io.Discard, resp.Body)
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return fmt.Errorf("bunny put %s: unexpected status %d", sha, resp.StatusCode)
}
return nil
} }
// Get fetches sha into dest via temp file + rename, re-hashing and deleting // Get fetches sha into dest via temp file + rename, re-hashing and deleting
+15 -2
View File
@@ -21,6 +21,7 @@ import (
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/buildinfo" "git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/buildinfo"
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/depot" "git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/depot"
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/menu" "git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/menu"
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/nwsync"
) )
// Exit codes follow the sysexits(3) convention so CI can distinguish // Exit codes follow the sysexits(3) convention so CI can distinguish
@@ -104,6 +105,16 @@ var Registry = []Builder{
}, },
Wired: true, Wired: true,
}, },
{
Name: "nwsync",
Bin: "crucible-nwsync",
Summary: "publish NWSync blobs and manifests (emit/assemble)",
Commands: []Command{
{Name: "emit", Summary: "explode one artifact into blobs plus its own NSYM manifest", Usage: "crucible nwsync emit <artifact> --out DIR"},
{Name: "assemble", Summary: "merge per-artifact NSYM manifests into one", Usage: "crucible nwsync assemble --order NAMES --entries DIR --out DIR [--group-id N]"},
},
Wired: true,
},
{ {
Name: "assets", Name: "assets",
Bin: "crucible-assets", Bin: "crucible-assets",
@@ -241,7 +252,7 @@ var Registry = []Builder{
"2da-to-module [flags] <input.2da> [output.json]", "2da-to-module [flags] <input.2da> [output.json]",
"json-to-2da <input.json> <output.2da>", "json-to-2da <input.json> <output.2da>",
}, },
Aliases: []CommandAlias{{Name: "convert-topdata"}}, Aliases: []CommandAlias{{Name: "convert-topdata"}},
}, },
}, },
Wired: true, Wired: true,
@@ -392,7 +403,7 @@ func runBuilder(name string, args []string, out, errw io.Writer) int {
} }
} }
if b.Wired { if b.Wired {
// depot and assets are self-contained builders: they parse their own // depot, assets and nwsync are self-contained builders: they parse their own
// subcommands and own their exit contract, bypassing the // subcommands and own their exit contract, bypassing the
// b.Commands/delegateLegacy legacy routing entirely. // b.Commands/delegateLegacy legacy routing entirely.
switch b.Name { switch b.Name {
@@ -400,6 +411,8 @@ func runBuilder(name string, args []string, out, errw io.Writer) int {
return depot.Run(args, out, errw, os.Getenv) return depot.Run(args, out, errw, os.Getenv)
case "assets": case "assets":
return assets.Run(args, out, errw, os.Getenv) return assets.Run(args, out, errw, os.Getenv)
case "nwsync":
return nwsync.Run(args, out, errw)
} }
} }
if !b.Wired { if !b.Wired {
+9 -4
View File
@@ -148,6 +148,10 @@ func TestBuilderHelpIsOK(t *testing.T) {
} }
} }
// selfContained builders parse their own subcommands instead of delegating to
// the legacy internal/app surface.
var selfContained = map[string]bool{"depot": true, "assets": true, "nwsync": true}
func TestCanonicalCommandSurface(t *testing.T) { func TestCanonicalCommandSurface(t *testing.T) {
want := map[string][]string{ want := map[string][]string{
"depot": {"status", "push", "verify", "get", "pull"}, "depot": {"status", "push", "verify", "get", "pull"},
@@ -156,6 +160,7 @@ func TestCanonicalCommandSurface(t *testing.T) {
"module": {"build", "extract", "validate", "compare", "manifest"}, "module": {"build", "extract", "validate", "compare", "manifest"},
"topdata": {"validate", "build", "package", "compare", "convert"}, "topdata": {"validate", "build", "package", "compare", "convert"},
"wiki": {"build", "deploy"}, "wiki": {"build", "deploy"},
"nwsync": {"emit", "assemble"},
} }
for _, builder := range Registry { for _, builder := range Registry {
got := builder.subcommands() got := builder.subcommands()
@@ -173,10 +178,10 @@ func TestRegistryCommandNamesAndAliasesAreUnambiguous(t *testing.T) {
for _, builder := range Registry { for _, builder := range Registry {
seen := map[string]bool{} seen := map[string]bool{}
for _, command := range builder.Commands { for _, command := range builder.Commands {
// depot and assets parse their own subcommands and bypass AppCommand // depot, assets and nwsync parse their own subcommands and bypass
// routing entirely (see the self-contained-builder branch in // AppCommand routing entirely (see the self-contained-builder branch
// runBuilder), so their Commands carry no AppCommand. // in runBuilder), so their Commands carry no AppCommand.
requireAppCommand := builder.Name != "depot" && builder.Name != "assets" requireAppCommand := !selfContained[builder.Name]
if command.Name == "" || command.Summary == "" || command.Usage == "" || (requireAppCommand && command.AppCommand == "") { if command.Name == "" || command.Summary == "" || command.Usage == "" || (requireAppCommand && command.AppCommand == "") {
t.Errorf("%s has incomplete command metadata: %#v", builder.Name, command) t.Errorf("%s has incomplete command metadata: %#v", builder.Name, command)
} }
+92 -45
View File
@@ -311,63 +311,110 @@ func Write(w io.Writer, archive Archive) error {
return nil return nil
} }
// IndexEntry locates one resource inside an archive without holding its
// payload. Streaming callers read one payload at a time from these, so peak
// memory tracks the largest resource instead of the whole archive.
type IndexEntry struct {
Name string
Type uint16
Offset int64
Size int64
}
// Index is the header plus the resource table of an ERF: everything except the
// payloads.
type Index struct {
FileType string
Version string
Entries []IndexEntry
}
// ReadIndex parses the tables of an ERF of the given size, reading only the
// header, the key list and the resource list.
func ReadIndex(r io.ReaderAt, size int64) (Index, error) {
if size < headerSize {
return Index{}, fmt.Errorf("erf file too small: %d bytes", size)
}
var hdr header
if err := binary.Read(io.NewSectionReader(r, 0, headerSize), binary.LittleEndian, &hdr); err != nil {
return Index{}, fmt.Errorf("decode erf header: %w", err)
}
if int64(hdr.KeyListOffset)+int64(hdr.EntryCount)*24 > size {
return Index{}, fmt.Errorf("erf key list exceeds file bounds")
}
keys := make([]keyEntry, hdr.EntryCount)
keyReader := io.NewSectionReader(r, int64(hdr.KeyListOffset), int64(hdr.EntryCount)*24)
if err := binary.Read(keyReader, binary.LittleEndian, &keys); err != nil {
return Index{}, fmt.Errorf("decode key list: %w", err)
}
if int64(hdr.ResourceListOffset)+int64(hdr.EntryCount)*8 > size {
return Index{}, fmt.Errorf("erf resource list exceeds file bounds")
}
entries := make([]resourceEntry, hdr.EntryCount)
entryReader := io.NewSectionReader(r, int64(hdr.ResourceListOffset), int64(hdr.EntryCount)*8)
if err := binary.Read(entryReader, binary.LittleEndian, &entries); err != nil {
return Index{}, fmt.Errorf("decode resource list: %w", err)
}
index := Index{
FileType: string(hdr.FileType[:]),
Version: string(hdr.Version[:]),
Entries: make([]IndexEntry, 0, hdr.EntryCount),
}
for position, key := range keys {
entry := entries[position]
if int64(entry.Offset)+int64(entry.Size) > size {
return Index{}, fmt.Errorf("resource %d exceeds file bounds", position)
}
index.Entries = append(index.Entries, IndexEntry{
Name: string(bytes.TrimRight(key.ResRef[:], "\x00")),
Type: key.ResourceType,
Offset: int64(entry.Offset),
Size: int64(entry.Size),
})
}
return index, nil
}
// ReadPayload returns one resource's bytes.
func ReadPayload(r io.ReaderAt, entry IndexEntry) ([]byte, error) {
payload := make([]byte, entry.Size)
if _, err := r.ReadAt(payload, entry.Offset); err != nil {
return nil, fmt.Errorf("read resource %q: %w", entry.Name, err)
}
return payload, nil
}
// Read materialises a whole archive. Payloads are subslices of the buffer the
// archive was read into, so nothing is copied twice: a caller must not mutate
// Data. Callers that only need one resource at a time should use ReadIndex
// instead, which never holds the archive at all.
func Read(r io.Reader) (Archive, error) { func Read(r io.Reader) (Archive, error) {
data, err := io.ReadAll(r) data, err := io.ReadAll(r)
if err != nil { if err != nil {
return Archive{}, fmt.Errorf("read erf: %w", err) return Archive{}, fmt.Errorf("read erf: %w", err)
} }
if len(data) < headerSize { index, err := ReadIndex(bytes.NewReader(data), int64(len(data)))
return Archive{}, fmt.Errorf("erf file too small: %d bytes", len(data)) if err != nil {
return Archive{}, err
} }
var hdr header resources := make([]Resource, 0, len(index.Entries))
if err := binary.Read(bytes.NewReader(data[:headerSize]), binary.LittleEndian, &hdr); err != nil { for _, entry := range index.Entries {
return Archive{}, fmt.Errorf("decode erf header: %w", err)
}
keyStart := int(hdr.KeyListOffset)
keyEnd := keyStart + int(hdr.EntryCount)*24
if keyEnd > len(data) {
return Archive{}, fmt.Errorf("erf key list exceeds file bounds")
}
keys := make([]keyEntry, hdr.EntryCount)
if err := binary.Read(bytes.NewReader(data[keyStart:keyEnd]), binary.LittleEndian, &keys); err != nil {
return Archive{}, fmt.Errorf("decode key list: %w", err)
}
resourceStart := int(hdr.ResourceListOffset)
resourceEnd := resourceStart + int(hdr.EntryCount)*8
if resourceEnd > len(data) {
return Archive{}, fmt.Errorf("erf resource list exceeds file bounds")
}
entries := make([]resourceEntry, hdr.EntryCount)
if err := binary.Read(bytes.NewReader(data[resourceStart:resourceEnd]), binary.LittleEndian, &entries); err != nil {
return Archive{}, fmt.Errorf("decode resource list: %w", err)
}
resources := make([]Resource, 0, hdr.EntryCount)
for index, key := range keys {
entry := entries[index]
start := int(entry.Offset)
end := start + int(entry.Size)
if end > len(data) {
return Archive{}, fmt.Errorf("resource %d exceeds file bounds", index)
}
resref := string(bytes.TrimRight(key.ResRef[:], "\x00"))
payload := make([]byte, entry.Size)
copy(payload, data[start:end])
resources = append(resources, Resource{ resources = append(resources, Resource{
Name: resref, Name: entry.Name,
Type: key.ResourceType, Type: entry.Type,
Data: payload, Data: data[entry.Offset : entry.Offset+entry.Size],
Size: int64(entry.Size), Size: entry.Size,
}) })
} }
return Archive{ return Archive{
FileType: string(hdr.FileType[:]), FileType: index.FileType,
Version: string(hdr.Version[:]), Version: index.Version,
Resources: resources, Resources: resources,
}, nil }, nil
} }
+128
View File
@@ -0,0 +1,128 @@
package nwsync
import (
"crypto/sha1"
"encoding/json"
"fmt"
"path"
)
// AssembleOptions describes one merged manifest.
type AssembleOptions struct {
// ArtifactKeys are the depot keys of the artifacts to merge, in
// Mod_HakList order — highest priority first. Each one's index is read
// from the key beside it.
ArtifactKeys []string
TLKKey string // the TLK's key, if the manifest carries one
OutDir string // write locally instead of uploading — the conformance path
GroupID int // 1 = current, 2 = testing; 0 means absent
ModuleName string
Description string
Sink sink // test seam; nil means OutDir or the zone
}
// AssembleResult reports what one assemble run produced.
type AssembleResult struct {
SHA1 string
ManifestPath string
Entries int
}
// Assemble merges the per-artifact NSYM manifests named by Order into one
// manifest. It reads no bulk data at all — only the small index files.
//
// Merge rule is resref shadowing, not concatenation: a resref present in more
// than one artifact resolves to the earliest artifact in Order, which is how
// the game resolves it (upstream's resman adds haks in reverse and lets the
// last one win). Get this backwards and the wrong texture ships silently.
func Assemble(options AssembleOptions) (AssembleResult, error) {
if len(options.ArtifactKeys) == 0 {
return AssembleResult{}, fmt.Errorf("assemble: no artifact keys given")
}
target, err := openSink(options.OutDir, options.Sink)
if err != nil {
return AssembleResult{}, err
}
// The TLK carries no precedence — it is not a hak and shadows nothing —
// so it merges last, after every hak has had its say.
keys := append([]string{}, options.ArtifactKeys...)
if options.TLKKey != "" {
keys = append(keys, options.TLKKey)
}
merged := make([]Entry, 0, 1024)
winner := make(map[Identity]bool, 1024)
var onDiskBytes int64
for _, artifactKey := range keys {
key, err := resolveIndexKey(artifactKey, options.OutDir)
if err != nil {
return AssembleResult{}, err
}
data, sidecarBody, err := target.getIndex(key)
if err != nil {
// An artifact with no index is an artifact whose emit never
// finished. Publishing a manifest without it would ship a release
// missing a hak, so this fails closed.
return AssembleResult{}, fmt.Errorf("assemble: no index for %s: %w", artifactKey, err)
}
entries, err := readManifest(data)
if err != nil {
return AssembleResult{}, fmt.Errorf("%s: %w", target.describe(key), err)
}
sidecar, err := parseSidecar(target.describe(key), sidecarBody)
if err != nil {
return AssembleResult{}, err
}
// Two producers of blobs means a skewed emitter can write blobs the
// merged manifest quietly disagrees with. Refuse to merge across
// mismatched emitter versions.
if sidecar.EmitterVersion != emitterVersion {
return AssembleResult{}, fmt.Errorf(
"assemble: emitter version mismatch: %s was emitted by emitter %q, this is emitter %q",
artifactKey, sidecar.EmitterVersion, emitterVersion)
}
// on_disk_bytes overcounts by the handful of cross-artifact
// duplicates. It is a display statistic; no dedupe pass for it.
onDiskBytes += sidecar.OnDiskBytes
for _, entry := range entries {
identity := entry.identity()
if winner[identity] {
continue
}
winner[identity] = true
merged = append(merged, entry)
}
}
if len(merged) == 0 {
return AssembleResult{}, fmt.Errorf("assemble: merged manifest is empty")
}
data, err := writeManifest(merged)
if err != nil {
return AssembleResult{}, err
}
sha1Hex := fmt.Sprintf("%x", sha1.Sum(data))
manifestKey := path.Join("manifests", sha1Hex)
sidecar := Sidecar{
ModuleName: options.ModuleName,
Description: options.Description,
GroupID: options.GroupID,
}
if err := putManifestPair(target, manifestKey, data, merged, onDiskBytes, sidecar); err != nil {
return AssembleResult{}, err
}
return AssembleResult{SHA1: sha1Hex, ManifestPath: target.describe(manifestKey), Entries: len(merged)}, nil
}
func parseSidecar(where string, data []byte) (Sidecar, error) {
var sidecar Sidecar
if err := json.Unmarshal(data, &sidecar); err != nil {
return Sidecar{}, fmt.Errorf("%s: %w", where, err)
}
return sidecar, nil
}
+80
View File
@@ -0,0 +1,80 @@
package nwsync
import (
"bytes"
"encoding/binary"
"fmt"
"github.com/klauspost/compress/zstd"
)
// NWCompressedBuffer framing, as upstream's neverwinter/compressedbuf.nim
// writes it for NWSync blobs. All fields are little-endian uint32:
//
// magic "NSYC", version 3, algorithm 2 (zstd), uncompressed size,
// zstd header version 1, dictionary 0, then the raw zstd frame.
const (
blobMagic = 0x4359534E // "NSYC" little-endian
blobVersion = 3
algorithmZstd = 2
zstdHeaderVer = 1
zstdDictionary = 0
blobHeaderBytes = 24
)
// EncodeAll/DecodeAll are single-threaded per call, so the default pool of one
// encoder per CPU only buys idle memory: each holds a window-sized history, so
// on a 24-core runner that is ~200 MB of live heap doing nothing. Concurrency 1
// produces byte-identical output.
var (
blobEncoder, _ = zstd.NewWriter(nil, zstd.WithEncoderConcurrency(1))
blobDecoder, _ = zstd.NewReader(nil, zstd.WithDecoderConcurrency(1))
)
// compressBlob wraps data in NWCompressedBuffer framing.
func compressBlob(data []byte) []byte {
var out bytes.Buffer
header := []uint32{blobMagic, blobVersion, algorithmZstd, uint32(len(data)), zstdHeaderVer, zstdDictionary}
for _, field := range header {
_ = binary.Write(&out, binary.LittleEndian, field)
}
out.Write(blobEncoder.EncodeAll(data, nil))
return out.Bytes()
}
// decompressBlob unwraps NWCompressedBuffer framing. It exists so a blob this
// package wrote — or one upstream wrote — can be compared by its uncompressed
// bytes, which is the only comparison that is meaningful across zstd
// implementations.
func decompressBlob(blob []byte) ([]byte, error) {
if len(blob) < blobHeaderBytes {
return nil, fmt.Errorf("blob too small: %d bytes", len(blob))
}
header := make([]uint32, 6)
if err := binary.Read(bytes.NewReader(blob[:blobHeaderBytes]), binary.LittleEndian, header); err != nil {
return nil, fmt.Errorf("decode blob header: %w", err)
}
switch {
case header[0] != blobMagic:
return nil, fmt.Errorf("invalid blob magic: %#x", header[0])
case header[1] != blobVersion:
return nil, fmt.Errorf("unsupported blob version: %d", header[1])
case header[2] != algorithmZstd:
return nil, fmt.Errorf("unsupported compression algorithm: %d", header[2])
case header[4] != zstdHeaderVer:
return nil, fmt.Errorf("unsupported zstd header version: %d", header[4])
case header[5] != zstdDictionary:
return nil, fmt.Errorf("zstd dictionaries are not supported")
}
if header[3] == 0 {
return nil, nil
}
data, err := blobDecoder.DecodeAll(blob[blobHeaderBytes:], nil)
if err != nil {
return nil, fmt.Errorf("decompress blob: %w", err)
}
if uint32(len(data)) != header[3] {
return nil, fmt.Errorf("blob size mismatch: header says %d, got %d", header[3], len(data))
}
return data, nil
}
+278
View File
@@ -0,0 +1,278 @@
package nwsync
import (
"context"
"crypto/sha1"
"fmt"
"io"
"os"
"path"
"path/filepath"
"slices"
"sort"
"strconv"
"strings"
"time"
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/buildinfo"
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/erf"
)
// fileSizeLimit matches upstream's --limit-file-size default of 15 MB. A
// resource over it is a hard failure, not a skip: upstream quit(1)s and so do
// we. Our largest resource today is 13.66 MiB, so the headroom is thin.
const fileSizeLimit = 15 * 1024 * 1024
// skippedTypes are never published, matching upstream's GobalResTypeSkipList.
var skippedTypes = resTypes("nss", "ndb", "gic")
// emitterVersion identifies the blob/manifest byte format this package
// produces. assemble refuses to merge indexes that disagree on it, because two
// producers of blobs mean a skewed emitter can otherwise write blobs the merged
// manifest quietly disagrees with. Bump it only when emitted bytes change — it
// is deliberately not the build revision, which would invalidate every
// published index on every unrelated commit.
const emitterVersion = "1"
// serverTypes are loaded only server-side; a manifest holding nothing else
// has no client contents. Mirrors upstream's GlobalResTypeServerList, whose
// trailing 0 is RESTYPE_INVALID.
var serverTypes = append(resTypes(
"are", "dlg", "fac", "gic", "git", "ifo", "itp", "jrl", "ncs", "ndb",
"nss", "ptm", "utc", "utd", "ute", "uti", "utm", "utp", "uts", "utt", "utw",
), 0)
func resTypes(extensions ...string) []uint16 {
types := make([]uint16, 0, len(extensions))
for _, extension := range extensions {
restype, ok := erf.ResourceTypeForExtension(extension)
if !ok {
panic("nwsync: unknown restype " + extension)
}
types = append(types, restype)
}
return types
}
// EmitResult reports what one emit run produced.
type EmitResult struct {
Name string // artifact name, without extension
ManifestPath string
Entries int
BlobsWritten int
}
// EmitOptions describes one emit run.
type EmitOptions struct {
ArtifactKey string // depot key of the artifact; the NSYM key is derived from it
ArtifactPath string // the file on disk
As string // name override, for a TLK whose filename is not its published name
OutDir string // write locally instead of uploading — the conformance path
Sink sink // test seam; nil means OutDir or the zone
}
// Emit explodes one artifact — a .hak/.erf or a loose file such as the TLK —
// into NWSync blobs plus a NSYM manifest describing only that artifact.
//
// Blobs go up as they are produced and the index lands last, so the presence of
// an index is the publication marker: an artifact whose emit died halfway has
// real blobs in the zone and no index, which is unambiguous. Blob names are
// content hashes, so re-running skips whatever already landed.
func Emit(options EmitOptions) (EmitResult, error) {
artifact, err := os.Open(options.ArtifactPath)
if err != nil {
return EmitResult{}, fmt.Errorf("read artifact: %w", err)
}
defer artifact.Close()
info, err := artifact.Stat()
if err != nil {
return EmitResult{}, fmt.Errorf("read artifact: %w", err)
}
// A section reader, not the file itself: hashing must not move the file
// offset out from under everything that reads the artifact afterwards.
if err := checkArtifactKey(options.ArtifactKey, io.NewSectionReader(artifact, 0, info.Size())); err != nil {
return EmitResult{}, err
}
name := options.As
if name == "" {
name = path.Base(options.ArtifactKey)
}
extension := path.Ext(name)
name = strings.TrimSuffix(name, extension)
key, err := resolveIndexKey(options.ArtifactKey, options.OutDir)
if err != nil {
return EmitResult{}, err
}
index, err := readArtifactIndex(options.ArtifactPath, artifact, info.Size(), name)
if err != nil {
return EmitResult{}, err
}
target, err := openSink(options.OutDir, options.Sink)
if err != nil {
return EmitResult{}, err
}
entries, blobs, onDiskBytes, err := emitResources(artifact, index, target)
if err != nil {
return EmitResult{}, err
}
if len(entries) == 0 {
return EmitResult{}, fmt.Errorf("%s: nothing to index (no publishable resources)", options.ArtifactPath)
}
data, err := writeManifest(entries)
if err != nil {
return EmitResult{}, err
}
if err := putManifestPair(target, key, data, entries, onDiskBytes, Sidecar{ModuleName: name}); err != nil {
return EmitResult{}, err
}
return EmitResult{Name: name, ManifestPath: target.describe(key), Entries: len(entries), BlobsWritten: blobs}, nil
}
// openSink returns the zone sink, or a local directory when outDir is set.
func openSink(outDir string, injected sink) (sink, error) {
if injected != nil {
return injected, nil
}
if outDir != "" {
return dirSink{root: outDir}, nil
}
return newZoneSink(context.Background(), os.Getenv)
}
// readArtifactIndex locates the resources of an ERF/HAK/MOD, or the single
// resource a loose file represents, without reading any payload. Upstream's
// resman does the same dispatch on the file's first three bytes. name is the
// artifact's published name, which for a loose file is also its resref.
func readArtifactIndex(path string, artifact io.ReaderAt, size int64, name string) ([]erf.IndexEntry, error) {
magic := make([]byte, 3)
if size >= 3 {
if _, err := artifact.ReadAt(magic, 0); err != nil {
return nil, fmt.Errorf("%s: %w", path, err)
}
}
switch string(magic) {
case "ERF", "HAK":
index, err := erf.ReadIndex(artifact, size)
if err != nil {
return nil, fmt.Errorf("%s: %w", path, err)
}
return index.Entries, nil
case "MOD":
// A persistent world never publishes module contents, so the .mod
// contributes no bytes to a manifest — it only says which haks and
// which TLK the manifest covers.
return nil, fmt.Errorf("%s: a module is never emitted; a manifest is haks plus the TLK", path)
}
extension := filepath.Ext(filepath.Base(path))
restype, ok := erf.ResourceTypeForExtension(extension)
if !ok {
return nil, fmt.Errorf("%s: unknown resource type %q", path, extension)
}
return []erf.IndexEntry{{Name: name, Type: restype, Offset: 0, Size: size}}, nil
}
// emitResources hashes, compresses and stores one resource at a time, reading
// each payload from the artifact only when its turn comes. Peak memory
// therefore tracks the largest single resource, not the archive: a 2 GB hak
// must emit inside a runner's few spare GB.
func emitResources(artifact io.ReaderAt, index []erf.IndexEntry, target sink) ([]Entry, int, int64, error) {
// A resref appearing twice inside one artifact resolves to the last one,
// the way resman lets the last container added win.
order := make([]Identity, 0, len(index))
latest := make(map[Identity]erf.IndexEntry, len(index))
var tooBig []string
for _, entry := range index {
if _, ok := erf.ExtensionForResourceType(entry.Type); !ok {
return nil, 0, 0, fmt.Errorf("resref %s is not resolvable (unknown restype %d)", entry.Name, entry.Type)
}
if slices.Contains(skippedTypes, entry.Type) {
continue
}
if entry.Size > fileSizeLimit {
tooBig = append(tooBig, fmt.Sprintf("%s: %d bytes > %d", entry.Name, entry.Size, fileSizeLimit))
continue
}
identity := Identity{ResRef: strings.ToLower(entry.Name), ResType: entry.Type}
if _, seen := latest[identity]; !seen {
order = append(order, identity)
}
latest[identity] = entry
}
if len(tooBig) > 0 {
sort.Strings(tooBig)
return nil, 0, 0, fmt.Errorf("resources exceed the file size limit:\n %s", strings.Join(tooBig, "\n "))
}
entries := make([]Entry, 0, len(order))
var blobs int
var onDiskBytes int64
for _, identity := range order {
payload, err := erf.ReadPayload(artifact, latest[identity])
if err != nil {
return nil, 0, 0, err
}
sum := sha1.Sum(payload)
entries = append(entries, Entry{
SHA1: sum,
Size: uint32(len(payload)),
ResRef: identity.ResRef,
ResType: identity.ResType,
})
written, err := target.putBlob(fmt.Sprintf("%x", sum), func() []byte { return compressBlob(payload) })
if err != nil {
return nil, 0, 0, err
}
if written > 0 {
blobs++
onDiskBytes += written
}
}
return entries, blobs, onDiskBytes, nil
}
// created is the sidecar timestamp. SOURCE_DATE_EPOCH pins it so a build can
// be reproduced byte for byte; the manifest itself is deterministic already.
func created() int64 {
if raw := os.Getenv("SOURCE_DATE_EPOCH"); raw != "" {
if seconds, err := strconv.ParseInt(raw, 10, 64); err == nil {
return seconds
}
}
return time.Now().Unix()
}
// putManifestPair stores a NSYM manifest and its .json sidecar at key. The
// caller supplies the sidecar fields it knows; the rest are derived from the
// entries. data must be the serialised form of entries.
func putManifestPair(target sink, key string, data []byte, entries []Entry, onDiskBytes int64, sidecar Sidecar) error {
var totalBytes int64
clientContents := false
for _, entry := range entries {
totalBytes += int64(entry.Size)
if !slices.Contains(serverTypes, entry.ResType) {
clientContents = true
}
}
sidecar.Version = manifestVersion
sidecar.SHA1 = fmt.Sprintf("%x", sha1.Sum(data))
sidecar.HashTreeDepth = hashTreeDepth
sidecar.IncludesModuleContents = false
sidecar.IncludesClientContents = clientContents
sidecar.TotalFiles = len(entries)
sidecar.TotalBytes = totalBytes
sidecar.OnDiskBytes = onDiskBytes
sidecar.Created = created()
sidecar.CreatedWith = buildinfo.String()
sidecar.EmitterVersion = emitterVersion
body, err := marshalSidecar(sidecar)
if err != nil {
return err
}
return target.putIndex(key, data, body)
}
+200
View File
@@ -0,0 +1,200 @@
package nwsync
import (
"bytes"
"encoding/binary"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"path/filepath"
"sort"
"strings"
)
// NSYM manifest, version 3, exactly as upstream neverwinter/nwsync.nim writes
// it. Everything is little-endian:
//
// "NSYM", uint32 version, uint32 entry count, uint32 mapping count,
// entries: byte[20] raw sha1, uint32 size, char[16] resref, uint16 restype
// mappings: uint32 entry index, char[16] resref, uint16 restype
//
// Entries are sorted by lowercase sha1 hex then resref, and a resource whose
// sha1 was already written becomes a mapping instead of a second entry.
const (
manifestVersion = 3
hashTreeDepth = 2
resRefBytes = 16
)
// Entry is one resource in a manifest.
type Entry struct {
SHA1 [20]byte
Size uint32
ResRef string // lowercase, no extension, at most 16 characters
ResType uint16
}
func (e Entry) sha1Hex() string { return hex.EncodeToString(e.SHA1[:]) }
// Identity is what a resref resolves by: name plus type. It is the merge key
// inside one artifact and across artifacts alike.
type Identity struct {
ResRef string
ResType uint16
}
func (e Entry) identity() Identity { return Identity{ResRef: e.ResRef, ResType: e.ResType} }
// writeManifest serialises entries into NSYM v3 bytes.
func writeManifest(entries []Entry) ([]byte, error) {
sorted := make([]Entry, len(entries))
copy(sorted, entries)
sort.SliceStable(sorted, func(i, j int) bool {
left, right := sorted[i].sha1Hex(), sorted[j].sha1Hex()
if left != right {
return left < right
}
return sorted[i].ResRef < sorted[j].ResRef
})
var body, mappings bytes.Buffer
seen := make(map[string]uint32, len(sorted))
var entryCount, mappingCount uint32
for _, entry := range sorted {
padded, err := padResRef(entry.ResRef)
if err != nil {
return nil, err
}
if index, ok := seen[entry.sha1Hex()]; ok {
_ = binary.Write(&mappings, binary.LittleEndian, index)
mappings.Write(padded)
_ = binary.Write(&mappings, binary.LittleEndian, entry.ResType)
mappingCount++
continue
}
seen[entry.sha1Hex()] = entryCount
entryCount++
body.Write(entry.SHA1[:])
_ = binary.Write(&body, binary.LittleEndian, entry.Size)
body.Write(padded)
_ = binary.Write(&body, binary.LittleEndian, entry.ResType)
}
var out bytes.Buffer
out.WriteString("NSYM")
for _, field := range []uint32{manifestVersion, entryCount, mappingCount} {
_ = binary.Write(&out, binary.LittleEndian, field)
}
out.Write(body.Bytes())
out.Write(mappings.Bytes())
return out.Bytes(), nil
}
// readManifest parses NSYM v3 bytes. Mappings are expanded back into entries,
// the way upstream's reader does, so a caller sees one entry per resref.
func readManifest(data []byte) ([]Entry, error) {
reader := bytes.NewReader(data)
magic := make([]byte, 4)
if _, err := io.ReadFull(reader, magic); err != nil || string(magic) != "NSYM" {
return nil, fmt.Errorf("not a manifest (invalid magic bytes)")
}
var version, entryCount, mappingCount uint32
for _, field := range []*uint32{&version, &entryCount, &mappingCount} {
if err := binary.Read(reader, binary.LittleEndian, field); err != nil {
return nil, fmt.Errorf("truncated manifest header: %w", err)
}
}
if version != manifestVersion {
return nil, fmt.Errorf("unsupported manifest version %d", version)
}
entries := make([]Entry, 0, entryCount+mappingCount)
for i := uint32(0); i < entryCount; i++ {
var entry Entry
if _, err := io.ReadFull(reader, entry.SHA1[:]); err != nil {
return nil, fmt.Errorf("truncated entry %d: %w", i, err)
}
if err := binary.Read(reader, binary.LittleEndian, &entry.Size); err != nil {
return nil, fmt.Errorf("truncated entry %d: %w", i, err)
}
resref, restype, err := readResRef(reader)
if err != nil {
return nil, fmt.Errorf("truncated entry %d: %w", i, err)
}
entry.ResRef, entry.ResType = resref, restype
entries = append(entries, entry)
}
for i := uint32(0); i < mappingCount; i++ {
var index uint32
if err := binary.Read(reader, binary.LittleEndian, &index); err != nil {
return nil, fmt.Errorf("truncated mapping %d: %w", i, err)
}
if index >= entryCount {
return nil, fmt.Errorf("mapping %d references non-existent entry %d", i, index)
}
resref, restype, err := readResRef(reader)
if err != nil {
return nil, fmt.Errorf("truncated mapping %d: %w", i, err)
}
target := entries[index]
entries = append(entries, Entry{SHA1: target.SHA1, Size: target.Size, ResRef: resref, ResType: restype})
}
return entries, nil
}
func readResRef(reader *bytes.Reader) (string, uint16, error) {
raw := make([]byte, resRefBytes)
if _, err := io.ReadFull(reader, raw); err != nil {
return "", 0, err
}
var restype uint16
if err := binary.Read(reader, binary.LittleEndian, &restype); err != nil {
return "", 0, err
}
return strings.ToLower(string(bytes.TrimRight(raw, "\x00"))), restype, nil
}
func padResRef(resref string) ([]byte, error) {
if len(resref) > resRefBytes {
return nil, fmt.Errorf("resref %q exceeds %d characters", resref, resRefBytes)
}
padded := make([]byte, resRefBytes)
copy(padded, strings.ToLower(resref))
return padded, nil
}
// Sidecar is the .json file written next to every manifest. Clients fetch it
// (nwn_nwsync_fetch.nim), so the field names and order match upstream.
// Upstream omits an integer meta field whose value is 0, hence group_id's
// omitempty: 0 means absent, not "group zero".
type Sidecar struct {
Version int `json:"version"`
SHA1 string `json:"sha1"`
HashTreeDepth int `json:"hash_tree_depth"`
ModuleName string `json:"module_name"`
Description string `json:"description"`
IncludesModuleContents bool `json:"includes_module_contents"`
IncludesClientContents bool `json:"includes_client_contents"`
TotalFiles int `json:"total_files"`
TotalBytes int64 `json:"total_bytes"`
OnDiskBytes int64 `json:"on_disk_bytes"`
Created int64 `json:"created"`
CreatedWith string `json:"created_with"`
EmitterVersion string `json:"emitter_version,omitempty"`
GroupID int `json:"group_id,omitempty"`
}
func marshalSidecar(sidecar Sidecar) ([]byte, error) {
body, err := json.MarshalIndent(sidecar, "", " ")
if err != nil {
return nil, err
}
// Upstream terminates the file with CRLF; match it.
return append(body, '\r', '\n'), nil
}
// blobPath is the data store path for a blob, hash tree depth 2.
func blobPath(root, sha1Hex string) string {
return filepath.Join(root, "data", "sha1", sha1Hex[0:2], sha1Hex[2:4], sha1Hex)
}
+150
View File
@@ -0,0 +1,150 @@
package nwsync
import (
"bytes"
"fmt"
"math/rand"
"os"
"path/filepath"
"runtime"
"runtime/debug"
"testing"
"time"
"github.com/klauspost/compress/zstd"
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/erf"
)
// TestSingleThreadedEncoderMatchesDefault pins the claim the blob encoder's
// concurrency setting rests on: it saves memory only, and a published blob is
// the same bytes either way.
func TestSingleThreadedEncoderMatchesDefault(t *testing.T) {
standard, err := zstd.NewWriter(nil)
if err != nil {
t.Fatal(err)
}
defer standard.Close()
body := make([]byte, 4<<20)
random := rand.New(rand.NewSource(1))
random.Read(body[:len(body)/2])
for _, size := range []int{0, 1, 4 << 10, len(body)} {
if !bytes.Equal(blobEncoder.EncodeAll(body[:size], nil), standard.EncodeAll(body[:size], nil)) {
t.Fatalf("%d bytes compress differently at concurrency 1", size)
}
}
}
// resourceSize is one payload in the memory fixtures. Real haks hold a few MB
// per resource, and peak memory is meant to track that, not the archive.
const resourceSize = 1 << 20
// writeStreamedHak builds a hak of count resources without ever holding the
// archive in memory, so the fixture itself does not decide the measurement.
// Payloads are distinct, so no blob is deduplicated away.
func writeStreamedHak(t *testing.T, path string, count int) {
t.Helper()
payload := filepath.Join(t.TempDir(), "payload.bin")
body := make([]byte, resourceSize)
for index := range body {
body[index] = byte(index)
}
resources := make([]erf.Resource, 0, count)
for index := range count {
// A distinct first byte per resource is enough to give every payload
// its own sha1 while still streaming from one file per resource.
unique := filepath.Join(filepath.Dir(payload), fmt.Sprintf("p%d.bin", index))
body[0] = byte(index)
body[1] = byte(index >> 8)
if err := os.WriteFile(unique, body, 0o644); err != nil {
t.Fatal(err)
}
resources = append(resources, erf.Resource{
Name: fmt.Sprintf("res%05d", index),
Type: restype(t, "tga"),
SourcePath: unique,
Size: resourceSize,
})
}
file, err := os.Create(path)
if err != nil {
t.Fatal(err)
}
defer file.Close()
if err := erf.Write(file, erf.New("HAK", resources)); err != nil {
t.Fatalf("write hak: %v", err)
}
}
// peakHeapDuring runs work while sampling the heap, and returns the largest
// live heap it saw.
func peakHeapDuring(work func()) uint64 {
runtime.GC()
done := make(chan struct{})
peak := make(chan uint64, 1)
go func() {
var highest uint64
var stats runtime.MemStats
for {
select {
case <-done:
peak <- highest
return
default:
}
runtime.ReadMemStats(&stats)
if stats.HeapAlloc > highest {
highest = stats.HeapAlloc
}
time.Sleep(time.Millisecond)
}
}()
work()
close(done)
return <-peak
}
// TestEmitPeakMemoryDoesNotScaleWithArtifactSize is the regression check for
// the OOM kills on large haks: emit used to hold the whole archive (twice), so
// a 2 GB hak needed about 10 GB. Emitting an archive 8× bigger must not cost
// meaningfully more memory.
func TestEmitPeakMemoryDoesNotScaleWithArtifactSize(t *testing.T) {
if testing.Short() {
t.Skip("writes a 64 MB fixture")
}
// A lazy GC lets garbage pile up in proportion to the live heap, which
// hides the thing under test. Collecting eagerly makes the sampled heap
// track what emit actually holds.
defer debug.SetGCPercent(debug.SetGCPercent(10))
measure := func(count int) uint64 {
dir := t.TempDir()
hak := filepath.Join(dir, "big.hak")
writeStreamedHak(t, hak, count)
// The key is computed outside the measurement: the test helper reads
// the whole file to hash it, which emit itself no longer does.
options := EmitOptions{
ArtifactKey: artifactKey(t, hak),
ArtifactPath: hak,
As: filepath.Base(hak),
OutDir: filepath.Join(dir, "out"),
}
return peakHeapDuring(func() {
if _, err := Emit(options); err != nil {
t.Fatalf("emit %d resources: %v", count, err)
}
})
}
small := measure(8) // 8 MB
large := measure(64) // 64 MB
const slack = 24 << 20
t.Logf("peak heap: 8 MB hak %d bytes, 64 MB hak %d bytes", small, large)
if large > small+slack {
t.Fatalf("peak heap scaled with artifact size: 8 MB hak peaked at %d bytes, 64 MB hak at %d", small, large)
}
}
+476
View File
@@ -0,0 +1,476 @@
package nwsync
import (
"bytes"
"crypto/sha1"
"crypto/sha256"
"encoding/binary"
"encoding/hex"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/erf"
)
func restype(t *testing.T, extension string) uint16 {
t.Helper()
value, ok := erf.ResourceTypeForExtension(extension)
if !ok {
t.Fatalf("unknown restype %q", extension)
}
return value
}
// writeHak builds a HAK fixture from resref.ext => body pairs.
func writeHak(t *testing.T, path string, contents map[string][]byte) {
t.Helper()
resources := make([]erf.Resource, 0, len(contents))
for name, body := range contents {
stem, extension, _ := strings.Cut(name, ".")
resources = append(resources, erf.Resource{
Name: stem,
Type: restype(t, extension),
Data: body,
Size: int64(len(body)),
})
}
var out bytes.Buffer
if err := erf.Write(&out, erf.New("HAK", resources)); err != nil {
t.Fatalf("write hak: %v", err)
}
if err := os.WriteFile(path, out.Bytes(), 0o644); err != nil {
t.Fatalf("write hak file: %v", err)
}
}
// artifactKey is the depot key a file would be published under: the sha256 of
// its bytes, hash-tree depth 2, keeping the extension.
func artifactKey(t *testing.T, path string) string {
t.Helper()
body, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
sum := sha256.Sum256(body)
digest := hex.EncodeToString(sum[:])
return "artifacts/haks/sha256/" + digest[0:2] + "/" + digest[2:4] + "/" + digest + filepath.Ext(path)
}
// emitLocal emits one artifact into a local tree, the conformance path.
func emitLocal(t *testing.T, path, out string) (EmitResult, error) {
t.Helper()
return Emit(EmitOptions{
ArtifactKey: artifactKey(t, path),
ArtifactPath: path,
As: filepath.Base(path),
OutDir: out,
})
}
func TestBlobFramingRoundTrips(t *testing.T) {
data := []byte("the quick brown fox jumps over the lazy dog, repeatedly and at length")
blob := compressBlob(data)
header := make([]uint32, 6)
if err := binary.Read(bytes.NewReader(blob[:blobHeaderBytes]), binary.LittleEndian, header); err != nil {
t.Fatalf("read header: %v", err)
}
want := []uint32{blobMagic, 3, 2, uint32(len(data)), 1, 0}
for i := range want {
if header[i] != want[i] {
t.Errorf("header field %d = %d, want %d", i, header[i], want[i])
}
}
if string(blob[:4]) != "NSYC" {
t.Errorf("magic bytes = %q, want NSYC", blob[:4])
}
got, err := decompressBlob(blob)
if err != nil {
t.Fatalf("decompress: %v", err)
}
if !bytes.Equal(got, data) {
t.Errorf("round trip mismatch: %q", got)
}
}
func TestManifestBytesMatchUpstreamLayout(t *testing.T) {
shared := sha1.Sum([]byte("shared"))
other := sha1.Sum([]byte("other"))
// Deliberately out of order, and with two resrefs sharing one sha1: the
// second one must become a mapping, not a second entry.
entries := []Entry{
{SHA1: other, Size: 5, ResRef: "zzz", ResType: 1},
{SHA1: shared, Size: 6, ResRef: "bbb", ResType: 2},
{SHA1: shared, Size: 6, ResRef: "aaa", ResType: 3},
}
data, err := writeManifest(entries)
if err != nil {
t.Fatalf("write manifest: %v", err)
}
if string(data[:4]) != "NSYM" {
t.Fatalf("magic = %q", data[:4])
}
var version, entryCount, mappingCount uint32
reader := bytes.NewReader(data[4:16])
for _, field := range []*uint32{&version, &entryCount, &mappingCount} {
_ = binary.Read(reader, binary.LittleEndian, field)
}
if version != 3 || entryCount != 2 || mappingCount != 1 {
t.Fatalf("header = version %d, %d entries, %d mappings; want 3/2/1", version, entryCount, mappingCount)
}
wantSize := 16 + int(entryCount)*(20+4+16+2) + int(mappingCount)*(4+16+2)
if len(data) != wantSize {
t.Fatalf("manifest is %d bytes, want %d", len(data), wantSize)
}
// Sorted by sha1 hex then resref, so the shared hash's "aaa" is the entry
// and "bbb" is demoted to a mapping.
round, err := readManifest(data)
if err != nil {
t.Fatalf("read manifest: %v", err)
}
if len(round) != 3 {
t.Fatalf("round trip returned %d entries, want 3", len(round))
}
byResRef := map[string]Entry{}
for _, entry := range round {
byResRef[entry.ResRef] = entry
}
for _, entry := range entries {
got, ok := byResRef[entry.ResRef]
if !ok {
t.Fatalf("resref %q lost in round trip", entry.ResRef)
}
if got != entry {
t.Errorf("resref %q = %+v, want %+v", entry.ResRef, got, entry)
}
}
if round[0].ResRef != "aaa" && round[1].ResRef != "aaa" {
t.Errorf("entries are not sorted by sha1 then resref: %+v", round)
}
}
func TestEmitWritesBlobsAndManifest(t *testing.T) {
dir := t.TempDir()
hak := filepath.Join(dir, "sow_test_01.hak")
body := []byte("texture bytes")
writeHak(t, hak, map[string][]byte{
"bloodstain1.tga": body,
"copy1.txi": body, // same content, different resref: one blob
"script1.nss": []byte("void main() {}"),
"debug1.ndb": []byte("debug"),
"comment1.gic": []byte("comment"),
})
out := filepath.Join(dir, "out")
result, err := emitLocal(t, hak, out)
if err != nil {
t.Fatalf("emit: %v", err)
}
if result.Entries != 2 {
t.Errorf("emitted %d entries, want 2 (nss/ndb/gic are always skipped)", result.Entries)
}
if result.BlobsWritten != 1 {
t.Errorf("wrote %d blobs, want 1 (identical content shares a blob)", result.BlobsWritten)
}
// The blob is named by the sha1 of the uncompressed bytes, under a depth-2
// hash tree, and decompresses back to exactly those bytes.
sum := sha1.Sum(body)
name := hex.EncodeToString(sum[:])
path := filepath.Join(out, "data", "sha1", name[0:2], name[2:4], name)
blob, err := os.ReadFile(path)
if err != nil {
t.Fatalf("blob missing at %s: %v", path, err)
}
got, err := decompressBlob(blob)
if err != nil {
t.Fatalf("decompress blob: %v", err)
}
if !bytes.Equal(got, body) {
t.Errorf("blob decompressed to %q, want %q", got, body)
}
entries := readEmitted(t, result.ManifestPath)
for _, entry := range entries {
if entry.ResType == restype(t, "nss") || entry.ResType == restype(t, "ndb") || entry.ResType == restype(t, "gic") {
t.Errorf("skipped restype leaked into the manifest: %+v", entry)
}
}
var sidecar Sidecar
body2, err := os.ReadFile(result.ManifestPath + ".json")
if err != nil {
t.Fatalf("sidecar missing: %v", err)
}
if err := json.Unmarshal(body2, &sidecar); err != nil {
t.Fatalf("sidecar json: %v", err)
}
if sidecar.TotalFiles != 2 || sidecar.HashTreeDepth != 2 || sidecar.Version != 3 {
t.Errorf("sidecar = %+v", sidecar)
}
if sidecar.IncludesModuleContents {
t.Error("sidecar claims module contents; a published manifest never has them")
}
if strings.Contains(string(body2), "group_id") {
t.Error("per-artifact sidecar should omit group_id (0 means absent)")
}
if _, err := os.Stat(filepath.Join(out, "latest")); err == nil {
t.Error("a latest file was written; there must never be one")
}
}
func readEmitted(t *testing.T, indexPath string) []Entry {
t.Helper()
data, err := os.ReadFile(indexPath)
if err != nil {
t.Fatalf("read emitted manifest: %v", err)
}
entries, err := readManifest(data)
if err != nil {
t.Fatalf("parse emitted manifest: %v", err)
}
return entries
}
func TestEmitFailsClosedOnOversizeResource(t *testing.T) {
dir := t.TempDir()
hak := filepath.Join(dir, "big.hak")
writeHak(t, hak, map[string][]byte{"huge1.tga": make([]byte, fileSizeLimit+1)})
if _, err := emitLocal(t, hak, filepath.Join(dir, "out")); err == nil {
t.Fatal("emit accepted a resource over the 15 MB limit")
}
}
func TestEmitLooseFile(t *testing.T) {
dir := t.TempDir()
tlk := filepath.Join(dir, "sow_tlk.tlk")
if err := os.WriteFile(tlk, []byte("TLK V3.0 payload"), 0o644); err != nil {
t.Fatal(err)
}
out := filepath.Join(dir, "out")
result, err := emitLocal(t, tlk, out)
if err != nil {
t.Fatalf("emit tlk: %v", err)
}
entries := readEmitted(t, result.ManifestPath)
if len(entries) != 1 || entries[0].ResRef != "sow_tlk" || entries[0].ResType != restype(t, "tlk") {
t.Fatalf("tlk emitted as %+v", entries)
}
if result.BlobsWritten != 1 {
t.Errorf("wrote %d blobs, want 1", result.BlobsWritten)
}
}
// emitFixture emits two haks that share a resref, so the merge rule is
// observable: "top" holds the winning body, "assets" the shadowed one.
func emitFixture(t *testing.T) (out string, keys map[string]string, topBody, assetBody []byte) {
t.Helper()
dir := t.TempDir()
topBody = []byte("2da from sow_top")
assetBody = []byte("2da from the asset hak")
writeHak(t, filepath.Join(dir, "sow_top.hak"), map[string][]byte{"appearance.2da": topBody})
writeHak(t, filepath.Join(dir, "sow_core_01.hak"), map[string][]byte{
"appearance.2da": assetBody,
"bloodstain1.tga": []byte("blood"),
})
out = filepath.Join(dir, "out")
keys = map[string]string{}
for _, name := range []string{"sow_top", "sow_core_01"} {
path := filepath.Join(dir, name+".hak")
keys[name] = artifactKey(t, path)
if _, err := emitLocal(t, path, out); err != nil {
t.Fatalf("emit %s: %v", name, err)
}
}
return out, keys, topBody, assetBody
}
func TestAssembleShadowsByOrder(t *testing.T) {
entriesDir, keys, topBody, assetBody := emitFixture(t)
result, err := Assemble(AssembleOptions{
ArtifactKeys: []string{keys["sow_top"], keys["sow_core_01"]},
OutDir: entriesDir,
GroupID: 2,
})
if err != nil {
t.Fatalf("assemble: %v", err)
}
if result.Entries != 2 {
t.Fatalf("merged %d entries, want 2 (appearance.2da is shadowed, not duplicated)", result.Entries)
}
data, err := os.ReadFile(result.ManifestPath)
if err != nil {
t.Fatalf("read merged manifest: %v", err)
}
merged := sha1.Sum(data)
if hex.EncodeToString(merged[:]) != result.SHA1 || filepath.Base(result.ManifestPath) != result.SHA1 {
t.Errorf("manifest is not named by its own sha1: %s", result.ManifestPath)
}
entries, err := readManifest(data)
if err != nil {
t.Fatalf("parse merged manifest: %v", err)
}
for _, entry := range entries {
if entry.ResRef != "appearance" {
continue
}
if entry.SHA1 != sha1.Sum(topBody) {
t.Errorf("appearance.2da resolved to the wrong hak; want the earliest in --order")
}
if entry.SHA1 == sha1.Sum(assetBody) {
t.Error("appearance.2da resolved to the shadowed hak")
}
}
sidecar := Sidecar{}
body, err := os.ReadFile(result.ManifestPath + ".json")
if err != nil {
t.Fatalf("merged sidecar missing: %v", err)
}
if err := json.Unmarshal(body, &sidecar); err != nil {
t.Fatalf("merged sidecar json: %v", err)
}
if sidecar.GroupID != 2 {
t.Errorf("group_id = %d, want 2 (testing)", sidecar.GroupID)
}
if sidecar.SHA1 != result.SHA1 {
t.Errorf("sidecar sha1 = %s, want %s", sidecar.SHA1, result.SHA1)
}
if sidecar.TotalFiles != 2 {
t.Errorf("total_files = %d, want 2", sidecar.TotalFiles)
}
}
func TestAssembleReversedOrderPicksTheOtherHak(t *testing.T) {
entriesDir, keys, topBody, assetBody := emitFixture(t)
result, err := Assemble(AssembleOptions{
ArtifactKeys: []string{keys["sow_core_01"], keys["sow_top"]},
OutDir: entriesDir,
})
if err != nil {
t.Fatalf("assemble: %v", err)
}
data, _ := os.ReadFile(result.ManifestPath)
entries, err := readManifest(data)
if err != nil {
t.Fatalf("parse merged manifest: %v", err)
}
for _, entry := range entries {
if entry.ResRef == "appearance" && entry.SHA1 != sha1.Sum(assetBody) {
t.Errorf("appearance.2da did not follow --order; still resolves to %x", entry.SHA1)
}
}
_ = topBody
}
func TestAssembleRefusesMismatchedEmitterVersions(t *testing.T) {
entriesDir, keys, _, _ := emitFixture(t)
index, err := resolveIndexKey(keys["sow_core_01"], entriesDir)
if err != nil {
t.Fatal(err)
}
path := filepath.Join(entriesDir, index+".json")
var sidecar Sidecar
body, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
if err := json.Unmarshal(body, &sidecar); err != nil {
t.Fatal(err)
}
sidecar.EmitterVersion = "0"
patched, err := marshalSidecar(sidecar)
if err != nil {
t.Fatal(err)
}
if err := os.WriteFile(path, patched, 0o644); err != nil {
t.Fatal(err)
}
_, err = Assemble(AssembleOptions{
ArtifactKeys: []string{keys["sow_top"], keys["sow_core_01"]},
OutDir: entriesDir,
})
if err == nil || !strings.Contains(err.Error(), "emitter version mismatch") {
t.Fatalf("assemble merged across emitter versions: %v", err)
}
}
func TestEmitHonoursSourceDateEpoch(t *testing.T) {
t.Setenv("SOURCE_DATE_EPOCH", "1700000000")
dir := t.TempDir()
hak := filepath.Join(dir, "pinned.hak")
writeHak(t, hak, map[string][]byte{"one1.tga": []byte("body")})
result, err := emitLocal(t, hak, filepath.Join(dir, "out"))
if err != nil {
t.Fatalf("emit: %v", err)
}
body, err := os.ReadFile(result.ManifestPath + ".json")
if err != nil {
t.Fatal(err)
}
var sidecar Sidecar
if err := json.Unmarshal(body, &sidecar); err != nil {
t.Fatal(err)
}
if sidecar.Created != 1700000000 {
t.Errorf("created = %d, want the pinned SOURCE_DATE_EPOCH", sidecar.Created)
}
}
func TestEmitRejectsAModule(t *testing.T) {
dir := t.TempDir()
path := filepath.Join(dir, "sow.mod")
var out bytes.Buffer
if err := erf.Write(&out, erf.New("MOD", []erf.Resource{
{Name: "module", Type: restype(t, "ifo"), Data: []byte("ifo"), Size: 3},
})); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(path, out.Bytes(), 0o644); err != nil {
t.Fatal(err)
}
if _, err := emitLocal(t, path, filepath.Join(dir, "out")); err == nil {
t.Fatal("emit accepted a .mod; a manifest never carries module contents")
}
}
func TestAssembleFailsClosedOnMissingIndex(t *testing.T) {
entriesDir, keys, _, _ := emitFixture(t)
missing := "artifacts/haks/sha256/00/11/" + strings.Repeat("0", 64) + ".hak"
_, err := Assemble(AssembleOptions{
ArtifactKeys: []string{keys["sow_top"], missing},
OutDir: entriesDir,
})
if err == nil || !strings.Contains(err.Error(), missing) {
t.Fatalf("assemble did not fail closed and name the missing artifact: %v", err)
}
}
func TestRunUsageErrors(t *testing.T) {
cases := [][]string{
nil,
{"nope"},
{"emit"},
{"emit", "artifact-key.hak"},
{"emit", "a", "b", "c"},
{"assemble", "--out", "y"},
}
for _, args := range cases {
var out, errw bytes.Buffer
if code := Run(args, &out, &errw); code != exitUsage {
t.Errorf("Run(%v) exit=%d, want %d", args, code, exitUsage)
}
}
}
+140
View File
@@ -0,0 +1,140 @@
// Package nwsync publishes NWSync repository data: blobs and a per-artifact
// NSYM manifest at each artifact's birth (emit), and one merged manifest at
// module release (assemble).
//
// The split exists because upstream nwn_nwsync_write wants every hak, the TLK
// and the module present in one run on one disk, which our build hosts cannot
// hold. Upstream stays the conformance oracle: manifests compare byte for
// byte, blobs compare after decompression.
package nwsync
import (
"flag"
"fmt"
"io"
)
const (
exitOK = 0
exitUsage = 64
exitInternal = 70
)
// Run executes an nwsync subcommand. args[0] is the subcommand (emit|assemble);
// returns the process exit code.
func Run(args []string, stdout, stderr io.Writer) int {
if len(args) == 0 {
printRunUsage(stderr)
return exitUsage
}
switch args[0] {
case "emit":
return runEmit(args[1:], stdout, stderr)
case "assemble":
return runAssemble(args[1:], stdout, stderr)
case "-h", "--help", "help":
printRunUsage(stdout)
return exitOK
default:
fmt.Fprintf(stderr, "nwsync: unknown subcommand %q\n\n", args[0])
printRunUsage(stderr)
return exitUsage
}
}
func printRunUsage(w io.Writer) {
fmt.Fprint(w, `usage:
nwsync emit [--as NAME] [--out DIR] <artifact-key> <file>
nwsync assemble --group-id N [--tlk-key KEY] [--out DIR] <artifact-key>...
emit explodes one .hak/.erf or one loose file (the TLK) into NWSync blobs plus
a NSYM index covering only that artifact, and uploads both. assemble merges
those indexes into one manifest, reading no bulk data. Artifact keys are depot
keys; an index lives beside its artifact, with the extension replaced.
--out DIR writes to a local repository tree instead of uploading, which is the
conformance path against upstream nwn_nwsync_write. Without it, the zone comes
from NWSYNC_STORAGE_ZONE, NWSYNC_STORAGE_PASSWORD and BUNNY_STORAGE_HOST.
`)
}
// parseArgs parses flags that may appear before, after or between positionals.
// Go's flag package stops at the first non-flag argument, which turns
// `emit <key> <file> --out DIR` into a confusing arity error.
func parseArgs(fs *flag.FlagSet, args []string) ([]string, error) {
var positional []string
for {
if err := fs.Parse(args); err != nil {
return nil, err
}
rest := fs.Args()
if len(rest) == 0 {
return positional, nil
}
positional = append(positional, rest[0])
args = rest[1:]
}
}
func runEmit(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("emit", flag.ContinueOnError)
fs.SetOutput(stderr)
as := fs.String("as", "", "published name of the artifact, when it differs from the key")
out := fs.String("out", "", "write to a local repository tree instead of uploading")
positional, err := parseArgs(fs, args)
if err != nil {
return exitUsage
}
if len(positional) != 2 {
fmt.Fprintf(stderr, "nwsync emit: <artifact-key> and <file> are both required\n")
return exitUsage
}
result, err := Emit(EmitOptions{
ArtifactKey: positional[0],
ArtifactPath: positional[1],
As: *as,
OutDir: *out,
})
if err != nil {
fmt.Fprintf(stderr, "nwsync emit: %v\n", err)
return exitInternal
}
fmt.Fprintf(stdout, "emitted %s: %d resources, %d new blobs, index %s\n",
result.Name, result.Entries, result.BlobsWritten, result.ManifestPath)
return exitOK
}
func runAssemble(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("assemble", flag.ContinueOnError)
fs.SetOutput(stderr)
tlkKey := fs.String("tlk-key", "", "depot key of the TLK, which shadows nothing and merges last")
out := fs.String("out", "", "write to a local repository tree instead of uploading")
groupID := fs.Int("group-id", 0, "NWSync group id (1 = current, 2 = testing; 0 omits it)")
moduleName := fs.String("module-name", "", "module name recorded in the sidecar")
description := fs.String("description", "", "description recorded in the sidecar")
positional, err := parseArgs(fs, args)
if err != nil {
return exitUsage
}
if len(positional) == 0 {
fmt.Fprintf(stderr, "nwsync assemble: at least one artifact key is required\n")
return exitUsage
}
result, err := Assemble(AssembleOptions{
ArtifactKeys: positional,
TLKKey: *tlkKey,
OutDir: *out,
GroupID: *groupID,
ModuleName: *moduleName,
Description: *description,
})
if err != nil {
fmt.Fprintf(stderr, "nwsync assemble: %v\n", err)
return exitInternal
}
fmt.Fprintf(stdout, "assembled manifest %s: %d resources, %s\n",
result.SHA1, result.Entries, result.ManifestPath)
return exitOK
}
+212
View File
@@ -0,0 +1,212 @@
package nwsync
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
"io"
"os"
"path"
"path/filepath"
"strings"
"git.westgate.pw/ShadowsOverWestgate/sow-tools/internal/depot"
)
// sink is where an emit or assemble run puts what it produces. The zone is the
// production sink; a local directory exists only as the conformance path, so
// upstream's output and ours can be diffed on a developer machine.
type sink interface {
// putBlob stores one NWCompressedBuffer blob under its sha1 name and
// returns the bytes stored, or 0 if the blob was already there. Blob names
// are content hashes, so an existing name is existing content — which is
// why body is a thunk: compression is the expensive part of emit and a
// blob that is already stored must not pay for it.
putBlob(sha1Hex string, body func() []byte) (int64, error)
// putIndex stores a NSYM manifest and its sidecar under key, which is
// either an artifact-derived object key or a local path.
putIndex(key string, manifest, sidecar []byte) error
// getIndex reads back a NSYM manifest and its sidecar.
getIndex(key string) (manifest, sidecar []byte, err error)
// describe names the sink for messages.
describe(key string) string
}
// dirSink writes a local NWSync repository tree.
type dirSink struct{ root string }
func (s dirSink) putBlob(sha1Hex string, body func() []byte) (int64, error) {
blob := blobPath(s.root, sha1Hex)
if _, err := os.Stat(blob); err == nil {
return 0, nil
}
if err := os.MkdirAll(filepath.Dir(blob), 0o755); err != nil {
return 0, fmt.Errorf("create blob directory: %w", err)
}
data := body()
if err := os.WriteFile(blob, data, 0o644); err != nil {
return 0, fmt.Errorf("write blob: %w", err)
}
return int64(len(data)), nil
}
func (s dirSink) putIndex(key string, manifest, sidecar []byte) error {
target := filepath.Join(s.root, filepath.FromSlash(key))
if err := os.MkdirAll(filepath.Dir(target), 0o755); err != nil {
return fmt.Errorf("create manifest directory: %w", err)
}
if err := os.WriteFile(target, manifest, 0o644); err != nil {
return fmt.Errorf("write manifest: %w", err)
}
if err := os.WriteFile(target+".json", sidecar, 0o644); err != nil {
return fmt.Errorf("write sidecar: %w", err)
}
return nil
}
func (s dirSink) getIndex(key string) ([]byte, []byte, error) {
target := filepath.Join(s.root, filepath.FromSlash(key))
manifest, err := os.ReadFile(target)
if err != nil {
return nil, nil, fmt.Errorf("read index: %w", err)
}
sidecar, err := os.ReadFile(target + ".json")
if err != nil {
return nil, nil, fmt.Errorf("read sidecar: %w", err)
}
return manifest, sidecar, nil
}
func (s dirSink) describe(key string) string {
return filepath.Join(s.root, filepath.FromSlash(key))
}
// zoneSink uploads straight to the NWSync storage zone. Nothing bulky is ever
// written to the runner's disk: the working set is one resource at a time.
type zoneSink struct {
store depot.KeyStore
ctx context.Context
zone string
}
func (s zoneSink) putBlob(sha1Hex string, body func() []byte) (int64, error) {
key := path.Join("data", "sha1", sha1Hex[0:2], sha1Hex[2:4], sha1Hex)
// A throttled probe must never be read as "missing, re-upload" or as
// "present, skip", so only a confirmed Present skips the upload.
state, _, err := s.store.ProbeKey(s.ctx, key)
if err != nil {
return 0, fmt.Errorf("probe blob %s: %w", sha1Hex, err)
}
if state == depot.Present {
return 0, nil
}
data := body()
if err := s.put(key, data); err != nil {
return 0, fmt.Errorf("upload blob %s: %w", sha1Hex, err)
}
return int64(len(data)), nil
}
func (s zoneSink) putIndex(key string, manifest, sidecar []byte) error {
// The manifest lands last: its presence is the publication marker, so it
// must never appear before the blobs it names.
if err := s.put(key+".json", sidecar); err != nil {
return fmt.Errorf("upload sidecar %s: %w", key, err)
}
if err := s.put(key, manifest); err != nil {
return fmt.Errorf("upload index %s: %w", key, err)
}
return nil
}
func (s zoneSink) getIndex(key string) ([]byte, []byte, error) {
manifest, err := s.store.GetKey(s.ctx, key)
if err != nil {
return nil, nil, fmt.Errorf("read index: %w", err)
}
sidecar, err := s.store.GetKey(s.ctx, key+".json")
if err != nil {
return nil, nil, fmt.Errorf("read sidecar: %w", err)
}
return manifest, sidecar, nil
}
func (s zoneSink) describe(key string) string { return s.zone + "/" + key }
func (s zoneSink) put(key string, body []byte) error {
sum := sha256.Sum256(body)
return s.store.PutReader(s.ctx, key, bytes.NewReader(body), int64(len(body)), hex.EncodeToString(sum[:]))
}
// newZoneSink builds the upload sink from the environment. NWSync data lives
// in its own zone, separate from the asset depot, so it has its own zone and
// credential; only the host is shared, and Crucible has no default host.
func newZoneSink(ctx context.Context, getenv func(string) string) (sink, error) {
cfg := depot.LoadConfig(getenv)
cfg.StorageZone = getenv("NWSYNC_STORAGE_ZONE")
cfg.WriteKey = getenv("NWSYNC_STORAGE_PASSWORD")
cfg.ReadKey = cfg.WriteKey
if cfg.StorageZone == "" {
return nil, fmt.Errorf("NWSYNC_STORAGE_ZONE is unset (or pass --out DIR to write locally)")
}
if cfg.WriteKey == "" {
return nil, fmt.Errorf("NWSYNC_STORAGE_PASSWORD is unset (or pass --out DIR to write locally)")
}
if cfg.StorageHost == "" {
return nil, fmt.Errorf("BUNNY_STORAGE_HOST is unset")
}
store, err := depot.NewKeyStore(cfg)
if err != nil {
return nil, err
}
return zoneSink{store: store, ctx: ctx, zone: cfg.StorageZone}, nil
}
// indexKey is where an artifact's NSYM lives: beside the artifact itself, with
// the final extension replaced. emit and assemble must agree on this one rule,
// so it lives here and nowhere else.
//
// artifacts/haks/sha256/30/46/3046….hak -> artifacts/haks/sha256/30/46/3046….nsym
func indexKey(artifactKey string) (string, error) {
extension := path.Ext(artifactKey)
if extension == "" {
return "", fmt.Errorf("artifact key %q has no extension", artifactKey)
}
return strings.TrimSuffix(artifactKey, extension) + ".nsym", nil
}
// resolveIndexKey is where emit writes an artifact's index and where assemble
// reads it from. On the zone that is beside the artifact; locally the indexes
// sit flat beside the data tree, so upstream's output and ours diff directly.
func resolveIndexKey(artifactKey, outDir string) (string, error) {
key, err := indexKey(artifactKey)
if err != nil {
return "", err
}
if outDir != "" {
return path.Base(key), nil
}
return key, nil
}
// checkArtifactKey fails closed when the key's embedded digest is not the
// digest of the bytes being emitted. Publishing an index under the wrong key
// silently pairs a manifest with the wrong artifact.
// artifact is hashed by streaming, so a multi-gigabyte hak is never resident.
func checkArtifactKey(artifactKey string, artifact io.Reader) error {
base := path.Base(artifactKey)
digest := strings.TrimSuffix(base, path.Ext(base))
if len(digest) != 64 {
return fmt.Errorf("artifact key %q does not name a sha256", artifactKey)
}
hash := sha256.New()
if _, err := io.Copy(hash, artifact); err != nil {
return fmt.Errorf("hash artifact: %w", err)
}
if got := hex.EncodeToString(hash.Sum(nil)); got != digest {
return fmt.Errorf("artifact key %q names digest %s but the file hashes to %s", artifactKey, digest, got)
}
return nil
}
+249
View File
@@ -0,0 +1,249 @@
package nwsync
import (
"crypto/sha1"
"crypto/sha256"
"encoding/hex"
"io"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"strings"
"sync"
"testing"
)
// fakeZone is a Bunny-shaped object store: PUT stores, GET reads, and the
// Checksum header is verified the way Bunny verifies it.
type fakeZone struct {
mu sync.Mutex
objects map[string][]byte
puts []string
failOn func(key string) bool // when true, the PUT fails
}
func newFakeZone(t *testing.T) (*fakeZone, func(string) string) {
t.Helper()
zone := &fakeZone{objects: map[string][]byte{}}
server := httptest.NewServer(zone)
t.Cleanup(server.Close)
getenv := func(name string) string {
switch name {
case "NWSYNC_STORAGE_ZONE":
return "sow-nwsync"
case "NWSYNC_STORAGE_PASSWORD":
return "write-key"
case "BUNNY_STORAGE_HOST":
return server.URL
}
return ""
}
return zone, getenv
}
func (z *fakeZone) ServeHTTP(w http.ResponseWriter, r *http.Request) {
key := strings.TrimPrefix(r.URL.Path, "/sow-nwsync/")
switch r.Method {
case http.MethodPut:
if z.failOn != nil && z.failOn(key) {
http.Error(w, "boom", http.StatusInternalServerError)
return
}
body, err := io.ReadAll(r.Body)
if err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
sum := sha256.Sum256(body)
if want := strings.ToUpper(hex.EncodeToString(sum[:])); r.Header.Get("Checksum") != want {
http.Error(w, "checksum mismatch", http.StatusBadRequest)
return
}
z.mu.Lock()
z.objects[key] = body
z.puts = append(z.puts, key)
z.mu.Unlock()
w.WriteHeader(http.StatusCreated)
case http.MethodGet:
z.mu.Lock()
body, ok := z.objects[key]
z.mu.Unlock()
if !ok {
http.Error(w, "not found", http.StatusNotFound)
return
}
_, _ = w.Write(body)
default:
http.Error(w, "unsupported", http.StatusMethodNotAllowed)
}
}
func (z *zoneSinkFixture) emit(t *testing.T, path string) EmitResult {
t.Helper()
result, err := Emit(EmitOptions{
ArtifactKey: artifactKey(t, path),
ArtifactPath: path,
Sink: z.sink,
})
if err != nil {
t.Fatalf("emit %s: %v", path, err)
}
return result
}
type zoneSinkFixture struct {
zone *fakeZone
sink sink
}
func newZoneFixture(t *testing.T) *zoneSinkFixture {
t.Helper()
zone, getenv := newFakeZone(t)
target, err := newZoneSink(t.Context(), getenv)
if err != nil {
t.Fatalf("zone sink: %v", err)
}
return &zoneSinkFixture{zone: zone, sink: target}
}
func sha1Of(body []byte) [20]byte { return sha1.Sum(body) }
func TestEmitUploadsBlobsThenIndex(t *testing.T) {
fixture := newZoneFixture(t)
dir := t.TempDir()
hak := filepath.Join(dir, "sow_test_01.hak")
body := []byte("texture bytes")
writeHak(t, hak, map[string][]byte{"bloodstain1.tga": body, "copy1.txi": body})
key := artifactKey(t, hak)
result := fixture.emit(t, hak)
index, err := indexKey(key)
if err != nil {
t.Fatal(err)
}
if _, ok := fixture.zone.objects[index]; !ok {
t.Fatalf("no index at %s; zone holds %v", index, fixture.zone.puts)
}
if result.BlobsWritten != 1 {
t.Errorf("uploaded %d blobs, want 1 (identical content shares a blob)", result.BlobsWritten)
}
// The index is the publication marker, so it must land after every blob it
// names — including its own sidecar.
last := fixture.zone.puts[len(fixture.zone.puts)-1]
if last != index {
t.Errorf("index landed at position %d of %d; it must be last", len(fixture.zone.puts), len(fixture.zone.puts))
}
for _, key := range fixture.zone.puts[:len(fixture.zone.puts)-1] {
if strings.HasPrefix(key, "data/sha1/") || key == index+".json" {
continue
}
t.Errorf("unexpected object uploaded before the index: %s", key)
}
}
func TestEmitSkipsBlobsAlreadyInTheZone(t *testing.T) {
fixture := newZoneFixture(t)
dir := t.TempDir()
hak := filepath.Join(dir, "sow_test_01.hak")
writeHak(t, hak, map[string][]byte{"bloodstain1.tga": []byte("blood")})
first := fixture.emit(t, hak)
if first.BlobsWritten != 1 {
t.Fatalf("first emit uploaded %d blobs, want 1", first.BlobsWritten)
}
second := fixture.emit(t, hak)
if second.BlobsWritten != 0 {
t.Errorf("re-emit uploaded %d blobs, want 0 (a blob name is its content)", second.BlobsWritten)
}
}
func TestEmitLeavesNoIndexWhenAnUploadFails(t *testing.T) {
fixture := newZoneFixture(t)
fixture.zone.failOn = func(key string) bool { return strings.HasPrefix(key, "data/sha1/") }
dir := t.TempDir()
hak := filepath.Join(dir, "sow_test_01.hak")
writeHak(t, hak, map[string][]byte{"bloodstain1.tga": []byte("blood")})
_, err := Emit(EmitOptions{
ArtifactKey: artifactKey(t, hak),
ArtifactPath: hak,
Sink: fixture.sink,
})
if err == nil {
t.Fatal("emit reported success after an upload failed")
}
for key := range fixture.zone.objects {
if strings.HasSuffix(key, ".nsym") {
t.Errorf("a half-emitted artifact published an index: %s", key)
}
}
}
func TestEmitRejectsAKeyThatDoesNotMatchTheFile(t *testing.T) {
fixture := newZoneFixture(t)
dir := t.TempDir()
hak := filepath.Join(dir, "sow_test_01.hak")
writeHak(t, hak, map[string][]byte{"bloodstain1.tga": []byte("blood")})
wrong := "artifacts/haks/sha256/00/11/" + strings.Repeat("0", 64) + ".hak"
_, err := Emit(EmitOptions{ArtifactKey: wrong, ArtifactPath: hak, Sink: fixture.sink})
if err == nil || !strings.Contains(err.Error(), "hashes to") {
t.Fatalf("emit published under a key that names another artifact: %v", err)
}
}
func TestAssembleReadsIndexesFromTheZone(t *testing.T) {
fixture := newZoneFixture(t)
dir := t.TempDir()
topBody := []byte("2da from sow_top")
assetBody := []byte("2da from the asset hak")
top := filepath.Join(dir, "sow_top.hak")
core := filepath.Join(dir, "sow_core_01.hak")
writeHak(t, top, map[string][]byte{"appearance.2da": topBody})
writeHak(t, core, map[string][]byte{"appearance.2da": assetBody, "bloodstain1.tga": []byte("blood")})
tlkPath := filepath.Join(dir, "sow_tlk.tlk")
if err := os.WriteFile(tlkPath, []byte("TLK V3.0 payload"), 0o644); err != nil {
t.Fatal(err)
}
fixture.emit(t, top)
fixture.emit(t, core)
if _, err := Emit(EmitOptions{
ArtifactKey: artifactKey(t, tlkPath),
ArtifactPath: tlkPath,
As: "sow_tlk.tlk",
Sink: fixture.sink,
}); err != nil {
t.Fatalf("emit tlk: %v", err)
}
result, err := Assemble(AssembleOptions{
ArtifactKeys: []string{artifactKey(t, top), artifactKey(t, core)},
TLKKey: artifactKey(t, tlkPath),
GroupID: 2,
Sink: fixture.sink,
})
if err != nil {
t.Fatalf("assemble: %v", err)
}
if result.Entries != 3 {
t.Fatalf("merged %d entries, want 3 (appearance.2da is shadowed, the TLK adds one)", result.Entries)
}
manifest, ok := fixture.zone.objects["manifests/"+result.SHA1]
if !ok {
t.Fatalf("no merged manifest in the zone; it holds %v", fixture.zone.puts)
}
entries, err := readManifest(manifest)
if err != nil {
t.Fatalf("parse merged manifest: %v", err)
}
for _, entry := range entries {
if entry.ResRef == "appearance" && entry.SHA1 != sha1Of(topBody) {
t.Errorf("appearance.2da resolved to the shadowed hak, not the first one given")
}
}
}
+1 -1
View File
@@ -18,7 +18,7 @@ bin=bin
# Keep in sync with internal/dispatch.Registry (Wired flag). # Keep in sync with internal/dispatch.Registry (Wired flag).
unwired=() unwired=()
wired=(assets depot hak module topdata wiki) wired=(assets depot hak module nwsync topdata wiki)
exit_of() { set +e; "$@" >/dev/null 2>&1; local c=$?; set -e; echo "${c}"; } exit_of() { set +e; "$@" >/dev/null 2>&1; local c=$?; set -e; echo "${c}"; }