Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,78 @@
# DockerHub discovery retry coalescing fails for an unbounded retry time
## Status
Open in the accepted live server runtime. Reproduced from the complete DockerHub
producer state and PostgreSQL authority on 2026-09-25. A minimal source fix and
PostgreSQL regression coverage have been added locally but have not been deployed
to the live runtime during the worker artifact validation.
## Impact
When DockerHub discovery cannot acquire a search page and delegates work without
an `available_after` timestamp, an already pending retry row cannot be coalesced.
The producer reports a generic `DockerHub discovery retry delegation failed`, the
source cycle fails, and the configured query does not advance. This can keep the
managed DockerHub producer in a restart loop while other sources remain healthy.
The complete producer history contained 265 failed DockerHub cycles with this
masked message. A protected diagnostic cycle using the captured zero-available-auth
state reproduced the original database exception exactly.
## Raw evidence
- Producer log: `build/live-trace-20260925/dockerhub-discovery.log`
- Captured runner state: `build/live-trace-20260925/runner_state_dockerhub.json`
- Complete retry/pass/cycle/lock capture:
`build/live-trace-20260925/dockerhub-discovery-db-raw.json`
- Unmasked exception:
`build/live-trace-20260925/dockerhub-unmasked-cycle-failure.json`
The unmasked exception is PostgreSQL `IndeterminateDatatype`, SQLSTATE `42P18`:
`could not determine data type of parameter $1`. It originates in
`ScannerDB.enqueue_discovery_retry()` while coalescing an existing `pending` row.
## Root cause
`app/scanner_db.py` used an untyped nullable placeholder in the PostgreSQL
expression:
```sql
WHEN available_after IS NULL OR ? IS NULL THEN NULL
```
When `available_after=None`, PostgreSQL has no typed expression from which it can
infer the placeholder type. SQLite accepts the same query, so existing SQLite
coalescing coverage did not expose the problem. Existing PostgreSQL integration
coverage inserted a retry but did not coalesce the same row with a null retry time.
The outer discovery helper intentionally replaced the original database exception
with a generic delegation error, which obscured the SQLSTATE in normal managed logs.
## Local correction
The placeholder is now explicitly typed as the schema's text timestamp form:
```sql
WHEN available_after IS NULL OR CAST(? AS TEXT) IS NULL THEN NULL
```
`tests/test_pipeline_postgres_integration.py` now coalesces the same retry with no
`available_after` value and checks that the existing row is reused.
The corrected statement was also executed against the live PostgreSQL schema in a
transaction and rolled back. It matched one pending row and completed without an
exception. Local focused results:
- DockerHub incremental discovery tests: 27 passed.
- SQLite retry lifecycle regression: 1 passed.
- PostgreSQL integration test: skipped locally because no disposable PostgreSQL DSN
was configured; the corrected SQL shape was verified transactionally against the
live schema without persisting a change.
## Operational note
The configured DockerHub credential pool was independently exhausted at capture
time: ten entries were invalid and the remaining entry was rate-limited. Correcting
retry coalescing preserves the failed work and stops this database error, but it does
not make an unavailable credential pool healthy.
@@ -0,0 +1,50 @@
# Terminal assignment status loses the durable scan deadline
Status: open and reproducible on the live `sec` validation cohort.
## Impact
`GET /api/v1/worker/assignments/{reservation_id}` can return
`deadlines.scan_deadline_at: null` for a resolved assignment even though the
durable terminal receipt contains the concrete scan deadline. The same response
labels the deadline set as immutable, so replay no longer faithfully exposes the
receipt that was committed at bundle acceptance.
Receipt, payload, scan-event, diagnostic, and reservation identities remain
correct. The loss is limited to terminal status readback of the scan deadline.
## Live evidence
The accepted Linux assignment used for this check had:
- a durable `bundle_accepted` receipt with a concrete `scan_deadline_at`;
- a later persisted `awaiting_receipt` progress event with
`scan_deadline_at: null`;
- a successful authenticated status response whose terminal receipt fields all
matched PostgreSQL, except that `scan_deadline_at` had become null.
The complete API response is retained locally as
`D:\truf\worker-linux-terminal-status-raw.json` with SHA-256
`0954e40b4ccb6921bbd572abb3d4898ee56e2955345de21ed7df5f85177536d1`.
The durable receipt is retained in
`build/live-trace-20260925/raw-evidence-expanded.json`.
## Root cause
`ScannerDB.remote_assignment_status()` loads the durable receipt and then calls
`result.update(observability)` (`app/scanner_db.py` around lines 16707-16710).
The observability object reconstructs its complete `deadlines` object from the
latest progress row. Its `scan_deadline_at` comes only from
`event.get('scan_deadline_at')` (`app/scanner_db.py` around lines 16588-16605).
Consequently, a later progress event with a null scan deadline replaces the
entire deadline object stored in the terminal receipt. This is a readback merge
problem; the persisted receipt itself is unchanged and correct.
## Expected correction
For resolved assignments, preserve the durable receipt's immutable deadline
object while overlaying only live observability fields such as latest progress,
known reason, and diagnostic authority. Add PostgreSQL regression coverage where
the accepted receipt has a concrete scan deadline and a later progress event has
a null deadline.
@@ -0,0 +1,72 @@
# Defect: Windows scan timestamps lose their UTC offset
## Status
Open and reproducible in the accepted Windows worker package validated on
2026-09-25. The live validation did not modify product code so that the tested
package remained identical to the accepted artifact.
## Symptom
Windows `scan.result_error` diagnostics can be stored with `occurred_at` about
the machine's local UTC offset in the future. In the validation environment the
offset was approximately three hours. The same naive timestamps also populate
Windows `target_scans.started_at` and `target_scans.ended_at`.
This can misorder diagnostics, distort time-window filters, and make the admin
panel show a scan event later than the server receipt that contains it.
## Live evidence
The accepted Windows package ran with a fresh state directory and one slot. In
the expanded live cohort, all 45 Windows `scan.result_error` diagnostics had an
`occurred_at` to server `received_at` delta between approximately 10,816 and
10,819 seconds. All 135 normal Windows scan-result rows used naive local start
and end timestamps with the same approximately three-hour displacement when
treated as UTC. The four Windows timeout-path rows and all 102 Linux rows had
normal small timing deltas.
The retained raw scanner material contained an explicit `+03:00` timestamp.
The persisted diagnostic retained the same wall-clock digits but labeled them
as UTC with `Z`. Monotonic scan durations and worker progress transport
timestamps remained correct.
Sensitive raw targets, scanner output, and credentials are retained only in the
restricted live evidence file and are intentionally not reproduced here.
## Root cause
`app/scanner.py` creates scan timestamps with naive local datetimes:
- `scan_target_result()` uses `datetime.now().isoformat()` for
`scan_started_at` and `timestamp` near lines 16339 and 16416.
- fallback result construction in `scan_single_target()` does the same near
lines 16844, 16863, and 16877.
Diagnostic construction parses those values and, when no timezone is present,
uses `occurred.replace(tzinfo=timezone.utc)` near line 16581. That operation
relabels local wall-clock time as UTC instead of converting it. The error is
visible on non-UTC hosts and is hidden on UTC Linux hosts.
## Expected behavior
All persisted protocol timestamps must identify a real UTC instant. Scanner
result timestamps should be emitted as timezone-aware UTC values, and legacy
naive values must not be silently reinterpreted as known UTC instants.
## Suggested correction and regression coverage
Emit `datetime.now(timezone.utc).isoformat()` at every result-construction site
and preserve the offset through serialization. Add a non-UTC-host regression
test that verifies:
- scanner start/end timestamps identify the actual UTC instant;
- diagnostic `occurred_at` precedes or closely tracks server `received_at`;
- Windows and Linux admin time-window filters return the same logical events;
- monotonic duration fields remain unchanged.
## Validation artifact
The expanded unrestricted evidence is retained at
`build/live-trace-20260925/raw-evidence-expanded.json`. It contains sensitive
raw internals and must not be published as a general operator report.
@@ -0,0 +1,50 @@
# Worker network OSError is logged as local I/O
Status: open, reproducible from the error-classification path.
## Observed behavior
The accepted Linux worker logged the following safe summary during the live
operator-experience validation:
```text
2026-09-25T13:10:05.909Z worker slot 0: local I/O operation failed
```
The complete worker journal shows that reservation 1689 had already entered
`assigned` at `13:09:46.641Z`. No runner phase had started. After the configured
error delay, the worker retried the same durable assignment, entered
`preparing` at `13:10:22.215Z`, completed it, and received one
`bundle_accepted` receipt at `13:10:28Z`.
The only operation between the successful `assigned` event and runner startup
is the authenticated assignment-status request in
`WorkerSlot.step()` (`app/remote_worker_client.py`, around lines 2141-2153).
The HTTPS client lets socket and transport `OSError` exceptions propagate.
`safe_worker_error_summary()` (`app/remote_worker_client.py`, around lines
128-135) maps every `OSError` to `local I/O operation failed`, including network
socket failures.
## Impact
- No assignment, result, or progress data was lost in the observed incident.
- The durable retry behavior worked and did not rescan the target.
- The operator-facing message misclassifies a transient network failure as a
local storage/filesystem problem, which can send diagnosis in the wrong
direction.
- The original exception type is not retained in the safe worker log, so the
transport subtype cannot be recovered after the fact.
## Evidence
- `build/live-trace-20260925/linux-worker-state-interim.tar.gz`
- Worker event sequences 1292-1301 and the matching history row for reservation
1689
- `app/remote_worker_client.py` status-read and safe-summary paths
## Expected correction
Classify network/socket failures before the broad `OSError` branch and emit a
safe transport-specific summary. Keep filesystem/storage `OSError` failures as
local I/O. Add a regression test covering an `OSError` raised by
`WorkerHTTPClient.status()` and verify that retry behavior remains unchanged.
@@ -0,0 +1,316 @@
# Scanner End-to-End Validation Evidence: 2026-09-22
## Verdict
The bounded production validation passed on the approved `sec` deployment.
It exercised the real PostgreSQL queue, protocol-2 remote assignment, existing
Windows/WSL worker, TruffleHog execution, bundle upload, durable receipt,
transactional ingestion, normalized findings/errors, and JSONL compatibility
projection paths.
The evidence consists of:
- 36 ordinary public-target scans under realistic production backlog;
- one separately managed non-live synthetic GitLab fixture scan proving the
positive finding path;
- exact append-region validation for `scan_results.jsonl` and
`found_secrets.jsonl` against PostgreSQL reconstruction;
- byte-identical restoration of the original production config;
- cleanup of private validation target files; and
- audited reopening of discovery and dispatch.
This is strong bounded production evidence, not a claim that every source,
failure mode, platform, detector, scale, or deployment environment is proven.
## Safety Envelope
- Only `sec` was used. `prod` was never touched.
- Raw targets, raw findings, credentials, device tokens, runtime YAML, worker
argv, the protected admin prefix, and edge markers were not printed.
- Configuration changes used managed Preview -> Save candidate -> Apply.
- Discovery and dispatch were paused and the runtime was drained before every
apply.
- Long operations and monitors ran detached and were observed with bounded
status polls.
- Host Caddy and X-UI remained outside the managed lifecycle.
- PostgreSQL remained the sole authority; JSONL was treated as a rebuildable
compatibility projection.
## Original Baseline
- Original active config SHA-256:
`f055a9f2506ab4fffa6953a95c2f6c07b1202e6558f4ce52bbf1b463ed6b1781`.
- Drained baseline high-water IDs:
- target queue: 1,534,069;
- result reservations: 880;
- target scans: 879;
- findings: 7;
- errors: 1,913.
- Runtime controls were revision 26, paused/paused, `drained`, blockers zero.
- One active worker device had contacted the server recently.
- All baseline orphan and referential invariants were zero.
Root-only baseline evidence:
- `/opt/truf-remote-server/staging/scanner-validation-pre.json`;
- `/opt/truf-remote-server/staging/scanner-validation-drained.json`;
- `/opt/truf-remote-server/staging/scanner-validation-pre-discovery-v4.json`;
- `/opt/truf-remote-server/staging/scanner-validation-pre-dispatch-newest-v5.json`.
## Defects Found and Corrected
The validation exposed defects that synthetic tests had not modeled precisely.
Each failure was contained by pause/drain, rollback, or failed-hold behavior
before dispatch was opened.
### Protected Config Parent
Discovery required the parent of a private file to be runtime-owned mode 0700,
while the deployed contract intentionally uses root-owned mode 0755
`/data/config` with runtime-owned mode 0600 documents. Writable runtime
directories still require private runtime ownership. Sensitive file parents now
also accept a non-link root-owned directory with no group/other write bits and
effective-user search access. The private file itself remains strictly checked.
### Supervisor Startup Locking
PostgreSQL readiness previously launched core children and discovery producers
while the supervisor held `control_lock`, then could perform another PostgreSQL
query under that lock. Child bootstrap/entrypoint authentication needed the same
lock and had bounded deadlines. Pipeline status refresh now happens before the
lock, structured snapshots use cached-only status, core children start before
source admission, and discovery producers use the source dependency gate.
### Strict Discovery Health
Ordinary Docker health intentionally tolerates periodic producer waits. Managed
lifecycle health now additionally uses explicit
`--require-discovery-producers` and rejects enabled producers that are absent,
blocked, never run, runtime-blocked, or waiting after a nonzero exit. Waiting
after a successful exit remains valid.
### Transient Strict-Health Probe
The first fixture apply encountered one bounded HuggingFace PostgreSQL
connection timeout after every core worker had started successfully. The host
lifecycle formerly performed only one strict probe after Docker health became
healthy. It now retries only health-category strict failures inside the existing
240-second runtime-health deadline. Identity and metadata errors remain
immediate failures, and persistent health failure still rolls back.
### Managed Claim Order
The remote assignment path already supported `oldest`, `newest`, and `balanced`
PostgreSQL admission, but the exact managed template omitted this field for the
three core sources. The optional field is now represented and semantically
validated. Temporary `newest` ordering allowed recent bounded discoveries to be
tested against the real 1.5-million-row queue without direct SQL mutation or
mass-hiding historical backlog. The restored original config omits the optional
field and therefore uses the normal `oldest` default.
### Candidate Base Authority
Candidate preparation originally used editor text that could represent an old
candidate rather than active config. This inherited an earlier intentionally
disabled Worker API setting. Candidate tools now read and hash-bind active
config bytes explicitly before deriving changes.
## Runtime Deployment Evidence
The corrected runtime was built as small derived immutable images rather than
modifying a running container. The final validation image ID was:
`sha256:5b9c86f68719d8c1f2358e0c4565dd2795ed608482968747feb961cf14908b5a`
Prior images remain under rollback tags. Image Entrypoint, Cmd, User, source
hashes, and in-image compilation were checked. An official lifecycle restart on
the final image completed `succeeded/succeeded`, reconciled, without a safe
category or failed hold.
Relevant local regression evidence accumulated during the run:
- runtime-document and worker-assignment tests: 45 passed;
- host lifecycle after transient-health retry: 34 passed, 4 platform skips;
- combined ACL/supervisor/health-focused suite: 246 passed, 9 platform skips;
- authenticated supervisor control class: 20 passed;
- focused compiles and `git diff --check`: passed.
## Bounded Discovery
The temporary candidate enabled one-page/one-result search settings for GitLab
and DockerHub and a four-item private custom file for HuggingFace. Dispatch
remained paused. Five successful cycles for each source completed before the
monitor's conservative time limit; no source cycle failed.
Because source cycles do not map directly to queue rows and uniqueness conflicts
consume sequence values, queue high-water deltas were not treated as exact
cohort membership. Eight new queue rows were observed: five GitLab pending and
three DockerHub deferred. No direct queue updates were made.
## Realistic 36-Scan Cohort
The worker processed exactly 36 new remote reservations, IDs 881 through 916,
while discovery remained paused. A fail-closed monitor paused dispatch at the
target and started drain. Final source mix:
| Source | Scans |
|---|---:|
| DockerHub | 11 |
| GitLab | 14 |
| HuggingFace | 11 |
| Total | 36 |
All 36 reservations were remote, resolved, acknowledged, and
`bundle_accepted`. They had 36 distinct queue IDs, bundle IDs, and scan event
IDs, and every reservation had a receipt, payload hash, and execution-snapshot
hash.
### Results
| Source | Result summary |
|---|---|
| DockerHub | 8 clean, 3 degraded |
| GitLab | 12 clean, 1 retryable API error, 1 permanent not-found |
| HuggingFace | 11 clean |
- Queue completion: 34 done, one deferred, one failed; no row remained fenced.
- Findings: zero, a valid outcome for random public targets.
- Errors: exactly two GitLab errors with queue dispositions matching their
retryable/permanent categories.
- Quarantine: zero new rows.
- Bundle/projection capacity after completion: zero items and zero bytes.
- Existing unrelated keycheck capacity was unchanged.
### Bundle and Projection Invariants
- 36 acknowledged bundles contained 110 frames.
- All bundle identities and counts matched their reservations and scans.
- Acknowledged physical `.trb` files were absent only after both pipeline
artifact records reached durable `deleted` state, as designed.
- 36 scans used `raw_result_storage=normalized_v2`.
- 36 compatibility rows used expected bounded reconstruction.
- Exactly 36 projection jobs completed, one per scan, without duplicates or
errors; all projection capacity was released.
- Physical append evidence covered 36 `scan_results` records and two
`scan_errors` records.
- Every registered append generation/offset/length existed and matched its
payload SHA-256, record count, and required JSON structure.
- `scan_results.jsonl` grew by exactly 83,752 bytes.
- `found_secrets.jsonl` did not change, matching zero random-target findings.
- All global queue/reservation and orphan invariants remained zero.
Root-only evidence:
- post snapshot:
`/opt/truf-remote-server/staging/scanner-validation-post-dispatch-newest-v5.json`,
SHA-256
`45f9db81e88dbcbe4d6a04094dd1d892df06dd3f7cdd05eb9280c19da54aed94`;
- aggregate report:
`/opt/truf-remote-server/staging/scanner-validation-cohort-report-v5.json`,
SHA-256
`e0826587e4ba667de10a994d5842aae5fae8fa11dc071210c202b38b0e643bf4`.
## Controlled Positive Fixture
Random public targets produced no finding, so a separate one-target run used a
public GitLab project whose README declares that its secret examples are
generated and non-live. No detector or verification behavior was weakened.
- Fixture queue ID: 1,534,100.
- Reservation ID: 917.
- The immutable Git plan bound the approved exact head commit
`2a09bd6767d39b95cf39ce4b5fd210721275d503`.
- The reservation became acknowledged with `bundle_accepted` and a durable
receipt.
- Queue completion was `done` with no remaining reservation fence.
- Target scan status was `found` with 116 findings and zero errors.
- All 116 findings used the existing OpenAI detector.
- Verified count was zero, consistent with the unchanged no-verification policy.
- All findings had distinct finding UIDs, nonempty identities, private raw
material, redaction different from raw material, correct secret hashes, and
complete non-omitted compatibility payloads.
- No raw finding value was emitted by validation tooling.
- Bundle retirement and both pipeline artifact tombstones were correct.
- The single projection job completed and released capacity.
- The registered `scan_results` region contained one record with exactly 116
findings and zero errors.
- The registered `found_secrets` region contained exactly 116 records.
- Both physical append regions matched the database payload SHA-256 and were
byte-identical to fresh PostgreSQL compatibility reconstruction.
- No new quarantine row was created.
Root-only fixture report:
`/opt/truf-remote-server/staging/scanner-validation-fixture-report-v2.json`
SHA-256:
`243a16b24bf9ae898bfdeb8f857c56ef1cf78e12e637730ff5b0674a250a4984`
## Restoration and Final State
The original 36,354 config bytes were passed through managed Preview, saved as
a candidate, and applied through the host agent. Preview preserved the exact
original SHA-256 and reported 169 semantic reversions.
- Restore Save operation:
`af2a3cc4-90ae-5831-b932-bbe78ceb2cab`.
- Restore Apply operation:
`b12a2fc4-202a-578e-8dcf-b0a88cb028ef`.
- Apply terminal state: `succeeded/succeeded`, reconciled, category `None`.
- Active and candidate config SHA-256 both equal the original
`f055a9f2506ab4fffa6953a95c2f6c07b1202e6558f4ce52bbf1b463ed6b1781`.
- Lifecycle preflight and strict Worker API/discovery health passed.
- Runtime and edge were healthy; no failed hold existed.
- All private validation target/evidence files were removed.
- The root-only original backup was retained for audit.
The drained post-restore snapshot is root-only at
`/opt/truf-remote-server/staging/scanner-validation-post-restore-drained-v1.json`,
SHA-256
`de4bd98d728dc551b10712a4afb6be2db0ac9111428d46fa5dbb16f8d2d611ca`.
Final audited control transitions advanced revision 46 to 49 in this order:
1. cancel drain;
2. resume discovery;
3. resume dispatch.
Final state was discovery open, dispatch open, drain `normal`. The existing
worker contacted the server within five minutes and immediately received normal
production work. A live assignment after reopening is expected and is not a
drain blocker because drain is no longer requested.
The final post-resume snapshot had zero orphan/referential invariants and
preserved the original config SHA-256:
`/opt/truf-remote-server/staging/scanner-validation-post-resume-final-v1.json`
SHA-256:
`42e19edae09550693d563b74631430cb1d2c1b807d636ccff20e545cebec3c2d`
External route checks through existing host Caddy returned:
- invalid Worker API authentication: 401;
- unauthenticated protected admin route: 401;
- unrelated path: 404.
Host-agent, Caddy, and X-UI services remained active. Caddy and X-UI were not
lifecycle targets.
## Residual Limits
This validation does not prove:
- long-duration soak or high-concurrency behavior;
- every detector and verification provider;
- every source mode, browser, OS, architecture, or network failure;
- every secrets/config mutation and rotation case;
- HA or multi-server operation;
- resistance to an independent penetration test; or
- correctness of arbitrary unsupported Compose, ingress, or proxy layouts.
Within its declared scope, the real queue, worker, scanner, ingestion,
findings, error, compatibility, restoration, and resumed-production paths all
produced internally consistent durable evidence.
+339
View File
@@ -0,0 +1,339 @@
# End-to-End Scanner Validation
This runbook validates the real discovery, remote-worker, result-ingestion, and
compatibility-projection path with a bounded cohort of 30-40 targets. It is an
evidence procedure, not a claim that every repository feature and environment
has been proven correct.
The completed 2026-09-22 production evidence is recorded in
[`end-to-end-scanner-validation-2026-09-22.md`](end-to-end-scanner-validation-2026-09-22.md).
## Validation Questions
The run must answer all of the following:
1. Does bounded discovery create the expected immutable queue identities?
2. Do remote workers receive each authoritative assignment with the correct
source, target identity, execution snapshot, and lease fencing?
3. Does each worker run the intended TruffleHog scan and upload a canonical
protocol-2 result bundle?
4. Does the server accept a result exactly once and make receipt replay
idempotent?
5. Does the ingester transactionally connect the reservation, queue row,
target scan, findings, errors, and bundle record?
6. Does the JSONL projector reproduce the Windows-compatible
`scan_results.jsonl` and `found_secrets.jsonl` structures without becoming
a second source of truth?
7. Are naturally found or controlled-canary findings stored with safe identity,
location, detector, verification, redaction, and provenance fields?
8. After the test, are queues settled, projections caught up, no pipeline item
quarantined, and the normal production configuration restored?
## Authorities and Expected Data Flow
The authoritative sequence is:
```text
discovery cycle
-> target_queue
-> result_reservation / immutable remote assignment
-> worker TruffleHog execution
-> canonical .trb upload
-> durable accepted receipt
-> result ingester transaction
-> target_scans + findings + errors + queue completion
-> projection_jobs
-> scan_results.jsonl + found_secrets.jsonl
```
PostgreSQL is authoritative. Result bundles are durable pipeline artifacts.
JSONL files are rebuildable compatibility projections and may lag briefly.
`accepted` proves durable server receipt; `ingested` proves the database
transaction completed. These states must not be treated as synonyms.
## Safety Rules
- Use only the approved test server and approved worker devices.
- Never print device tokens, source credentials, raw secrets, active runtime
YAML, worker argv, or complete unredacted findings into a terminal/log.
- Query finding structure using IDs, hashes, redacted values, lengths, booleans,
detector names, and location metadata. Review any raw secret only through the
already protected admin workflow if explicitly required.
- Apply and restore configuration through Preview -> Save candidate -> Apply.
Do not edit the active runtime document in place.
- Pause discovery and dispatch and drain before each runtime-document apply.
- Record the original config SHA-256 and require byte-identical restoration at
the end.
- Use a unique test-run label and database high-water marks. Never infer the
cohort from wall-clock time alone.
- Do not delete queue, bundle, finding, projection, or receipt evidence to make
a failed test look clean.
## Bounded Test Configuration
Use a temporary candidate derived from the active document. Preserve all
secrets and unrelated settings. The exact candidate must be reviewed before it
is applied.
### DockerHub
- `mode: search`
- `pages: 1`
- `per_page: 1`
- `docker_images_per_repository: 1`
- Use a reviewed finite query list for the test window.
One query is consumed per source cycle. An already-known or unsuitable search
result can produce no new queue row, so the number of cycles is not the cohort
size.
### GitLab
- `mode: search`
- `pages: 1`
- `per_page: 1`
- Use a reviewed finite query list for the test window.
- Keep current age, commit-boundary, exact-ref, visibility, and history-depth
safety controls unless the test explicitly records a different expectation.
### HuggingFace
HuggingFace recent discovery does not support a real `per_page: 1` keyword
test. Its API runner fetches newest-modified Spaces and the current API page
size is fixed at 100; the configured query is only a rotation placeholder.
For a bounded cohort, use `mode: custom` with a reviewed private `target_file`
containing a small list of Space IDs. Do not claim that changing `per_page` to
1 bounded this source when it did not.
### Recommended Cohort
Target 36 authoritative terminal scans:
- 16 DockerHub immutable digest targets;
- 16 GitLab exact-ref/commit-planned targets;
- 4 HuggingFace custom Space targets.
The exact split may vary between 30 and 40 when discovery deduplicates known
targets or a target becomes permanently inaccessible. Continue only until the
recorded cohort reaches the agreed bound. Do not inflate discovery simply to
hit an exact aesthetic number.
## Positive-Finding Requirement
A random public cohort may correctly produce zero findings. Zero findings
cannot validate the finding-storage and `found_secrets.jsonl` path.
Include at least one separately identified, non-live controlled fixture that is
expected to trigger an already approved detector. The fixture must contain no
usable credential. Record its expected detector and identity before scanning.
Do not weaken verification, introduce a new detector, or publish a real secret
merely to force a positive result.
If no approved positive fixture is available, report the finding path as
unverified by this run even if all zero-finding scans succeed.
## Phase 1: Baseline
With runtime healthy, record a secret-safe baseline:
- active config SHA-256 and semantic config SHA-256;
- runtime-control revision and open/paused/drain state;
- enabled source set and active worker package manifests;
- remote worker/device count, recent contact, and package capability match;
- high-water IDs for `target_queue`, `result_reservations`, `target_scans`,
`findings`, `errors`, `result_bundles`, and `projection_jobs`;
- queue counts by source and status;
- active reservation count and oldest age;
- pipeline worker readiness and capacity counters;
- pending/leased/quarantined bundle, projection, and keycheck counts;
- current projection stream/cursor identity;
- byte size and final complete-line identity of active JSONL files.
The baseline collector must print aggregates and hashes only. It must not emit
targets, assignment payloads, tokens, raw findings, or runtime documents.
## Phase 2: Apply the Test Candidate
1. Pause discovery.
2. Pause dispatch.
3. Start drain and wait for blocker count zero and `drained`.
4. Preview the bounded candidate and review the semantic diff.
5. Save and apply the candidate through the host-agent lifecycle.
6. Require reconciled `succeeded`, no failed hold, strict runtime health, edge
health, admin health, and Worker API health.
7. Cancel drain, then resume discovery and dispatch in that order.
Do not continue if the lifecycle operation rolls back or enters failed hold.
## Phase 3: Build and Freeze the Cohort
Record the baseline `target_queue.id` high-water mark. Let the bounded sources
cycle until 30-40 new eligible queue rows have been created after that mark.
Then:
1. Pause discovery so the cohort cannot grow.
2. Leave dispatch open until the selected queue rows settle.
3. Record cohort queue IDs and only their safe identities: source, normalized
target hash, query hash, immutable planning kind, and creation order.
4. Separate deduplicated, permanently inaccessible, deferred, retried, and
actually assigned items. Do not count an API result as a scan.
The authoritative cohort is a fixed set of queue IDs, not "whatever completed
during the same hour."
## Phase 4: Observe Remote Execution
For every cohort queue ID, verify:
- no more than one current authoritative reservation;
- assignment package/platform capability matches the registered worker;
- lease token and execution snapshot are bound but never printed;
- Docker targets are immutable `repo@sha256` identities;
- GitLab targets have the intended exact planning/ref identity;
- HuggingFace targets use the direct Space execution kind;
- terminal report classification is success, permanent target failure, or
retryable provider failure as designed;
- retries preserve queue identity and increment attempts without creating a
second authoritative acceptance;
- accepted receipt replay returns the same durable result.
Physical work can repeat after a lease expiry or network partition. Correctness
means fencing permits one authoritative acceptance and one queue completion,
not that duplicate physical execution is impossible.
## Phase 5: Validate Bundles and PostgreSQL Structure
For each accepted result, validate without dumping body content:
- bundle exists at the registered private relative path;
- bundle byte count and SHA-256 match database metadata;
- bundle schema/version, event ID/hash, reservation ID, queue ID, source,
normalized target identity, execution snapshot identity, and scan policy are
internally consistent;
- result is ingested exactly once;
- `target_queue.target_scan_id` references the corresponding `target_scans.id`;
- queue completion is applied once with a terminal disposition;
- `target_scans.queue_id` and claim lease identity refer back to the cohort row;
- `target_scans.findings_count` and `error_count` equal actual child-row counts;
- every finding/error references the same target scan, source, cycle, and run;
- no legacy `raw_result_json`, publication outbox row, or orphan relation is
introduced;
- ingested bundle credit and pipeline-capacity counters are released exactly
according to the durable state machine;
- no cohort item enters `pipeline_quarantine`.
Aggregate checks must cover the entire cohort. Additionally inspect a small
redacted structural sample from every source and every terminal disposition.
## Phase 6: Validate Findings
For every finding in the cohort, inspect structure only:
- stable `finding_uid` and finding fingerprint;
- detector name/type and verification flag;
- source, target hash, file path, line/commit/source timestamp where applicable;
- redacted secret and secret/detector hashes;
- provider and credential-kind enrichment;
- required-context and raw-payload-omitted flags;
- bounded source metadata and enrichment JSON decode successfully;
- no unexpected raw-secret exposure in logs, queue rows, assignment metadata,
admin list views, or compatibility scan summaries.
For the controlled positive fixture, require the expected finding to exist in
PostgreSQL and to project once to `found_secrets.jsonl`.
## Phase 7: Validate Windows-Compatible Files
Use the active configured `global.results_dir`. The relevant compatibility
outputs are:
- `scan_results.jsonl` for one sanitized scan event per projected scan;
- `found_secrets.jsonl` for projected finding events;
- their publication ledgers, active stream metadata, and rotated segments;
- per-service keycheck result files only if keychecks run for the finding.
For the cohort, verify:
- every required `projection_job` reaches `completed`;
- projector cursor and append ledger advance monotonically;
- each cohort scan event appears exactly once by `scan_event_id`;
- each cohort finding appears exactly once by `finding_uid`;
- JSON lines parse and match the current compatibility schema;
- scan summaries match PostgreSQL counts and terminal status;
- finding projections are redacted as designed and preserve safe provenance;
- active and rotated segments together contain the events; checking only the
active file is insufficient when rotation occurs;
- no torn-tail quarantine, duplicate append, skipped cursor, or unpublished
completed job exists.
These files should have the same logical structure as the Windows deployment,
but path separators and host/container root paths are platform-specific.
## Phase 8: Queue and Pipeline Closure
After all cohort rows settle, require:
- 30-40 cohort queue rows accounted for by terminal, deferred, or explicitly
classified retry state;
- no expired active reservation remains unreaped;
- no queue row has multiple authoritative accepted results;
- no accepted result remains un-ingested beyond the bounded pipeline window;
- no completed scan remains unprojected beyond the bounded projector window;
- no stale pipeline lease or capacity leak;
- no unexpected quarantine;
- runtime, Worker API, edge, host-agent, host Caddy, and unrelated host service
health remain good.
The final report must show counts for discovered, deduplicated, assigned,
retried, accepted, ingested, projected, succeeded, skipped/permanent,
retryable/deferred, findings, errors, and quarantines.
## Phase 9: Restore Production Configuration
1. Pause discovery and dispatch.
2. Drain to zero blockers.
3. Apply the exact original runtime document through the normal lifecycle.
4. Require byte-identical original config SHA-256, reconciled lifecycle success,
strict health, and no failed hold.
5. Cancel drain, resume discovery, then resume dispatch according to the
original control state.
6. Confirm worker contact and normal post-test assignment flow.
Do not restore by manually editing YAML or replacing files behind the
host-agent.
## Pass Criteria
The run passes only when:
- at least 30 and at most 40 fixed-cohort rows are fully accounted for;
- all accepted cohort results ingest exactly once;
- all required cohort projections complete exactly once;
- queue/reservation/scan/finding/error/bundle relationships are consistent;
- the controlled positive finding reaches PostgreSQL and
`found_secrets.jsonl`, or the report explicitly marks positive-finding
validation incomplete because no approved fixture existed;
- no unexplained retry, orphan, duplicate acceptance, capacity leak,
quarantine, failed hold, or projection gap remains;
- the original production config is restored exactly and services are healthy.
Any failure must retain its operation IDs, queue IDs, reservation IDs, hashes,
safe categories, and aggregate evidence for diagnosis. A partial pass must not
be reported as "100% scanner correctness."
## Evidence Report
Append or link a dated report containing:
- environment and worker package identities;
- original/test/restored config hashes;
- cohort definition and aggregate source split;
- lifecycle operation IDs for test apply and restore;
- queue and pipeline baseline/final aggregates;
- per-stage reconciliation counts;
- redacted structural examples for a scan, an error/skip, and a finding;
- JSONL/ledger reconciliation counts;
- deviations, retries, quarantines, and unresolved questions;
- final verdict with explicit tested and untested boundaries.
+105
View File
@@ -0,0 +1,105 @@
# Extended Live Validation
Date: 2026-09-26
Run: `e6ac8aec-ec20-4ba4-a924-abe5ee95d82c`
Conclusion: `pass_with_documented_deviations`
## Scope
The production validation ran one native Windows worker slot and one WSL/Docker
worker slot through discovery, assignment, execution, progress, diagnostics,
bundle ingestion, projection, capacity release, and scheduled keycheck. The run
started at `2026-09-26T00:29:30.500768Z`; terminal server evidence was captured
at `2026-09-26T03:54:04.573328Z`.
This report contains derived counts, classifications, and cryptographic
identities only. Raw provider values and worker payloads remain in private
evidence storage.
## Outcome
- 314 reservations reached terminal outcomes: 309 acknowledged and 5 refunded.
- All 309 accepted bundles were ingested in one attempt, projected, settled, and
released; no unresolved reservation or duplicate receipt remained.
- The accepted cohort covered DockerHub (105), GitLab (102), and Hugging Face
(107), split across Windows (167) and WSL (147).
- 3,042 progress events covered all 309 accepted reservations with no duplicate
reservation sequence.
- All 348 projection jobs completed and released in one attempt with no error.
- The sealed terminal cut at controller revision 143 had no live assignment,
pre-commit bundle, capacity use, publication outbox item, quarantine row,
waiting lock, or blocker.
The 309 scan outcomes were 174 clean, 75 degraded, 56 error, and 4 found. Five
findings were recorded and none were verified. The run recorded 100 classified
scan errors, led by 71 download failures; source distribution was GitLab 92,
DockerHub 7, and Hugging Face 1.
## Diagnostics And Keycheck
The validation captured 61 worker diagnostics: 52 result errors, 4 stage
timeouts, 4 runner protocol failures, and 1 client process failure. Fifty-one
were retryable and ten were nonretryable. The records remain available in the
private raw bundle for exact-body investigation.
The captured keycheck cohort contained 39 candidates. All completed in one
attempt, released capacity, linked to a result, and projected. Results were 35
invalid or revoked and 4 no-balance; 37 came from API execution and 2 from
cache. At the terminal cut the global keycheck queue had 184 completed and no
pending, leased, or deferred candidate.
Nine pending DockerHub discovery retry rows with zero attempts remained as
expected durable discovery backlog, not as leaked worker-pipeline work.
## Evidence Integrity
- Server chain: 202 valid records, sequence `-1..199`, with all 575 referenced
objects (1,245,586,992 bytes) verified.
- Server NDJSON SHA-256:
`a6beb1864c540fc5f22b2b647730a39e45dfddb10cf5ceff5e0181eb1a8f0bd8`.
- Server run SHA-256:
`5ae58df5f67a8d2a8d4e84d73262a6e7009e03dcbcc9ea176537b384d4e693c3`.
- Worker chain: 158 valid records, sequence `0..157`, with all 2,136 referenced
objects (7,733,582,955 bytes) verified.
- Worker NDJSON SHA-256:
`03a04a0da5a2db6bd02f2aaed41c191554b135ec287fbc1e2d9998d77da59898`.
- Worker final payload SHA-256:
`f652e5f98fb23f9aa2fb4e44e10168a931ea973be616fb09dbdb72001ce961ad`.
The machine-readable verification manifest is at
`build/extended-live-validation/runs/e6ac8aec-ec20-4ba4-a924-abe5ee95d82c/verification-manifest.json`.
The complete server evidence root is retained in private production storage;
the local run root contains compact server artifacts and complete worker
evidence.
## Observed Deviations
The authenticated `recheck all` operation was accidentally used where only two
pending candidates should have been rechecked. This caused 35 broad provider
checks contrary to `KEYCHECK-001`. The command processed all 35, skipped none,
returned success, produced no failed operation, and all resulting effects
settled. This is an operator-scope deviation, not an approved workflow change.
Six monitor samples encountered transient database statement timeouts. Every
monitor error recovered, and the evidence chains remained valid. The runtime
and edge each recorded zero restart during the validation.
## Remaining Defects
- Windows GitLab filename-too-long checkout recovery exists locally but is not
deployed.
- Invalid API-key classification has a local fix that is not deployed.
- A roughly 20-second WSL clock-domain monotonic failure remains open.
- Generic `WorkerContractError` diagnostics can lose structured field detail.
- Monitor aggregate queries can exceed their statement timeout under load.
Focused regression coverage for Git checkout recovery, Hugging Face long paths,
and direct remote credentials passed: 145 tests in 155.26 seconds.
## Restore
Production was restored with compare-and-swap transitions from revision 143 to
146: cancel the validation drain, resume discovery, then resume
dispatch. Final state was dispatch open, discovery open, drain normal, healthy
core services and producers, and exactly one live assignment on each worker.
No further validation probe is required.
@@ -0,0 +1,77 @@
# TRUF worker: шпаргалка Docker
Откройте Bash или PowerShell в корне bundle с `compose.yaml`, worker image archive
и helper scripts. Нужен Docker Engine с Compose v2 либо Docker Desktop.
## Один раз
Создайте локальный `worker-install.yaml` через текстовый редактор:
```yaml
server: https://pregnant.horsecock.store
token: PASTE_DEVICE_TOKEN_HERE
parallelism: 1
```
Bash:
```sh
chmod 600 ./worker-install.yaml
chmod 700 ./workerctl.sh
docker load -i truf-worker-linux-x86_64.tar.gz
docker compose run --rm -T worker install --config - < ./worker-install.yaml
rm -- ./worker-install.yaml
docker compose run --rm worker doctor --json
```
PowerShell:
```powershell
docker load -i .\truf-worker-linux-x86_64.tar.gz
Get-Content -Raw .\worker-install.yaml |
docker compose run --rm -T worker install --config -
Remove-Item -LiteralPath .\worker-install.yaml
docker compose run --rm worker doctor --json
```
## Каждый день
Bash:
```sh
docker compose up -d
./workerctl.sh status
./workerctl.sh attach --follow-seconds 300
./workerctl.sh watch --follow-seconds 300
./workerctl.sh stop --timeout 120 --json
docker compose down
```
PowerShell:
```powershell
docker compose up -d
.\workerctl.ps1 status
.\workerctl.ps1 attach --follow-seconds 300
.\workerctl.ps1 watch --follow-seconds 300
.\workerctl.ps1 stop --timeout 120 --json
docker compose down
```
`q` или `Ctrl-C` отсоединяет `attach`/`watch`, но не останавливает worker.
`docker compose down` без `--volumes` сохраняет identity, token, незавершённую
работу и history в `truf-worker-data`.
## Диагностика
```sh
./workerctl.sh status --json
./workerctl.sh logs --tail 200
./workerctl.sh logs --follow --follow-seconds 300
./workerctl.sh history --limit 50
./workerctl.sh doctor --json
```
Не используйте `docker compose down --volumes`. Clean stop должен вернуть
`state: stopped`, `drained: true`, `exit_code: 0`; только после этого выполняйте
`docker compose down`.
+70
View File
@@ -0,0 +1,70 @@
# TRUF worker: шпаргалка Linux без Docker
Нужны системный Python 3.12 и package-local Git, TruffleHog и Python dependencies
из проверенного artifact. Native executable authority должна находиться под
root-owned путём, поэтому один раз перенесите распакованный package в `/opt`:
## Один раз
```sh
sudo mv ./truf-worker-linux-x86_64 /opt/truf-worker
cd /opt/truf-worker
```
Подготовьте доверенные ownership и permissions package. Скрипт оставляет
application code приватным для текущего user, а native Git и TruffleHog -
неизменяемыми для него:
```sh
sudo ./prepare-worker.sh
```
Затем создайте `~/.config/truf/worker-install.yaml` через локальный текстовый
редактор:
```sh
install -d -m 700 ~/.config/truf
nano ~/.config/truf/worker-install.yaml
```
```yaml
server: https://pregnant.horsecock.store
token: PASTE_DEVICE_TOKEN_HERE
parallelism: 1
```
Затем выполните:
```sh
chmod 600 ~/.config/truf/worker-install.yaml
./truf-worker install --config ~/.config/truf/worker-install.yaml
rm -- ~/.config/truf/worker-install.yaml
./truf-worker doctor --json
```
## Каждый день
```sh
./truf-worker start --startup-timeout 30
./truf-worker status
./truf-worker attach --follow-seconds 300
./truf-worker watch --follow-seconds 300
./truf-worker stop --timeout 120 --json
```
`q` или `Ctrl-C` отсоединяет `attach`/`watch`, но не останавливает worker.
Portable package не устанавливает systemd unit: после перезагрузки выполните
`start` из того же package под тем же OS user.
## Диагностика
```sh
./truf-worker status --json
./truf-worker logs --tail 200
./truf-worker logs --follow --follow-seconds 300
./truf-worker history --limit 50
./truf-worker doctor --json
```
Clean stop должен вернуть `state: stopped`, `drained: true`, `exit_code: 0`.
Иначе сохраните state и диагностику; не удаляйте work или bundles.
@@ -0,0 +1,55 @@
# TRUF worker: шпаргалка Windows
Откройте PowerShell в корне распакованного Windows package. Обычный запуск не
требует прав администратора.
## Один раз
```powershell
Set-ExecutionPolicy -Scope Process Bypass
notepad .\worker-install.yaml
```
Заполните открытый файл:
```yaml
server: https://pregnant.horsecock.store
token: PASTE_DEVICE_TOKEN_HERE
parallelism: 1
```
Затем выполните:
```powershell
.\prepare-worker.ps1
.\truf-worker.cmd install --config .\worker-install.yaml
Remove-Item -LiteralPath .\worker-install.yaml
.\truf-worker.cmd doctor --json
```
## Каждый день
```powershell
.\truf-worker.cmd start --startup-timeout 30
.\truf-worker.cmd status
.\truf-worker.cmd attach --follow-seconds 300
.\truf-worker.cmd watch --follow-seconds 300
.\truf-worker.cmd stop --timeout 120 --json
```
`q` или `Ctrl-C` отсоединяет `attach`/`watch`, но не останавливает worker.
После перезагрузки снова выполните только `start`; token уже находится в
приватном installed config.
## Диагностика
```powershell
.\truf-worker.cmd status --json
.\truf-worker.cmd logs --tail 200
.\truf-worker.cmd logs --follow --follow-seconds 300
.\truf-worker.cmd history --limit 50
.\truf-worker.cmd doctor --json
```
Clean stop должен вернуть `state: stopped`, `drained: true`, `exit_code: 0`.
Иначе сохраните state и диагностику; не удаляйте work или bundles.
+454
View File
@@ -0,0 +1,454 @@
# Remote Scan Worker Operations
This runbook covers the opt-in remote worker boundary. Production defaults remain
disabled: `supervisor.worker_api.enabled` and
`supervisor.worker_api.admin.enabled` are both `false` in
`app/config.linux.yaml`. Enabling either one, applying schema changes, or starting
the production edge requires a separate reviewed rollout.
Remote workers are trusted clients. Their only operator-authored runtime settings
are the HTTPS server origin, one opaque device token, and a positive slot count
`N`. Target/source settings, immutable plans, scanner policy, limits, and only the
credentials needed for an assignment come from the server. Protocol-2 package
schema 3 manifests advertise exact `(source, platform, planning_kind)`
capabilities. The distributed core package contains GitLab `exact_git_v1`,
DockerHub `docker_direct_v1`, and HuggingFace `huggingface_space_v1`; GitHub is a
legacy optional capability and is not in the core profile. Detailed keycheck
remains server-only after bundle ingestion; clients must not receive keycheck
configuration or run keycheckers.
## Isolated verification
Run each verifier independently from the root of the isolated
`D:\truf-workers` checkout. Do not run plain `docker compose`, production
Compose files, import overrides, unrestricted pytest, or broad Docker cleanup.
Do not run these verifiers concurrently. They create random, ownership-labelled
resources and never use production data, credentials, volumes, or image tags.
### Main Docker E2E
From Linux or WSL, with the already-built local images
`truf-worker-test:runtime` and `truf-worker-test:test` and an already ignored
`docker/test-results/latest.json`:
```sh
python3 -I -S -B docker/verify.py
```
The verifier invokes only `compose.e2e.yaml`, does no build or pull, and normally
removes only its ownership-verified resources. It retains artifacts after a
failure. Use `--keep` only when a reviewed investigation needs stopped artifacts;
record the printed project name and never substitute prune, broad `down`, or
`down --volumes` commands.
### Packaged Windows/Linux client E2E
From Windows, with Windows Python, `wsl.exe`, passwordless `sudo -n docker` in the
selected WSL distribution, the built portable directory, and the already-built
Linux worker and test images:
```powershell
python -I -S -B docker/verify_packaged_workers.py `
--windows-artifact dist/truf-worker-windows-x86_64 `
--linux-image truf-remote-worker:linux-x86_64 `
--test-image truf-worker-test:test `
--wsl-distro Ubuntu-24.04
```
This single gate runs real packaged Windows and Linux scanners at `N=2`, tests a
server outage and restart recovery, and compares normalized cross-platform
evidence. It does not use Compose. Successful resources are removed unless
`--keep` is supplied; failures retain the labelled resources and the reported
`build/pwe-*` evidence directory.
### Edge E2E
From Windows with the same WSL Docker access and already-built
`truf-worker-test:test` and `truf-edge-e2e:test` images:
```powershell
python -I -S -B docker/verify_edge_e2e.py --wsl-distro Ubuntu-24.04
```
This gate uses only random labelled resources and a private synthetic backend. It
checks authenticated routes, two-failure admin bans, restart persistence,
automatic expiry, SSH-equivalent unban behavior, forwarded-header handling, and
worker availability from the banned admin IP. Success removes its resources;
failure reports the retained owned inventory and evidence path.
## Client bootstrap
Use a separately issued token for every device. The server stores only its
SHA-256 digest. Do not place a real token in documentation, source control,
shell transcripts, support output, or process diagnostics. The client validates
the server certificate and accepts only an HTTPS origin without credentials,
path, query, or fragment. `N` must be between 1 and 128; it bounds local occupied
slots but never overrides the user's server cap across devices.
Generated state, pending bundles, and work directories are recovery data, not
additional source/provider configuration. Preserve them across restarts until
the server authoritatively resolves the corresponding slots.
### Artifact source and distribution
Worker tools come from a reviewed release checkout; they are not installed
piecemeal on each client. The Windows builder and Linux worker package/image
targets assemble and verify the complete worker authority, Python runtime where
applicable, Git helper, TruffleHog binary, detector policy, CA roots, and
hash-locked dependencies. They deliberately exclude PostgreSQL tools,
server/runtime authority, provider implementations, detailed keycheck code,
server credentials, and database credentials.
There is no public worker download, image registry, installer, or automatic
updater. Build each release once in a controlled release environment, retain its
generated manifest/release metadata, and distribute the exact ZIP or image digest
through a trusted artifact channel. Do not rebuild independently on every worker.
Install the corresponding trusted package manifest beneath
`/etc/truf/worker-packages` on the server before allowing that artifact to claim
protocol-2 work.
Onboard a new worker in this order:
1. Create or select its server-side user and start with active-assignment cap `1`.
2. Create a distinct device identity and issue its token. The plaintext token is
shown once; never reuse it for another device.
3. Deliver the exact reviewed Windows ZIP, native Linux package, or Linux image,
and compare its package identity with the registered server manifest.
4. Create a private three-field YAML document with the HTTPS origin, token, and
conservative parallelism such as `1`, then run `install --config`. Delete the
input YAML after installation. This writes the existing private local
configuration; lifecycle commands never need the token on their command line.
5. Run `doctor`, start the supervisor, inspect `status` and `attach`, and confirm
server `last_contact_at`.
6. Reconcile one real assignment through accepted receipt, ingestion, settlement,
and projection before increasing either the user cap or local parallelism.
7. Preserve the private state tree across restart or outage. Drain work before
token rotation, revocation, artifact replacement, or state removal.
For a routine new device, artifact delivery and token issuance are the only
installation work. Building the artifact and registering its trusted manifest are
release-management operations and should not be delegated to the device operator.
The verified copy-paste command sheets are:
- `docs/remote-worker-cheatsheet-windows-ru.md`
- `docs/remote-worker-cheatsheet-linux-ru.md`
- `docs/remote-worker-cheatsheet-docker-ru.md`
### Windows portable client
Build the pinned amd64 package in the isolated checkout when producing a release:
```powershell
New-Item -ItemType Directory -Path dist -Force | Out-Null
python -B app/worker_package_builder.py windows `
--project-root . `
--output dist/truf-worker-windows-x86_64 `
--archive dist/truf-worker-windows-x86_64.zip `
--cache build/worker-cache
```
Verify the ZIP and adjacent release JSON through the release process. On the
client, extract to a private local directory and run `prepare-worker.ps1` once to
replace inherited ACLs. Create `worker-install.yaml` in that private directory and
put only `server`, `token`, and `parallelism` in it:
```powershell
.\prepare-worker.ps1
.\truf-worker.cmd install --config .\worker-install.yaml
Remove-Item -LiteralPath .\worker-install.yaml
.\truf-worker.cmd doctor
.\truf-worker.cmd start --startup-timeout 30
.\truf-worker.cmd status
```
By default, state and data are below the current user's `LOCALAPPDATA`. The
package verifies its manifest and application files before launch and bundles
Python, Git, TruffleHog, detector policy, and locked dependencies. There is no
automatic updater. Use the same Windows account for installation and operation;
the supervisor instance and private state belong to that account.
### Native Linux portable client
Install the reviewed package beneath a root-owned path such as
`/opt/truf-worker`, then run `sudo ./prepare-worker.sh` from that directory as the
worker OS user. The preparation keeps native Git and TruffleHog immutable and
root-owned while making application code exact-private to the worker user. It
does not install a systemd unit. Use the private YAML installation and lifecycle
commands in `docs/remote-worker-cheatsheet-linux-ru.md` from that package root.
### Linux client image
Build the worker-only image for the target architecture through the reviewed
`worker` target. For x86-64:
```sh
docker build --target worker -t truf-remote-worker:linux-x86_64 .
```
Install one device into a private persistent volume. Put the server, token, and
parallelism in a mode-0600 `worker-install.yaml`, pass it on standard input to the
short-lived install container, and delete it after success:
```sh
docker volume create truf-worker-device-a-data
chmod 600 worker-install.yaml
docker run --rm -i \
--mount type=volume,source=truf-worker-device-a-data,target=/data \
--env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
truf-remote-worker:linux-x86_64 install \
--config - < worker-install.yaml
rm -- worker-install.yaml
docker run --rm \
--mount type=volume,source=truf-worker-device-a-data,target=/data \
--env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
truf-remote-worker:linux-x86_64 doctor --json
```
Run the installed supervisor in the foreground under Tini while Docker supplies
detachment and restart policy:
```sh
docker run --detach --name truf-worker-device-a --restart unless-stopped \
--read-only --cap-drop ALL --security-opt no-new-privileges --pids-limit 256 \
--tmpfs /tmp:rw,nosuid,nodev,noexec,size=128m,mode=1777 \
--mount type=volume,source=truf-worker-device-a-data,target=/data \
--env XDG_DATA_HOME=/data/client --env XDG_STATE_HOME=/data/state-base \
truf-remote-worker:linux-x86_64 run
```
The image entrypoint supplies `python -u -I -S -B` and the integrity-checking
bootstrap. It contains no PostgreSQL client, server runtime, provider
implementations, or detailed keycheck code.
Ordinary private OS storage is supported on both platforms. Application-layer
encryption of the local workspace is not required. Operators may still use host
full-disk encryption according to their own endpoint policy; this is not a TRUF
protocol requirement.
## Daily worker operation
The commands below use `truf-worker.cmd` on Windows. On Linux without Docker, use
the generated `truf-worker` launcher. For a running Docker worker, define a local
helper that executes the same package bootstrap inside its container:
```sh
workerctl() {
docker exec truf-worker-device-a /usr/local/bin/python3 -u -I -S -B \
/opt/truf-worker/app/remote_worker_bootstrap.py -- "$@"
}
```
Use these commands for normal operation:
| Command | Purpose |
| --- | --- |
| `truf-worker status` | One current human-readable supervisor and slot snapshot. |
| `truf-worker status --json` | Versioned snapshot for automation. |
| `truf-worker attach --follow-seconds 300` | Follow the verified running instance for a bounded interval; detaching does not stop it. |
| `truf-worker watch --follow-seconds 300` | Human live-status alias over the same verified attach stream. |
| `truf-worker attach --ndjson --follow-seconds 300` | Machine-readable bounded event stream. |
| `truf-worker logs --tail 200` | Read bounded rotating supervisor logs. |
| `truf-worker logs --follow --follow-seconds 300` | Follow logs for a bounded interval. |
| `truf-worker history --limit 50` | Show terminal assignment outcomes and durations. |
| `truf-worker history --reservation <id> --json` | Retrieve one assignment's terminal local record. |
| `truf-worker doctor --json` | Validate package identity, directories, configuration, retention, instance state, and runtime prerequisites. |
Run `status` first when investigating. It reports configured/occupied slots,
current phase, phase age, scan deadline, assignment time remaining, child state,
last progress age, backoff/idle reason, pending retention data, and progress-outbox
cursor. It does not invent percentage completion.
### Phases and deadlines
| Phase | Operator interpretation |
| --- | --- |
| `idle`, `claiming` | Slot is available or asking the server for work. |
| `assigned` | Immutable assignment identity and deadlines are persisted locally. |
| `waiting_permit` | Assignment owns a slot but is waiting for the shared scanner permit. This time counts against the scan-stage deadline. |
| `preparing`, `resolving`, `downloading`, `cloning` | Source-specific preparation before scanning. |
| `scanning` | Scanner process tree is active. |
| `filtering`, `cleaning`, `bundling` | Findings are converted, work is cleaned, and deterministic result bytes are staged. These phases remain inside the hard scan-stage deadline. |
| `uploading`, `awaiting_receipt` | Staged result is being transferred or waiting for authoritative server acknowledgement. |
| `backoff` | A bounded retry delay is active; inspect the reason and next-claim time. |
| `draining`, `stopped` | No new local work is starting; existing work is resolving or shutdown completed. |
The scan-stage deadline starts before `waiting_permit` and covers preparation,
provider access, scanning, filtering, cleanup, and bundle staging. Crossing it
terminates the contained runner tree and produces a normal phase-specific timeout
result while the assignment upload window remains available. The server's
assignment deadline is fixed at issue time and is never renewed by progress,
polling, restart, or upload retries. `status` therefore presents both deadlines
separately.
### Diagnostics and retained evidence
`history` identifies the terminal receipt, prebundle, timeout, stale, or recovery
outcome. `logs` shows supervisor operation; diagnostic records carry the stable
phase/category/code and optional body/log material. The private state tree stores:
```text
events/worker-events.jsonl
history/worker-history.jsonl
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.json
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.body
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.log
logs/worker.log
```
Diagnostic JSON records original/stored sizes, hash, encoding, and truncation
state. A missing or truncated body must not be described as complete. The server
admin detail view separately shows the ordered progress/receipt/ingestion/
settlement/projection timeline and canonical diagnostic snapshot. Assignment
transport outcome, scan outcome, and diagnostics are independent fields.
### Graceful stop and drain
Request local stop before maintenance. The supervisor closes new claims for this
worker, leaves authentication and the Worker API available while existing work and
pending uploads resolve, and requires a clean drain receipt:
```powershell
.\truf-worker.cmd stop --timeout 120 --json
```
For Docker, run `workerctl stop --timeout 120 --json`; the foreground supervisor
then exits and Tini returns the shutdown code to Docker. A successful shutdown
receipt reports `drained: true` and `exit_code: 0`. Do not interpret `docker stop`,
process termination, or loss of contact as assignment cancellation.
### Recovery
After an OS restart, network outage, server outage, or unclean process exit:
1. Preserve the entire private state/volume; do not remove work, bundle, event,
history, or control files.
2. Run `doctor --json` and `status --json` from the exact installed package.
3. Start the same package with `start` on Windows/Linux, or restart the same Docker
container/volume. The supervisor replays its journal and recovers assigned,
staged, upload-retry, and awaiting-receipt slots.
4. Use `attach --ndjson --follow-seconds 300` and server assignment detail to
distinguish active recovery from backoff.
5. Keep the same device identity and state until recovered slots and
accepted-but-not-ingested bundles are reconciled. Escalate only if the instance
is unverifiable, a fixed assignment deadline has passed without server recovery,
or repeated startup validation fails.
Re-uploading identical accepted bytes returns the original receipt. Conflicting
bytes or stale ownership are rejected; never delete local bytes merely to silence
that signal.
### Update and rollback
1. Request local graceful stop, reconcile unresolved/precommit work, and obtain a
successful clean drain receipt.
2. Retain the complete state tree and previous exact artifact/digest.
3. Verify the new ZIP/image and its registered server manifest. On Windows run
`prepare-worker.ps1`, then `doctor`; for Docker recreate only the container and
mount the same volume.
4. Start at parallelism `1`, confirm package identity/contact and one complete
accepted-ingested-projected assignment, then restore the intended cap.
5. If validation fails, stop and return to the previous exact artifact with the
same state. Additive server progress/diagnostic records need no rollback.
Do not replace binaries beneath a running supervisor or switch packages while an
assignment runner is active.
### Device removal
1. Stop every device that uses the identity locally; leave authentication valid
while all unresolved assignments and precommit bundles reach zero.
2. Stop gracefully and retain the shutdown receipt and terminal history.
3. Revoke the device and disable its user only if that user is not shared by an
active device.
4. Confirm the old token no longer authenticates and no authoritative recovery
remains.
5. Remove the container/package. Remove its private volume/state only after the
server reconciliation evidence is retained and no rollback requires it.
## Server tuning
Make tuning changes in the reviewed private runtime configuration, not on the
client. Keep the raw worker service on loopback/private addressing and expose it
only through certificate-validating Caddy HTTPS.
| Control | Meaning |
| --- | --- |
| Client `--parallelism N` | Maximum locally occupied slots, one claim per free slot. |
| Typed admin assignment cap | Atomic positive active-assignment cap for one user across all devices; lowering it does not cancel existing assignments. |
| `assignment_ttl_seconds` | Fixed server-clock lifetime covering download, scan, and upload; default `86400`. API contact and restart do not renew it. |
| `bundle_body_timeout_seconds` | Upload body deadline; default `1800`. The assignment lifetime must exceed the largest configured source scan timeout plus this value plus 60 seconds. |
| `json_body_timeout_seconds` / `body_idle_timeout_seconds` | Request and idle transport bounds; defaults `60` and `30`. They do not renew ownership. |
| `reaper_interval_seconds` / `reaper_batch_size` | Expired-assignment recovery cadence and bounded batch; defaults `60` and `1000`. |
| `limit_concurrency` | Worker API request concurrency bound, not a replacement for user quotas; default `64`. |
Shortening the fixed lifetime can reject a valid long scan or upload and allow a
second physical execution after recovery. Lengthening it holds user quota,
target ownership, and dependent leases longer after a lost client. Tune it from
observed end-to-end duration plus upload headroom, not from HTTP polling cadence.
## Typed administration
Use only the authenticated random-prefix admin page configured by the production
edge. It provides CSRF/Origin-checked typed operations to create/enable/disable
users, set assignment caps, issue/rotate/revoke/unrevoke device tokens, and
requeue selected deferred queue IDs. It deliberately provides no shell or generic
supervisor command.
An issued or rotated token is shown once. Rotation replaces the stored digest, so
the old token stops authenticating; update that device without copying the token
to other devices. Rotation does not clear an existing revoked state. Revocation
blocks further API authentication but does not invent cancellation for assigned
work; plan for outstanding work to be completed before revocation or recovered at
its fixed expiry. Use local graceful stop to close claims before planned device
maintenance or removal while preserving authentication for pending work.
For an admin IP ban, use SSH and fail2ban first so fail2ban and Caddy agree:
```sh
sudo fail2ban-client set truf-admin-auth unbanip 203.0.113.10
sudo /usr/local/sbin/truf-caddy-admin-denylist status
sudo /usr/local/sbin/truf-caddy-admin-denylist expire
```
If fail2ban is unavailable, use the explicit admin-only updater:
```sh
sudo /usr/local/sbin/truf-caddy-admin-denylist unban 203.0.113.10
```
These commands change only the admin-route matcher. Never replace them with a
global port 443 firewall unban/ban. The full damaged-snippet recovery procedure
is in `deploy/edge/README.md`.
## Accepted custody and expiry
`accepted` means the server validated the canonical v2 `.trb`, durably published
it, and persisted ready/recovery state before returning a receipt. It does not
mean ingestion, projection, candidate handling, or detailed keycheck has
finished. Track `accepted` and `ingested` separately in the typed admin view.
The client keeps assignment identity and pending bytes until authoritative
acknowledgement. Retrying identical accepted bytes returns the original receipt,
including after ingestion, spool cleanup, restart, or the former deadline;
conflicting bytes are rejected. An interrupted, invalid, expired-before-first-
acceptance, or stale upload is not a successful scan.
An unfinished assignment expires at the original server-set deadline, normally
24 hours after issue. The periodic recovery pass reconciles quota, credits,
target ownership, and dependent plan/blob leases through existing retry policy.
Already accepted ready bundles are not requeued as unfinished. A crashed client
can therefore delay work for about a day, while a partitioned client can continue
physical scanning after the server has expired and reissued the target. Ownership
fencing guarantees one authoritative acceptance, not exactly-once physical work.
## Staged rollout and rollback
1. Obtain separate review for production schema/deployment changes. Drain protocol-1 work before replacing packages. Keep worker API and admin disabled while configuring exact protocol-2 capability profiles, task-specific auth entries, trusted package manifests, edge origin/marker, and per-user caps.
2. Pass the main Docker, packaged Windows/Linux, and edge gates with their pinned artifacts. Do not infer readiness from unit mocks or one platform.
3. Enable a small canary: one user, one device, cap `1`, and client `N=1`. Confirm assignment, acceptance, ingestion, projection, detailed server keycheck, expiry, and safe logs before increasing either cap or device count.
4. Expand caps and clients in stages while comparing unfinished/completed/failed/expired counts, last authenticated contact, accepted-versus-ingested state, durations, capacity, and stale/duplicate events. Contact age alone is not a liveness failure while a client holds long-running work and has not yet returned to claim polling.
5. To drain, request local graceful stop on every affected device. Leave authentication and the Worker API available so pending uploads and terminal reports can resolve. Wait for unfinished assignments to complete or pass through fixed-expiry recovery, and separately reconcile accepted bundles awaiting ingestion.
6. After the drain is authoritative, disable remote admission and return scheduling to local-only execution. Then revoke unused device tokens if required. Retain reservation metadata, accepted receipts, bundles/recovery state, and schema until reviewed reconciliation is complete.
7. Do not drop worker metadata, clear spool state, rotate/revoke tokens before a drain, switch server binaries underneath unfinished work, or treat a service stop as cancellation. The local execution path remains available and must not depend on a remote worker.
+98
View File
@@ -0,0 +1,98 @@
# TRUF worker: установка и работа
Worker получает задания от сервера, скачивает публичные targets и отправляет
только результат сканирования. Для каждого компьютера или Docker volume нужен
отдельный device token.
Сервер: `https://pregnant.horsecock.store`
## Что получить у администратора
1. Проверенный artifact для своей платформы и соседний файл с SHA-256.
2. Одноразово показанный device token. Не отправляйте его в чат, лог или снимок
экрана.
3. Подтверждение, что server-side User и Device включены и artifact зарегистрирован.
Администратор создаёт отдельные User и Device на странице `Workers / Dispatch`,
выдаёт token и назначает положительный assignment cap. Один token нельзя
использовать на нескольких устройствах.
## Приватный install YAML
Token не нужно передавать в аргументах процесса. Создайте локальный
`worker-install.yaml` в приватной папке через текстовый редактор:
```yaml
server: https://pregnant.horsecock.store
token: PASTE_DEVICE_TOKEN_HERE
parallelism: 1
```
`parallelism` задаёт число локальных занятых slots и должен быть от 1 до 128.
После `install` worker сохраняет настройки в своём приватном
`worker.config.json`; исходный YAML нужно удалить. При обычных `start`, `stop`,
`status`, `attach` и `watch` token больше не вводится.
## Шпаргалки
- [Windows](remote-worker-cheatsheet-windows-ru.md)
- [Linux без Docker](remote-worker-cheatsheet-linux-ru.md)
- [Docker Engine / Docker Desktop](remote-worker-cheatsheet-docker-ru.md)
Каждая шпаргалка начинается из корня распакованного artifact или Compose bundle
и содержит проверенные команды установки и lifecycle.
## Что делают lifecycle-команды
| Команда | Результат |
| --- | --- |
| `start` | Запускает установленный worker в фоне; повторный запуск не создаёт второй instance. |
| `stop --timeout 120` | Локально закрывает новые claims, завершает текущую работу и требует clean drain receipt. |
| `status` | Показывает instance, slots, текущие phases, deadlines и retained state. |
| `attach` | Подключает live status/event view; `q` или `Ctrl-C` только отсоединяет. |
| `watch --follow-seconds 300` | Запускает bounded live status view и затем отсоединяется. |
| `logs --follow --follow-seconds 300` | Показывает bounded live event/log stream. |
| `doctor --json` | Проверяет package, config, пути, native tools, TLS endpoint и singleton state. |
`stop` управляет только выбранным локальным worker. Менять server assignment cap
для обычного stop, restart или обновления не требуется. Не завершайте процесс и
не удаляйте state, пока clean drain receipt не подтверждён.
## Capacity и backpressure
- Один result bundle имеет hard limit 64 MiB.
- При выдаче remote assignment server резервирует baseline 2 MiB для bundle и
2 MiB для projection; это не новый hard limit.
- Валидный результат больше baseline атомарно расширяет reservation по фактическому
размеру. При временной нехватке capacity worker повторяет upload позже.
- Server допускает не более 50 unresolved remote assignments глобально и
одновременно применяет положительный per-user cap. Фактический предел равен
меньшему из доступной capacity, global limit и user cap.
- Client `parallelism` ограничивает только локальные slots и не повышает server cap.
## Частые состояния
- `idle` или `claiming`: slot свободен или запрашивает задание.
- `downloading`, `cloning`, `scanning`: выполняется задание.
- `uploading`, `awaiting_receipt`: результат отправляется или ждёт подтверждения.
- `backoff`: временная ошибка; причину и следующую попытку показывает `status`.
- `draining`: новые локальные claims закрыты, текущая работа завершается.
- `stopped`: clean shutdown завершён.
При проблеме сохраните вывод `doctor --json`, `status --json` и
`logs --tail 200`. Никогда не прикладывайте device token, install YAML или
приватный `worker.config.json`.
## Безопасное обновление
1. Выполните локальный `stop --timeout 120 --json` и получите `drained: true`,
`exit_code: 0`.
2. Сохраните предыдущий точный artifact/image и весь state/volume.
3. Проверьте SHA-256 и package identity новой версии.
4. Windows/Linux: распакуйте новую версию в отдельную папку. Docker: загрузите
новый image и пересоздайте только container с прежним volume.
5. Выполните `doctor`, `start`, `status` и bounded `watch`.
Не удаляйте локальные `state`, `work`, `bundles`, `events`, `history` или Docker
volume при ошибке и не используйте `docker compose down --volumes`. Они нужны
для безопасного продолжения и authoritative receipt recovery.
+78
View File
@@ -0,0 +1,78 @@
# Current State
Updated: 2026-09-26
Workspace: `D:\truf-workers`
## STATUS: Primary Objective Complete
Extended production validation with exactly one native Windows worker slot and
one WSL/Docker worker slot is complete. Run
`e6ac8aec-ec20-4ba4-a924-abe5ee95d82c` exercised discovery, assignment, worker
execution, progress, diagnostics, bundle upload and ingestion, projection,
capacity release, and scheduled keycheck.
The authoritative public-safe result is
`docs/extended-live-validation-2026-09-26.md`. The machine-readable aggregate
manifest is
`build/extended-live-validation/runs/e6ac8aec-ec20-4ba4-a924-abe5ee95d82c/verification-manifest.json`.
Conclusion: `pass_with_documented_deviations`.
## STATUS: Terminal Validation Cut
- Controller revision 143 was sealed with dispatch and discovery paused and the
validation drain complete.
- All 314 reservations were terminal: 309 acknowledged and 5 refunded.
- All 309 accepted bundles were ingested, projected, settled, and released.
- No unresolved reservation, live assignment, pre-commit bundle, capacity use,
publication outbox item, quarantine row, waiting lock, or blocker remained.
- All 348 projection jobs completed and released in one attempt without error.
- The captured 39-candidate keycheck cohort completed, linked, projected, and
released; the global queue had no pending, leased, or deferred candidate.
- Server and worker hash chains and all referenced objects verified.
## STATUS: Production Restored
Production restore completed through compare-and-swap revisions 143 to 146:
1. Cancel the validation drain: 143 to 144.
2. Resume discovery: 144 to 145.
3. Resume dispatch: 145 to 146.
Last verified state:
- Controller revision 146, actor `validation-final-restore`.
- Dispatch open, discovery open, drain normal.
- Exactly one live assignment on Windows and one on WSL/Docker.
- Runtime healthy with zero validation-run restarts.
- Edge running with zero validation-run restarts.
- Core services and discovery producers healthy.
This is a recorded final snapshot, not a substitute for a fresh query in a
future operational session.
## STATUS: Documented Deviation
The authenticated `recheck all` operation was accidentally used instead of
rechecking only the two pending candidates, causing 35 broad provider checks
contrary to `KEYCHECK-001`. All 35 were processed, none was skipped, the command
returned success, and all effects settled. Do not repeat this probe.
## STATUS: Remaining Defects
- Windows GitLab filename-too-long checkout recovery is local and not deployed.
- Invalid API-key classification is fixed locally and not deployed.
- A roughly 20-second WSL clock-domain monotonic failure remains open.
- Generic `WorkerContractError` diagnostics lose structured field detail.
- Monitor aggregate queries produced six recovered statement timeouts under
load.
Focused regression tests for the local checkout, long-path, and credential
changes passed: 145 tests in 155.26 seconds.
## NEXT: No Validation Work Pending
Do not repeat the completed live validation or broad keycheck. Future work, if
requested, is limited to reviewing and deploying the documented local fixes,
investigating the remaining defects, and archiving the completed OpenSpec
change. Do not archive OpenSpec without explicit user instruction.
+94
View File
@@ -0,0 +1,94 @@
# Durable Decisions
Updated: 2026-09-25
## DECISION: RAW-001 - Raw Analysis Is Mandatory
For analysis and live testing, it is forbidden to design or add token hiders,
maskers, redactors, secret-scrubbers, post-hoc sanitizers, lossy summaries, or
similar mechanisms that conceal the values being validated.
- Inspect and retain the original raw payloads, targets, findings, credentials,
provider responses, diagnostics, logs, database rows, and worker records needed
to prove correctness.
- The monitor must not replace, hash-only, truncate, partially reveal, or rewrite
a value before the analysis has consumed and preserved the raw record.
- If a UI already has a reveal control, use the revealed raw value for the test;
do not build another concealment layer.
- Raw evidence must remain in private local/protected storage and must not be
published to a public report, issue, commit, or chat excerpt. This storage rule
is not permission to hide data from the analysis itself.
- A later public/operator report may reference counts and hashes, but it must be
derived only after raw correctness has been checked.
This decision supersedes any old handoff wording that instructed the testing
session to analyze only sanitized aggregates. Historical sanitized reports remain
valid as reports; they are not sufficient evidence for the new run.
## DECISION: ENV-001 - Production Target
- Use the configured SSH server named `sec` only.
- Never call, connect to, or mutate the configured server named `prod`.
- Workspace is `D:\truf-workers`.
- Do not run the inherited native runtime launchers in this source-only workspace.
- Do not mount or mutate unrelated `D:\truf` runtime data.
## DECISION: RUN-001 - Dual Worker Bounds
- Native Windows: exactly one worker slot/thread.
- Docker under WSL: exactly one worker slot/thread.
- Expected maximum combined worker concurrency: two.
- Do not increase caps or parallelism to accelerate the observation window.
- Waits of up to ten minutes are allowed; several hours of observation are
explicitly authorized.
## DECISION: KEYCHECK-001 - Scheduled Validation
- Keycheck may be enabled for the run every 30 minutes (`1800` seconds).
- Verify candidate leases, provider execution, append-only results, current-state
selection, projection jobs/appends, and capacity release from raw records.
- Provider probes may have real external effects or cost; do not silently widen
service args or recheck policy beyond the active configuration.
## DECISION: CONFIG-001 - Stale Candidate Must Not Be Applied
Do not apply the stale config candidate with SHA-256
`c0966cac4f7f0610a813fa8732e91953f2e3e838ad880f91fd1a9437096925c7`.
It was based on an older active hash and would reduce
`global.keycheck_queue_max_items` from the retained `8192` to `4096`.
Always fetch the current active config identity and use the authenticated
fresh-hash/CAS workflow for any 1800-second keycheck edit.
## DECISION: ARCH-001 - Worker Authority
- The worker is final authority for real provider access.
- Server planning may bind immutable Git/Docker identity but must not add
per-target preflight/provider-access proof machinery.
- Do not add credential sandboxes, environment rewriting, durable access proofs,
or security-specific infrastructure without a separate explicit user decision
and OpenSpec requirement.
- Prefer bounded direct error classification. Authentication/access/not-found is
permanent when target-scoped; rate limits, network failures, and provider 5xx
are retryable.
The complete engineering decision is in `AGENTS.md`.
## DECISION: CHANGE-001 - Repository and OpenSpec
- This repository has no baseline commit; the full tree appears untracked.
Never use Git to revert or clean files and never treat `git diff` as complete.
- Preserve unrelated files and evidence directories.
- `add-worker-operator-experience` is complete but must not be archived without
an explicit request.
- Do not repeat the already completed 295-assignment production validation unless
a fresh verification proves its retained evidence invalid.
## DECISION: CONTEXT-001 - Session Continuity
- Use these files for continuation instead of recursive DCP summaries.
- Do not proactively invoke conversation compression in the new session.
- Batch searches and process large evidence in tools; avoid injecting raw
multi-megabyte files into the conversation context.
- The prohibition on injecting large evidence into chat does not permit masking
or omitting it from the private analysis artifact.
+140
View File
@@ -0,0 +1,140 @@
# Evidence Index
Updated: 2026-09-26
This file maps facts to their existing source. Do not duplicate the underlying
evidence in handoff prose.
## STATUS: Authoritative Reports
- `docs/extended-live-validation-2026-09-26.md`
- Final public-safe report for the two-slot extended live run, including
settlement, keycheck, evidence integrity, deviation, restore, and open
defects.
- `build/extended-live-validation/runs/e6ac8aec-ec20-4ba4-a924-abe5ee95d82c/verification-manifest.json`
- Machine-readable derived aggregates, classifications, evidence hashes,
terminal state, post-restore state, and test result.
- `docs/worker-operator-experience-live-trace-2026-09-25.md`
- 295-assignment Windows/Linux cohort, aggregate outcomes, recovery, admin UI,
pipeline state, tests, and evidence hashes.
- `docs/worker-operator-experience-validation-2026-09-24.md`
- Earlier bounded production/package acceptance and operator validation.
- `docs/remote-worker-operations.md`
- Canonical worker install, lifecycle, diagnostics, drain, update, and removal.
- `WORKER_OPERATOR_EXPERIENCE_HANDOFF.md`
- Historical implementation and reproducible artifact detail. Its stop point
predates the final live trace and defect-fix rollout.
- `openspec/changes/add-worker-operator-experience/tasks.md`
- Current completion authority: all 27 tasks checked.
## RAW: Private Live Evidence
- `build/extended-live-validation/runs/e6ac8aec-ec20-4ba4-a924-abe5ee95d82c`
- Local run root with compact server artifacts and complete worker evidence.
- Server NDJSON SHA-256:
`a6beb1864c540fc5f22b2b647730a39e45dfddb10cf5ceff5e0181eb1a8f0bd8`.
- Worker NDJSON SHA-256:
`03a04a0da5a2db6bd02f2aaed41c191554b135ec287fbc1e2d9998d77da59898`.
- Server run SHA-256:
`5ae58df5f67a8d2a8d4e84d73262a6e7009e03dcbcc9ea176537b384d4e693c3`.
- Private production server evidence root:
`/var/lib/docker/volumes/truf-remote-server-data/_data/extended-live-validation/e6ac8aec-ec20-4ba4-a924-abe5ee95d82c`
- Complete server evidence, approximately 1.2 GiB. Analyze in place and do not
copy raw values into public documents.
- `build/live-trace-20260925/raw-evidence-final.json`
- Complete server cohort evidence. Historical recorded SHA-256:
`4071e38a1dec540663bc6febd9538ff54bedafa1627250933045beb8f7d09ec5`.
- `build/live-trace-20260925/raw-evidence-expanded.json`
- Expanded unrestricted evidence used for defect diagnosis.
- `build/live-trace-20260925/monitor-final.ndjson`
- Time-series server monitor. Historical SHA-256:
`e07507b0b8f0f86d1a1c7aade7186297c372b1ff177067bea19dfd219f5027a9`.
- `build/live-trace-20260925/windows-localappdata/TRUF/RemoteWorker`
- Final retained Windows worker state, events, history, logs, and inactive
abandoned roots.
- `build/live-trace-20260925/linux-worker-state-final.tar.gz`
- Final retained Linux worker state. Historical SHA-256:
`b64e81c90b232f46b400a63ed08f5660f46e34fedb1f67b64afba079d8d36364`.
- `build/live-trace-20260925/dockerhub-discovery-db-raw.json`
- `build/live-trace-20260925/dockerhub-unmasked-cycle-failure.json`
- `build/live-trace-20260925/dockerhub-discovery.log`
- Raw DockerHub retry defect evidence.
These files may contain sensitive raw values. Analyze them in place; do not copy
their contents into a public document.
## STATUS: Defects and Current Treatment
- Extended-run open items:
- Windows GitLab filename-too-long checkout recovery is local and not
deployed.
- Invalid API-key classification is fixed locally and not deployed.
- A roughly 20-second WSL clock-domain monotonic failure remains open.
- Generic `WorkerContractError` diagnostics lose structured field detail.
- Six monitor aggregate-query statement timeouts recovered during the run.
- `KEYCHECK-001` deviation:
- An unintended authenticated broad recheck processed 35 provider checks
instead of only the two pending candidates. All effects settled; do not
repeat this probe.
- `docs/defect-windows-scan-timestamps-utc-2026-09-25.md`
- Still open. Local `app/scanner.py` continues to create naive timestamps.
- `docs/defect-dockerhub-discovery-retry-null-type-2026-09-25.md`
- Original report says local-only. Current source has the PostgreSQL
`CAST(? AS TEXT)` correction and the defect-fix rollout included it. Reverify
live retry coalescing during the extended run.
- `docs/defect-terminal-status-scan-deadline-readback-2026-09-25.md`
- Original report says open. Current source preserves durable deadlines with
`result.setdefault('deadlines', observability['deadlines'])`; included in the
defect-fix rollout. Reverify terminal readback.
- `docs/defect-worker-network-oserror-mislabeled-local-io-2026-09-25.md`
- Original report says open. Current source introduces `WorkerNetworkError`
before broad `OSError` classification; included in the defect-fix worker
artifacts. Reverify under a real transport failure.
## STATUS: Defect-Fix Rollout Artifacts
- `build/runtime-defect-fixes-v1/deploy-runtime.sh`
- Atomic runtime/manifests cutover and rollback logic.
- `build/runtime-defect-fixes-v1/windows-b.zip`
- `build/runtime-defect-fixes-v1/windows-b.zip.json`
- `build/runtime-defect-fixes-v1/linux-worker-package.json`
- `build/runtime-defect-fixes-v1/verify-production-worker.py`
- `build/runtime-defect-fixes-v1/restore-production-worker.py`
Recorded identities:
- Deployed runtime image:
`sha256:7d84fdf57a1cb9e6d38a571fbd3566b7549f1cda04ae02c4864cac70f74f2aaa`.
- Current WSL worker image:
`sha256:491b3a2343571072209a7f92e83399fe206006dcccf247a0c551c50fc9f35e30`.
- Rollback tag: `truf-local:runtime-pre-defect-fixes-v1`.
- New registered manifest file hashes from the deploy script:
- Linux: `b9d3594e4846a21ca12de5fc6973c04d9eea6f61fd0ecda83875426aa48c7b4b`.
- Windows: `4cc97d17c34f7d89150927719f131f426f1068ab99af7f3f2f156d8642ae539e`.
## STATUS: Historical Accepted Artifacts
The pre-defect-fix live trace used:
- Windows package manifest:
`78a962b2bd3fa411413c79e9a8ffb021608a08ff020b1ad851f4505ea634b2b6`.
- Linux package identity:
`45588f2cf406b41b239cfa3b8a9dc83fe84b587229bc997b2729016e1f0dde42`.
- Linux image:
`sha256:3a088f5743121d823aae132234a29730a84339cecbfda5fc601e8e942f9948c3`.
These remain valid historical evidence but are not the preferred artifacts for
the new defect-fix observation run.
## VERIFY: Fast Orientation Commands
Run from `D:\truf-workers`:
```powershell
openspec list --json
git status --short --branch
python -B -m pytest tests/test_worker_api.py tests/test_worker_api_runtime.py tests/test_worker_assignment.py tests/test_worker_assignment_runner.py tests/test_worker_cli.py tests/test_worker_contracts.py tests/test_worker_local_state.py tests/test_worker_observability_db.py tests/test_worker_package.py tests/test_worker_runner_handoff_linux.py tests/test_worker_supervisor.py tests/test_remote_worker_db.py tests/test_scan_execution.py tests/test_admin_api.py -q
```
Do not use an unrestricted repository-wide pytest run as the release gate. Do
not use Git clean/reset/checkout in this uncommitted snapshot.
+35
View File
@@ -0,0 +1,35 @@
# Session Handoff Index
Updated: 2026-09-26
This directory is the authoritative entry point for a new session working on
the live remote-worker validation. Read only these files first:
1. `CURRENT_STATE.md` - where work stopped and the exact next actions.
2. `DECISIONS.md` - binding user decisions and operational constraints.
3. `EVIDENCE_INDEX.md` - existing reports, raw captures, artifacts, and hashes.
The older root `WORKER_OPERATOR_EXPERIENCE_HANDOFF.md` is historical background.
It remains useful for implementation detail and artifact provenance, but its
"Immediate next actions" section is obsolete.
## Knowledge Layout
- A current fact has exactly one owner: `CURRENT_STATE.md`.
- A durable rule has exactly one owner: `DECISIONS.md`.
- Evidence is not copied into handoff prose; `EVIDENCE_INDEX.md` points to it.
- Dated reports are immutable history. Record later corrections here instead of
rewriting the original report.
- Replace stale current-state statements rather than appending contradictory
status paragraphs.
Useful grep tags are `STATUS:`, `NEXT:`, `BLOCKER:`, `DECISION:`, `VERIFY:`, and
`RAW:`.
## New Session Start
Use this prompt:
> Read `docs/session-handoff/README.md` and its three linked files. Continue the
> `NEXT:` work in `CURRENT_STATE.md` autonomously. Do not repeat completed live
> validation. Follow every `DECISION:` literally, especially RAW-001 and ENV-001.
@@ -0,0 +1,336 @@
# Worker Operator Experience Live Trace - 2026-09-25
## Result
The accepted Windows package and Linux image completed a fresh, concurrent live
cohort against the `sec` runtime. All 295 assignments were acknowledged,
ingested, projected, and settled. There were no unresolved assignments,
pre-commit bundles, expiries, pre-bundle failures, quarantine rows, append
failures, or publication-outbox rows at the final cut.
The worker transport and server pipeline acceptance result is **pass**. Four
product defects and two operational warnings were found. The defects did not
invalidate the one-authoritative-acceptance, ingestion, projection, recovery, or
shutdown guarantees demonstrated by this cohort, but they remain release inputs
and are listed below.
This report is sanitized. It intentionally excludes authentication values,
private routes, worker command lines, raw targets, raw findings, secrets, and
runtime configuration bodies. The protected evidence files referenced below
contain sensitive material and must not be published.
## Validated artifacts
The run used the previously accepted reproducible artifacts without modifying
their product code:
| Platform | Accepted identity |
| --- | --- |
| Windows x86-64 | package manifest `78a962b2bd3fa411413c79e9a8ffb021608a08ff020b1ad851f4505ea634b2b6` |
| Linux x86-64 | package identity `45588f2cf406b41b239cfa3b8a9dc83fe84b587229bc997b2729016e1f0dde42` |
| Linux image | `sha256:3a088f5743121d823aae132234a29730a84339cecbfda5fc601e8e942f9948c3` |
Both trusted manifests remained registered on the server. Each validation worker
ran one slot with local parallelism `1`. The measured combined concurrency
reached exactly `2`, proving concurrent accepted Windows and Linux execution
without increasing either worker's local parallelism.
## Final cohort
### Platform and source distribution
| Worker | DockerHub | GitLab | Hugging Face | Total |
| --- | ---: | ---: | ---: | ---: |
| Windows | 58 | 57 | 51 | 166 |
| Linux | 50 | 30 | 49 | 129 |
| Total | 108 | 87 | 100 | 295 |
Every history row ended with `bundle_accepted` and mapped to exactly one server
reservation and receipt. The Windows and Linux reservation sets were disjoint and
their union exactly matched the 295-row server cohort.
### Scan and queue outcomes
| Scan outcome | Count |
| --- | ---: |
| clean | 143 |
| degraded | 73 |
| error | 75 |
| found | 4 |
| Queue disposition | Count |
| --- | ---: |
| done | 220 |
| deferred | 74 |
| failed | 1 |
These are scanner/queue outcomes, not transport failures. All corresponding
result bundles were accepted and projected. In particular, timeout and scanner
error results remained normal uploadable terminal results.
### Persisted detail
The stable capture contains:
- 295 admission intents, reservations, bundles, target scans, compatibility rows,
and completed scan projection jobs;
- 2,905 ordered progress events;
- 75 structured diagnostics;
- 129 normalized scan errors;
- 4 normalized findings, 4 stable UID mappings, and 4 bounded compatibility
payloads;
- 4 pending keycheck candidates linked to 2 normalized credentials;
- 374 appended projection records across 3 streams;
- 956 pipeline artifacts, all deleted by the final snapshot; and
- 19 successful typed runtime operations with 38 chained audit events.
The diagnostic aggregate was:
| Worker | Scanner result errors | Stage timeouts | Total diagnostics |
| --- | ---: | ---: | ---: |
| Windows | 55 | 8 | 63 |
| Linux | 1 | 11 | 12 |
Of the Windows scanner-result errors, 54 were retryable and one was
non-retryable. The Linux scanner-result error was retryable. Complete safe
diagnostic envelopes, exception identities, bounded process-log representations,
occurrence times, receipt authority, and scan links were retained and checked.
## End-to-end integrity checks
`build/operator-experience-validation/analyze-live-trace-final-20260925.py`
executed 65,675 checks with zero failures. It verified, row by row:
- admission, reservation, queue, bundle, scan, compatibility, and projection
foreign-key relationships;
- receipt, payload, bundle, event, device, and deadline identities;
- canonical SHA-256 values for execution snapshots, progress events,
diagnostics, compatibility metadata, plans, projection events, and finding
payloads;
- exact bounded reconstruction of the four compatibility findings, including
their original numeric detector identity and explicit null mapped fields;
- every normalized error, finding, keycheck candidate, credential reference,
projection append, capacity release, and deleted artifact;
- all 19 operation-to-audit pairs; and
- the complete 38-event audit parent/hash chain.
`build/operator-experience-validation/analyze-worker-states-final-20260925.py`
executed a further 28,686 checks with zero failures. It parsed every retained
worker JSON/JSONL record and verified:
- 166 Windows and 129 Linux history rows against the server receipts;
- 2,338 contiguous Windows events and 1,672 contiguous Linux events;
- receipt payload, bundle, event, acceptance-time, and reservation identities;
- clean local shutdown with `drained=true`, `exit_code=0`, and empty progress
outboxes; and
- no active work root in either final worker snapshot.
The two analyzers therefore executed 94,361 deterministic checks without a
failure.
## Recovery and shutdown
The validation exercised durable recovery rather than only clean executions:
- transient server `502` responses were retained in both local worker logs and
recovered without duplicate authoritative acceptance;
- one Linux assignment survived a worker stop in the persistent volume, resumed
after restart, produced one accepted result, and released all capacity;
- a transient assignment-status network failure retried the same durable
assignment without rescanning or data loss; and
- both workers then drained and stopped cleanly.
The final local states were:
| Worker | State | Drained | Exit | Pending outbox |
| --- | --- | --- | ---: | ---: |
| Windows | stopped | true | 0 | 0 |
| Linux container | exited | true | 0 | 0 |
Retained `work/abandoned` roots are inactive evidence governed by normal worker
retention. They are not active assignments.
## Worker API validation
Authenticated worker API validation covered:
- device identity, package-manifest trust, and assignment-cap enforcement;
- claim, reservation replay, assignment status, and immutable execution snapshot;
- monotonic progress submission and latest-progress readback;
- bounded diagnostic body and process-log payloads;
- durable bundle upload, idempotent receipt replay, and accepted resolution;
- local restart recovery and terminal history;
- server ingestion, normalized scan authority, compatibility reconstruction, and
projection completion; and
- terminal capacity release and zero unresolved work.
The API preserved receipt, payload, scan-event, diagnostic, and reservation
identities throughout the cohort. A separate terminal readback defect affecting
only `scan_deadline_at` is documented below.
## Admin UI and operator workflow
The authenticated admin UI was exercised through a browser across the complete
operator surface:
- Workers / Dispatch: control state, users, devices, assignment caps, enable,
disable, revoke, unrevoke, and bounded worker detail;
- Overview: runtime health, producer lifecycle, pipeline workers, leases, queue
counts, capacity, controls, recent operations, and duration groups;
- Search: bounded assignment, scan, finding, error, diagnostic, and progress
lookup with safe empty and populated states;
- Supervisor: source start, restart, stop, lifecycle, and safe error rendering;
- Logs: bounded source/component/level/time filters and empty-result handling;
- Config and Secrets: active identities, stale-candidate warning, redacted
projections, validation, and non-secret operation results;
- Files: bounded listing, file identity/hash verification, and safe download
behavior;
- Operations: accepted/running/succeeded projections and filters; and
- Audit: accepted/succeeded event pairs, pagination, actor/action filters, and
parent/hash continuity.
The final UI snapshot showed:
- Supervisor `ACTIVE` and PostgreSQL `READY`;
- result ingester, JSONL projector, janitor, and worker API all running;
- ingester and projector leases ready;
- control revision `126`, discovery and dispatch open, drain state normal;
- zero active assignments and zero pre-commit bundles; and
- zero quarantined queue rows.
The current DockerHub producer safe state remained `runtime_error`; it is the
known retry-coalescing defect plus unavailable credential pool described below,
not an unclassified new failure.
## Server and pipeline final state
The final server snapshot at 2026-09-25 14:59 UTC confirmed:
- 295 issued, 295 accepted, 295 ingested, 295 projected, and 295 settled;
- 0 unresolved, expired, pre-bundle-failed, pre-commit, quarantine, and drain
blockers;
- bundle, projection, and quarantine capacity at zero;
- keycheck capacity at 137 items / 27,262,866 bytes, representing real pending
work rather than leaked assignment or projection capacity;
- runtime container healthy with zero restarts and no OOM;
- edge container running with zero restarts and no OOM; and
- dashboard health endpoint returning `200 ok`.
## Defects found
### 1. Windows scanner timestamps lose their UTC offset
Status: open in the accepted Windows package.
Naive local scanner timestamps are relabeled as UTC, producing approximately a
three-hour future displacement on the validation host. Progress transport and
monotonic durations remain correct, but diagnostic ordering, time filters, and
scan start/end instants are wrong.
Detailed evidence and correction:
`docs/defect-windows-scan-timestamps-utc-2026-09-25.md`.
### 2. DockerHub retry coalescing fails with a null retry time
Status: open in the live runtime; fixed locally but not deployed.
PostgreSQL cannot infer the type of a nullable retry placeholder while
coalescing an existing row, raising SQLSTATE `42P18`. The managed producer masks
that exception as a generic delegation failure. A minimal local correction casts
the placeholder to text, and regression coverage now exercises same-row null-time
coalescing.
Detailed evidence and local fix:
`docs/defect-dockerhub-discovery-retry-null-type-2026-09-25.md`.
### 3. Terminal status can lose the durable scan deadline
Status: open in the live runtime.
A later progress event with a null scan deadline can replace the durable
receipt's concrete immutable deadline in authenticated terminal status readback.
The persisted receipt and all other identities remain correct.
Detailed evidence:
`docs/defect-terminal-status-scan-deadline-readback-2026-09-25.md`.
### 4. Network `OSError` is mislabeled as local I/O
Status: open in the accepted workers.
The broad safe-summary branch classifies socket/transport `OSError` as
`local I/O operation failed`. Recovery worked and no data was lost, but the
operator message incorrectly points toward local storage.
Detailed evidence:
`docs/defect-worker-network-oserror-mislabeled-local-io-2026-09-25.md`.
## Operational warnings
- The server root filesystem was approximately 90% used with roughly 1 GiB free.
Containers remained healthy, but capacity should be reclaimed or expanded.
- The DockerHub credential pool was independently unavailable: ten credentials
were invalid and the remaining entry was rate-limited. Deploying the SQL fix
preserves retries correctly but cannot make an unavailable credential pool
healthy.
## Tests and specification gates
The retained gates are:
- worker/operator focused matrix: `355 passed, 3 skipped`;
- admin/runtime focused matrix: `124 passed, 2 warnings`;
- DockerHub incremental discovery matrix: `27 passed`;
- SQLite retry lifecycle regression: `1 passed`;
- corrected PostgreSQL statement verified transactionally against the live schema
and rolled back; and
- OpenSpec strict validation passed with all 27 implementation tasks complete.
The disposable PostgreSQL integration test remains skipped locally because no
disposable DSN was configured. The live transactional SQL verification did not
persist a change.
## Evidence manifest
| Artifact | Bytes | SHA-256 |
| --- | ---: | --- |
| `build/live-trace-20260925/raw-evidence-final.json` | 10,030,750 | `4071e38a1dec540663bc6febd9538ff54bedafa1627250933045beb8f7d09ec5` |
| `build/live-trace-20260925/monitor-final.ndjson` | 1,260,087 | `e07507b0b8f0f86d1a1c7aade7186297c372b1ff177067bea19dfd219f5027a9` |
| `build/live-trace-20260925/summary-final.json` | 3,669 | `19e5ec40cf354c6095a478b68a69f8e51fb3b8ea1c6f278c1d9c90bdaab9ebe2` |
| `build/live-trace-20260925/final-analysis-summary.json` | 946 | `2d17605ba7b730ad78a48bcc70e40e78f90637e3fd59004ebf5eb4a08241ba42` |
| `build/live-trace-20260925/final-worker-state-analysis-summary.json` | 1,376 | `b0ab428eb08587f13f59a1763f835125216f769713e259d85989ea8b7f8af95c` |
| `build/live-trace-20260925/linux-worker-state-final.tar.gz` | 134,380 | `b64e81c90b232f46b400a63ed08f5660f46e34fedb1f67b64afba079d8d36364` |
The final Windows state is retained under
`build/live-trace-20260925/windows-localappdata/TRUF/RemoteWorker`.
## Deliberately retained live state
The user requested that validation state not be restored. The following changes
therefore remain deliberate:
- accepted Windows and Linux trusted manifests remain installed;
- the keycheck queue maximum remains increased from 4,096 to 8,192 items;
- the Linux validation user remains enabled at assignment cap `0`;
- the Windows validation user remains enabled at assignment cap `1`;
- both validation workers themselves are stopped;
- normal global controls remain open at revision `126`; and
- the dashboard remains running.
The config editor still contains an intentionally stale candidate based on the
pre-validation active hash. It must not be applied without first rebasing it onto
the current active configuration.
## Release conclusion
The accepted worker artifacts passed live cross-platform execution, concurrent
dispatch, bounded progress/diagnostics, durable recovery, one-authoritative
receipt handling, normalized ingestion, projection, audit, local shutdown, and
operator UI validation. The complete cohort settled without leaked worker,
bundle, projection, or quarantine capacity.
Before broad rollout, deploy and revalidate the DockerHub SQL correction, decide
release treatment for the Windows timestamp and terminal-deadline defects, fix
network error classification, and address server disk pressure and DockerHub
credential health. The OpenSpec change is complete but remains unarchived until
explicitly requested.
@@ -0,0 +1,269 @@
# Worker Operator Experience Validation - 2026-09-24
## Scope and acceptance
This report closes the release and production-proof work for OpenSpec change
`add-worker-operator-experience`. Validation covered the shared worker event
contract, local supervisor and contained runner, progress and diagnostics APIs,
admin projections, reproducible Windows and Linux packages, packaged
cross-platform operation, bounded production behavior, and final restoration.
Acceptance required:
- the focused unit, integration, protocol, and package matrix to pass;
- independently reproducible Windows and Linux artifacts with documented and
registered package manifests;
- packaged Windows/Linux evidence for multi-slot operation, outage and restart
recovery, durable bundles and receipts, shutdown, and local cleanup;
- bounded production evidence for progress, a full scan-stage timeout,
diagnostics, reconciliation, and restoration; and
- a from-zero operator runbook covering acquisition through removal.
No production assignment was repeated to prepare this report. All production
facts below are derived from the retained validation snapshots. Raw targets,
findings, credentials, private routes, runtime configuration, and authenticated
worker command lines are intentionally excluded.
## Focused test matrix
The final 14-file worker-operator matrix completed on 2026-09-25:
```text
355 passed, 3 skipped in 49.54s
```
The matrix includes worker API/runtime, assignment and contained-runner,
CLI/contracts/local state, observability persistence, package, Linux handoff,
supervisor, remote database, scan execution, and admin API coverage. The three
independent-watchdog timing regressions also passed after their test setup bounds
were stabilized. The timing change did not alter product deadlines or watchdog
behavior.
An unrestricted repository-wide test run is not a release gate for this change:
the checkout has unrelated missing private/generated assets and platform
assumptions. The focused matrix, packaged E2E, production evidence, and strict
OpenSpec validation are the scoped gates.
## Reproducible artifacts
### Windows portable package
Final independently built archives:
- `build/operator-experience-validation/windows-i.zip`
- `build/operator-experience-validation/windows-j.zip`
Both archives have the following identical identities:
| Identity | Value |
| --- | --- |
| Archive bytes | `134850988` |
| Archive SHA-256 | `6ea9290736a059f1e17d8e89d9cf83506fa4abe2ba2f3731a7422a7b0f386e97` |
| Package manifest identity | `78a962b2bd3fa411413c79e9a8ffb021608a08ff020b1ad851f4505ea634b2b6` |
| Build-input identity | `6991ebbce6ae758c2bdd19a6ae934335aa585a50f86b18ccde8d88bca40ce436` |
| Raw `worker-package.json` SHA-256 | `e0b17d70fcb868fe39fac45ab6e05a17c6d40852e6034010fb63b6cab31f8a3c` |
Acceptance used the fresh extraction at `build/pwe-final-i-extracted`. Its own
`prepare-worker.ps1` established protected explicit ACLs before direct package
verification. Older G/H Windows archives are excluded because their preparation
script could leave packaged executables inaccessible.
### Linux worker image
Final independently built local tags:
- `truf-worker-test:operator-experience-final-3g`
- `truf-worker-test:operator-experience-final-3h`
Both provenance-disabled builds have the following identical identities:
| Identity | Value |
| --- | --- |
| Worker package identity | `45588f2cf406b41b239cfa3b8a9dc83fe84b587229bc997b2729016e1f0dde42` |
| Image manifest / accepted image ID | `sha256:3a088f5743121d823aae132234a29730a84339cecbfda5fc601e8e942f9948c3` |
| Config SHA-256 | `sha256:687a1c4c51c1b962c7fa7ea0cc4b04d159e7ba4f94ef347940c9fb225f7cb87d` |
| Raw `worker-package.json` SHA-256 | `ee926cce3c19e9e6094753f51fa902415bd7364c24fa649cd0c1b659c0aa4d60` |
The retained manifest snapshot is
`build/operator-experience-validation/linux-worker-package-g.json`.
### Trusted manifest registration
The accepted Linux and Windows manifests were registered after packaged E2E
acceptance. Their remote SHA-256 values match the raw manifest hashes above.
Both files are owned by `root:root` with mode `0644`; pre-change backups remain
intact and upload temporary files were removed. Registration required no runtime
restart or configuration mutation, and canonical health remained successful.
## Packaged Windows/Linux E2E
Run `35f3f52e232067c1` passed with the freshly extracted/prepared Windows I
package and Linux G image. The safe summary is
`build/pwe-35f3f52e232067c1/summary.json`.
The gate confirmed:
- real packaged Windows and Linux operation at two slots;
- server outage handling and restart recovery;
- durable and direct assignment bundle paths;
- authoritative receipt handling;
- graceful shutdown receipts;
- no active local work after completion while intentionally retained abandoned
roots remained inactive;
- matching normalized cross-platform evidence; and
- complete cleanup of owned resources with foreign Docker state unchanged.
## Production evidence
### Reconciliation
The retained snapshot records 34 issued assignments: 33 accepted and one
intentional expected expiry. All 33 accepted bundles were ingested, settled, and
projected. Final unresolved, precommit, quarantine, and drain-blocker counts were
zero.
Evidence sources:
- `build/operator-experience-validation/final-evidence.json`
- `build/operator-experience-validation/progress-v3-evidence.json`
- `build/operator-experience-validation/timeout-evidence.json`
- `build/operator-experience-validation/server-baseline.json`
### Duration percentiles
The table reports every retained end-to-end metric group. Values are seconds.
`Sufficient` means the server-side minimum sample count of five was met. Rows
below that minimum are retained observations, not statistically sufficient
percentile estimates.
| Platform | Source | Outcome | Samples | p50 | p95 | p99 | Sufficient |
| --- | --- | --- | ---: | ---: | ---: | ---: | --- |
| Linux | DockerHub | degraded | 1 | 305 | 305 | 305 | no |
| Linux | DockerHub | error | 1 | 19 | 19 | 19 | no |
| Linux | GitLab | error | 1 | 5371 | 5371 | 5371 | no |
| Linux | HuggingFace | expired | 1 | 7219 | 7219 | 7219 | no |
| Windows | DockerHub | clean | 2 | 31 | 31.9 | 31.98 | no |
| Windows | DockerHub | degraded | 4 | 275.5 | 443.45 | 462.29 | no |
| Windows | DockerHub | error | 7 | 619 | 1679.5 | 1731.1 | yes |
| Windows | GitLab | clean | 3 | 27 | 873 | 948.2 | no |
| Windows | GitLab | error | 6 | 342.5 | 1723.75 | 1916.75 | yes |
| Windows | HuggingFace | clean | 1 | 1019 | 1019 | 1019 | no |
| Windows | HuggingFace | error | 7 | 971 | 3092.6 | 3452.12 | yes |
The snapshot contains 11 Linux and 21 Windows phase/outcome metric groups in
total. Three Windows end-to-end error groups met the minimum; the other 29
phase/outcome groups did not. Rollout decisions must therefore preserve the
sample-count qualification rather than treating all reported percentiles as
stable capacity estimates.
### Progress and watchdog evidence
Reservation `1455` is the retained complete-stage progress reference. It was
acknowledged with an accepted bundle and persisted ten monotonic events spanning
`assigned`, `preparing`, `waiting_permit`, `scanning`, `filtering`, `cleaning`,
`bundling`, `uploading`, and `awaiting_receipt`. This demonstrates one coherent
server-visible sequence across the complete local execution and upload boundary.
Independent watchdog fault-injection coverage passed for blocked state
persistence, startup-gate persistence, and event draining. The contained runner
tests verify bounded process-tree termination rather than relying on scanner
cooperation. Packaged E2E additionally passed its watchdog, restart, durable
bundle, and cleanup gates.
### Natural full-stage timeout
Reservation `1453` is the retained natural timeout reference. The scan ended
with one `timeout / scan.stage_timeout / scanning` diagnostic after 603.367
seconds. The diagnostic was current and available, with no body or process-log
payload fabricated for the exception. The end-to-end assignment-resolution
duration was 619 seconds.
The scan outcome was `error`, while the transport outcome was independently
accepted: the bundle was acknowledged, the projection completed, and the
diagnostic was attached to the authoritative result. This confirms that a hard
scan-stage timeout remains a normal, uploadable terminal result and does not
collapse scan, transport, and projection outcomes into one status.
### Diagnostic and admin snapshots
The final diagnostic aggregate contains five grouped rows and 23 occurrences:
| Category | Code | Phase | Occurrences |
| --- | --- | --- | ---: |
| scanner | `scan.result_error` | `scanning` | 20 |
| timeout | `scan.stage_timeout` | `scanning` | 1 |
| assignment expiry | `assignment.deadline_expired` | `assigned` | 1 |
| network | `scan.result_error` | `scanning` | 1 |
The retained snapshots also confirm separate assignment and scan outcomes,
ordered progress, current diagnostic availability, accepted receipt state,
ingestion/settlement/projection completion, duration metrics, and zero unresolved
or precommit work. This is the durable machine-readable substitute for copying
private admin pages or unbounded diagnostic bodies into the report.
## Sanitized operator transcript
The release and validation sequence was:
1. Build the Windows package twice from the same reviewed inputs and compare the
archive, package-manifest, build-input, and raw-manifest identities.
2. Extract Windows I into a fresh directory, run its packaged
`prepare-worker.ps1`, and run direct package verification.
3. Build Linux G and H independently with provenance disabled and compare image,
config, package, and raw-manifest identities.
4. Run `docker/verify_packaged_workers.py` with the accepted Windows extraction,
Linux image, and isolated test image; retain only its safe summary and owned
evidence directory.
5. Register the two accepted trusted manifests through the reviewed deployment
path and recheck canonical runtime health.
6. Use typed operations to bound production dispatch, start at assignment cap
`1`, observe status/attach/history and server progress, exercise normal,
timeout, outage, and restart paths, and reconcile accepted, ingested, settled,
and projected counts.
7. Restore standard identities and normal/open controls, disable/revoke temporary
validation identities, and recheck runtime and edge health.
8. Run the exact focused pytest matrix recorded above.
Authentication values, worker argv, raw targets/findings, private route names,
and runtime configuration are omitted by design.
## Known limits
- There is no public ZIP download, image registry, installer, or automatic
updater. Release artifacts must move through a trusted channel and match a
server-registered manifest.
- Completed runner roots move under top-level `work/abandoned` and are retained
for at least 60 seconds; normal retention maintenance runs every 300 seconds.
They are inactive evidence, not live work.
- Most retained percentile groups have fewer than five samples. Their values are
useful validation observations but not stable performance baselines.
- Ownership fencing guarantees one authoritative acceptance, not exactly-once
physical execution across a long partition and server-side expiry/reissue.
- Local state remains recovery authority until the server resolves the slot.
Operators must not remove pending bundles or work trees to clear an alert.
## Rollout, rollback, and restoration
Rollout uses the exact accepted package identities, begins with one user/device at
server cap `1` and local parallelism `1`, and requires one accepted, ingested,
settled, and projected assignment before expansion. Caps and client count should
increase in stages while unresolved/precommit counts, diagnostic availability,
duration sample counts, and authenticated contact remain observable.
Rollback first sets the affected cap to `0`, allows pending uploads to resolve,
and obtains a graceful shutdown receipt. The operator then returns to the
previous exact artifact while preserving the same private state tree or volume.
Additive server progress and diagnostic records do not require schema rollback.
Final restored production state:
- operations controls normal/open at revision `126`;
- standard WSL production worker user enabled at assignment cap `1`;
- standard production device enabled and not revoked;
- temporary validation identities disabled/revoked;
- unresolved, precommit, quarantine, and drain blockers at zero;
- canonical runtime healthy; and
- edge service remained available.
Strict OpenSpec validation passed. The operationally validated change is ready
for archival.
@@ -0,0 +1,473 @@
# Worker Parallelism Validation, 2026-09-23
## Scope
This report records the production validation performed on `sec` using the local
Windows computer:
- bounded discovery for GitLab, DockerHub, and HuggingFace;
- a native Windows protocol-2 worker with client and server parallelism 3;
- a dual-worker run with native Windows parallelism 1 and the existing trusted
WSL production worker at cap 1;
- authoritative reconciliation from discovery through queue, reservation,
receipt, bundle ingestion, scan settlement, projection, and JSONL append;
- restoration of production source/capacity settings and normal controls.
No provider credential, device token, raw target, raw finding, runtime YAML, or
worker command line is included in this report.
## Result
The core validation passed.
- Native Windows reached exact unfinished concurrency 3 and never exceeded 3.
- The dedicated Windows cohort accounted for 180 accepted assignments across all
three sources.
- Windows and WSL independently reached concurrency 1 at the same time, for exact
combined concurrency 2, and never exceeded their individual caps.
- All accepted results were ingested, queue-settled, projected, and appended as
required.
- Final cohort and global lineage checks reported zero violations.
- No pipeline quarantine or failed hold was created.
- Production controls and the WSL worker were restored.
The originally requested 10-hour observation was not completed continuously.
After successful active parallelism-3 work, 14,105 seconds (about 3 hours 55
minutes) of detached idle stability observation completed before the user changed
the objective to the dual-worker test. The active concurrency evidence itself is
complete; the residual limit is soak duration, not functional coverage.
## Baseline
The fresh pre-test snapshot was stored root-only on `sec`:
- Path: `/opt/truf-remote-server/staging/windows-p3-before2.json`
- SHA-256: `469bcd69888cacce55523fef17508ee48191c4c879a6d39f2251abad4aa0efdb`
- Active config SHA-256:
`e48797689ab0ee7328d21cf5c23d0ffedb979dba492bf49d1da3afe4b075cf5b`
- Config bytes: 37,233
- Controls: revision 56, discovery open, dispatch open, drain normal
- Existing blocker: one stale production assignment, allowed to settle naturally
- Runtime image:
`sha256:3b4d6e19e29e85a32b75d64265d75e71100d3f7f00221d21de544256dadd25cf`
- Failed hold: absent
- All recorded orphan/mismatch invariants: zero
Baseline high-water values included reservation 929, scan 926, projection job
926, projection append 984, error 1919, finding 127, and zero quarantine rows.
The existing WSL assignment was never canceled or directly mutated. Dispatch was
paused and the lease was allowed to resolve before changing worker caps.
## Windows Package
The trusted run artifact was built from pinned local caches:
- Package directory:
`build/worker-parallelism3-20260922/package-fixed1`
- Archive:
`build/worker-parallelism3-20260922/truf-worker-windows-x86_64-fixed1.zip`
- Archive SHA-256:
`6ca460a47d31606c430e8c6f081443530892332cfc183fc0c6edfe7568a454cc`
- Manifest file SHA-256:
`cd55820e0a04a04ded9ba340b9c6e5c6259422aab80989180a6ba28ff3b77103`
- Canonical code-manifest SHA-256:
`e981da19283aad5971fc2865268bb8151b3540d71fbd057ae56d9d2379993374`
- Platform: `windows-x86_64`
- Protocol and bundle schema: v2
- Capabilities: GitLab `exact_git_v1`, DockerHub `docker_direct_v1`,
HuggingFace `huggingface_space_v1`
The registered Windows v2 manifest was stale relative to the current production
authority files. It was not overwritten. The fixed package was registered as the
new root-owned `windows-worker-package-v3.json` trust authority.
### Reproducibility defect
Two initial package builds differed only in four timestamp bytes in each
pip-generated Windows launcher executable and the corresponding RECORD hashes.
The package builder now normalizes the embedded launcher ZIP timestamps to the
DOS epoch and recomputes the RECORD rows.
Changed local source:
- `app/worker_package_builder.py`
- `tests/test_worker_package.py`
Verification:
- Two post-fix archives were byte-identical.
- `python -m pytest tests/test_worker_package.py -q`: 16 passed.
## Discovery
The temporary discovery candidate was derived from the exact active config.
GitLab used 16 reviewed keyword queries, one API page, one result per page, and
one target maximum per query. DockerHub used 16 reviewed keyword queries, one
repository result per query, and one image per repository.
HuggingFace was deliberately reported differently: the configured implementation
does not perform keyword search. It requests the newest-modified Spaces with a
fixed API limit of 100. `pages: 1` therefore means one real recent page, not one
keyword result. Repeated polls occurred while GitLab and DockerHub rotated their
queries; deduplication bounded the admitted rows.
Final bounded discovery evidence:
- GitLab: 16/16 distinct successful query rotations, 7 new queue rows
- DockerHub: 16/16 distinct successful query rotations, 10 new queue rows
- HuggingFace: 17 successful recent-page cycles, 31 new queue rows
- Discovery failures: 0
- Hard limits: GitLab 20, DockerHub 20, HuggingFace 120; none exceeded
## Native Windows Parallelism 3
An isolated temporary server user and device were created with typed, audited
`AdminService` operations. The one-time token was transferred privately, the
server transfer file was deleted after successful launch, and the native worker
used a dedicated protected `LOCALAPPDATA` state tree.
The first foreground launch exposed a local automation issue: the tool runner
retained the descendant process and did not return. The worker itself survived.
Subsequent starts used `Win32_Process.Create` so the worker was genuinely
detached. No process command line was inspected.
### Capacity findings
Client `--parallelism 3` and server user cap 3 were not by themselves enough to
permit three physical scans.
The first run reached only one active reservation because the production config
had:
- `global.max_active_scans = 1`
- 64 MiB maximum event size
- 256 MiB projection backlog maximum
- 128 MiB projection headroom
The capacity model reserves twice the event maximum per assignment, plus
headroom. Three slots therefore require 512 MiB. The temporary test candidate
changed only:
- `max_active_scans: 1 -> 3`
- `projection_backlog_max_bytes: 256 MiB -> 512 MiB`
A second run still reached only one assignment. Admission evidence showed
`pipeline_capacity_closed` on the keycheck axis. The stable keycheck backlog was
127 items and each assignment reserved up to 2,000 items. The production limit
of 4,096 allowed one reservation but not three. The temporary candidate changed:
- `keycheck_queue_max_items: 4096 -> 8192`
No unrelated limit was increased.
### Successful run
After both capacity corrections:
- Exact max unfinished concurrency: 3
- Never exceeded: 3
- Issued: 180
- Accepted: 180
- Prebundle failures: 0
- Expired: 0
- Unfinished after settlement: 0
- Accepted-not-ingested: 0
- Queue-not-settled: 0
- Projection-not-completed: 0
Source totals:
| Source | Accepted |
|---|---:|
| DockerHub | 63 |
| GitLab | 65 |
| HuggingFace | 52 |
Scan outcomes were transport-successful but not necessarily target-clean:
| Source | Outcome | Count |
|---|---|---:|
| DockerHub | clean | 44 |
| DockerHub | degraded | 18 |
| DockerHub | error | 1 |
| GitLab | clean | 57 |
| GitLab | error | 8 |
| HuggingFace | clean | 43 |
| HuggingFace | error | 9 |
These scan statuses are provider/target results. They are not lost transport,
ingestion, or projection records. All 180 authoritative results settled.
## Stability Observation
After the 180-assignment cap closed dispatch for the test identity, the worker
remained online at cap 0 under detached monitoring.
- Successful checks: 43
- Elapsed observation: 14,105 seconds
- Runtime restarts: 0
- Edge restarts: 0
- Native worker OOM/restart: none
- Pipeline debt: 0
- Quarantine: 0
- Global invariants: 0
A single strict-health probe can observe a discovery producer during its normal
short startup window. Monitors were corrected to require failure across three
attempts with 20-second delays rather than treating one transient sample as a
production defect.
## Dual Worker Validation
The user then requested a second topology:
- native Windows worker: real client parallelism 1 and server cap 1;
- existing trusted production WSL worker: cap 1;
- desired combined concurrency: 2.
The native worker was cleanly stopped at cap 0, relaunched from the same trusted
package and protected state with parallelism 1, and verified by PID/resource
evidence. The production WSL token and command line were never inspected.
### Monitor corrections
Several fail-closed attempts improved the monitoring policy without losing or
canceling work:
- A scanner subprocess can remain active without an HTTP contact until it
reports. Contact age alone must not fail an identity with unfinished work.
- A worker may honor a prior cap-0 `Retry-After` after caps change. Arming uses a
separate 600-second threshold; the 180-second idle threshold applies only
after dispatch has opened.
- Operation IDs are one-use and identity-bound. Every retry used a fresh audited
operation namespace.
- A transient PostgreSQL connection timeout while reopening a read-only monitor
connection terminated observability after caps and gates were already safely
closed. The monitor now retries database opens three times with 20-second
delays. A separate read-only settlement monitor completed the active run.
### Attempt 1 evidence
The first active dual run already proved max concurrency Windows 1, WSL 1,
combined 2. It fail-closed because of the old contact policy while WSL was doing
a long DockerHub scan.
- Windows: 19 accepted assignments
- WSL: 2 accepted HuggingFace assignments
- WSL: 1 DockerHub reservation naturally expired after its two-hour lease
- No assignment was canceled
### Final attempt 4
The final bounded run used baseline reservation ID 1131.
- Issued: 30
- Windows issued/accepted: 29/29
- WSL issued: 1
- WSL accepted: 0
- WSL expired/refunded: 1 long DockerHub assignment after its two-hour lease
- Max Windows concurrency: 1
- Max WSL concurrency: 1
- Max combined concurrency: 2
- Accepted-not-ingested: 0
- Queue-not-settled: 0
- Projection-not-completed: 0
- Cohort integrity violations: 0
- Global invariant violations: 0
- Failed holds in the cohort window: 0
- Quarantine bytes/items: 0/0
Windows source/outcome totals in the final run:
| Source | Outcome | Count |
|---|---|---:|
| DockerHub | clean | 12 |
| GitLab | clean | 7 |
| HuggingFace | clean | 9 |
| HuggingFace | error | 1 |
The WSL expiry does not invalidate the concurrency result: both independent
clients simultaneously held one real reservation and never exceeded cap. Earlier
in attempt 1, the same WSL path also completed two accepted HuggingFace results.
Authoritative dual evidence:
- `/opt/truf-remote-server/staging/windows-dual-worker-report.json`
- SHA-256:
`8fd03bf79960393406ef1a0f788d90db3ed3521d32ffa83f802c9d1c2c3509bf`
## Data Reconciliation
The validation distinguished assignment acceptance, bundle ingestion, queue
settlement, and projection completion.
Checks covered:
- accepted reservation has a durable receipt;
- accepted reservation has exactly one matching bundle;
- reservation/bundle/scan event IDs and hashes agree;
- target scan references the same queue row;
- queue completion is `applied`;
- projection job exists, matches the event identity, and is completed;
- required `scan_results` append exists;
- required `found_secrets` append exists when requested by the mask;
- projection append event identity matches the job;
- no orphan bundle, scan, finding, error, or projection job exists;
- JSONL files are present, newline-terminated, and cursor offsets match.
All cohort and global mismatch counters were zero at the final report points.
## Code Defects Fixed
### Integer cap zero
`AdminService._validate_cap()` used `str(value or '')`, converting integer zero
to an empty string even though cap 0 is valid. It now distinguishes `None` from
zero.
Changed local source:
- `app/admin_api.py`
- `tests/test_admin_api.py`
Verification:
- `python -m pytest tests/test_admin_api.py -q`: 62 passed.
The currently deployed runtime image predates this source fix. Run helpers used
canonical string `"0"` through the same typed/audited API; no direct database
mutation was used.
### Operational helpers
Run-specific helpers were added under `build/sec-deploy/` for config lifecycle,
controls, typed worker administration, monitoring, observation, local launch,
safe stop, settlement, and reporting. Python helpers passed `py_compile` and
Ruff; PowerShell helpers passed parser validation before use.
## Configuration Restoration
Temporary active config identities were:
| Stage | SHA-256 |
|---|---|
| Bounded discovery + v3 profile | `ec2ae20f9a23ee9ec423c28f7da0b45db3eaff17abb3f484e1a1f0c9227fada4` |
| Parallel scan/projection capacity | `4fee05f3a9ae001767f2de0388dcd5db75183dd954be8e024e98f47d4b64b6b1` |
| Parallel keycheck capacity | `b3509aacc53ab0bfa1e6b13dd789f889d9fea34daf0f680579d2a54697393bf9` |
| Restored production semantics + v3 trust | `c0966cac4f7f0610a813fa8732e91953f2e3e838ad880f91fd1a9437096925c7` |
The final config restores the original production discovery/source settings,
global scan capacity, projection capacity, keycheck capacity, and supervisor
interval. The only intentional permanent semantic difference from the initial
config is the Windows compatibility profile path changing from stale v2 to the
reproducible/current v3 manifest. Therefore the final config hash intentionally
differs from the initial hash while retaining the same 37,233-byte size.
All four managed config apply operations succeeded and reconciled. The final
apply operation was:
- `a9943967-372e-522a-8386-4f9465cc03f9`
The complete root-only operation export contains 69 operations for the run
actor, all with status `succeeded`:
- `/opt/truf-remote-server/staging/windows-p3-operations.json`
- SHA-256:
`318aaa56cc2ba348e905b72fd63d7795e24cbd1aa9e84d0a75fa341744ef0acf`
## Final Production State
The temporary native process was stopped only after its local bundle/work trees
and authoritative pipeline were empty.
Typed cleanup completed:
- temporary Windows cap: 0
- temporary device: revoked
- temporary user: disabled
- server plaintext token transfer: absent
- production WSL cap: 1
Controls after restoration:
- revision: 98
- discovery: open
- dispatch: open
- drain: normal
- effective gates: open
The production WSL container was idle but still honoring a previous cap-0 retry
delay. With zero unfinished work it was safely restarted using the same container,
token, and state. It immediately made a fresh server contact and received a new
production assignment, proving restored admission.
At the final snapshot:
- production WSL contact was fresh and one real production assignment was active;
- blocker count 1 was therefore expected, not a drain defect;
- WSL container running, OOM false, restart count 0;
- runtime healthy, restart count 0;
- edge running, restart count 0;
- host-agent, host Caddy, and X-UI active;
- strict runtime health: ACTIVE/READY with DockerHub, GitLab, HuggingFace,
janitor, JSONL projector, result ingester, and Worker API;
- failed hold absent.
Runtime image remained:
`sha256:3b4d6e19e29e85a32b75d64265d75e71100d3f7f00221d21de544256dadd25cf`
## Final Snapshot
- Path: `/opt/truf-remote-server/staging/windows-p3-final.json`
- SHA-256: `1e3ab6bdb50c5be1416974bf56235c220166a6b731e76a87d8abebb947aae074`
- Bytes: 6,883
- Config SHA-256:
`c0966cac4f7f0610a813fa8732e91953f2e3e838ad880f91fd1a9437096925c7`
- Reservation high-water: 1165
- Scan high-water: 1159
- Projection job high-water: 1159
- Projection append high-water: 1239
- Error high-water: 1945
- Finding high-water: 127
- Quarantine rows: 0
- All six recorded global orphan/mismatch invariants: 0
- `scan_results.jsonl`: 4,592,879 bytes, cursor exact, newline-terminated
- `found_secrets.jsonl`: 146,366 bytes, cursor exact,
newline-terminated
Final strict-health evidence:
- `/opt/truf-remote-server/staging/windows-p3-final-health.json`
- SHA-256:
`68c2b0dcd16591b2407fefcb5c3921aa7f58443011539ff912e2cb19b9fdef13`
## Retained Local Evidence
Per the user's explicit instruction on 2026-09-23, no additional local files were
deleted during final reporting.
The following remain intentionally retained:
- protected local one-time token file;
- dedicated native worker state directory;
- worker stdout/stderr and PID evidence;
- fixed package and both reproducibility builds;
- local run-specific helpers and reports.
The token is no longer usable because the server device is revoked and user is
disabled, but the local protected file remains until the user authorizes deletion.
## Residual Limits
- The continuous 10-hour soak was shortened by the later dual-worker request.
- HuggingFace configured discovery is recent-page discovery, not keyword search.
- Long DockerHub work twice reached the fixed two-hour WSL lease and expired;
expiry/refund behavior was correct, but those targets did not produce bundles.
- The local source fixes for package reproducibility and integer cap zero are
tested in this checkout but were not deployed as a new production runtime
image during this validation.
- Local sensitive/state evidence is deliberately retained and requires a later
explicit cleanup instruction.