Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-09-23
@@ -0,0 +1,206 @@
## Context
The protocol-2 remote worker has proven its authoritative data path under real production load, including native Windows concurrency 3 and simultaneous Windows/WSL execution. Its operator experience has not reached the same level: the package exposes only `--server`, `--token`, and `--parallelism`; successful work is mostly silent; scanner output is captured out of view; slot state remains `assigned` during synchronous execution; and the server sees no phase progress between claim and result.
Two production DockerHub assignments remained unresolved until the fixed 7,200-second assignment deadline even though the configured Docker target timeout was 600 seconds. The current watchdog terminates the owned TruffleHog process tree, but permit acquisition and surrounding Python work such as cleanup, filtering, result serialization, bundle staging, and handoff are not one hard-preemptible unit. Current evidence cannot identify which phase stalled.
The administration page compounds the problem by labeling `result_reservations.last_error_code` as `Error category`. That field describes assignment transport failures and expiry, while accepted scan errors live in `target_scans`, `errors`, and result metadata and are not shown as assignment diagnostics.
Legacy `supervisor.py --attach` demonstrates a useful interaction model: a verified background instance, an authenticated local control channel, a live status table, bounded logs, and detach without shutdown. The remote-worker package does not contain that supervisor and needs a smaller worker-specific implementation.
## Goals / Non-Goals
**Goals:**
- Deliver one coherent operator product rather than temporary UI over incomplete fields.
- Give owners a cross-platform lifecycle CLI with detached operation, attach, status, logs, history, diagnostics, and doctor commands.
- Define one phase/event model used locally, over the Worker API, in PostgreSQL, and in the admin UI.
- Apply a hard local scan-stage deadline to the complete assignment execution unit, not only its scanner child process.
- Preserve exact diagnostic material when available and explicitly describe size truncation or other transformations.
- Separate assignment transport outcome, scan outcome, and diagnostics in storage and UI.
- Make global and per-source assignment deadline policy visible and measurable with phase duration percentiles.
- Ship a from-zero operator guide and fault-injection coverage as part of the same change.
**Non-Goals:**
- Replacing PostgreSQL assignment authority, immutable assignment expiry, durable receipts, or bundle ingestion.
- Renewing assignment ownership from progress events.
- Reporting fabricated percentage completion when the scanner has no reliable denominator.
- Rebuilding the full server runtime supervisor inside the worker package.
- Adding a temporary compatibility-only error page that will be removed after diagnostics land.
- Adding a new diagnostic masking or redaction subsystem.
## Decisions
### 1. One worker supervisor owns lifecycle and authoritative local projections
The package will expose a `truf-worker` command with `install`, `run`, `start`, `stop`, `status`, `attach`, `logs`, `history`, and `doctor` subcommands. `run` executes the supervisor in the foreground; `start` launches the same supervisor detached and waits for a startup handshake. Docker continues to run the supervisor in the foreground under Tini while Docker supplies detachment.
The supervisor owns the singleton lock, slot controllers, local control endpoint, rotating human log, event JSONL, current status snapshot, and terminal local history. A verified instance record binds PID, process creation identity, executable, package identity, local control endpoint, and lifecycle state. `attach` reads status/events through the local control channel; it does not attach to arbitrary stdout and detaching never stops the worker.
The current three-flag invocation remains a foreground-run migration alias for already shipped package launch definitions, but all new documentation and generated launchers use explicit subcommands.
Alternative considered: add more prints to `remote_worker_client.py`. Rejected because prints cannot provide detached lifecycle control, reliable concurrent-slot rendering, machine output, history, or process identity.
### 2. A single append-only event model drives every projection
Each state transition emits a versioned event with a monotonic local sequence:
```json
{
"schema": 1,
"sequence": 42,
"timestamp": "2026-09-23T12:00:00Z",
"instance_id": "...",
"slot_id": 0,
"reservation_id": 123,
"source": "dockerhub",
"type": "slot.phase",
"phase": "scanning",
"phase_started_at": "...",
"scan_deadline_at": "...",
"assignment_deadline_at": "...",
"progress": {}
}
```
Canonical phases are `idle`, `claiming`, `assigned`, `waiting_permit`, `preparing`, `resolving`, `downloading`, `cloning`, `scanning`, `filtering`, `cleaning`, `bundling`, `uploading`, `awaiting_receipt`, `backoff`, `draining`, and `stopped`.
The local status JSON is a rebuildable projection of the event stream, not a second independently written state machine. Server progress uses the same event schema with a server-assigned receive timestamp. Terminal history records authoritative receipt/prebundle/stale outcomes plus durations.
Alternative considered: define separate local, API, and admin status shapes. Rejected because they would drift and recreate the current ambiguity.
### 3. Execute each scan stage in a supervised assignment runner process
Threads in the controller continue to own claim/recovery/upload state, but scanner execution moves into a package-local assignment runner subprocess. The controller writes one validated input record and starts the runner in a contained process tree with a dedicated work/output directory.
The hard scan-stage deadline starts before permit acquisition and covers:
- permit wait;
- assignment preparation and source resolution performed locally;
- downloads/clones;
- scanner subprocess execution;
- filtering and result conversion;
- cleanup;
- result and bundle staging.
The runner emits phase events over a local pipe/file protocol. At the deadline the controller terminates the runner process tree, atomically detaches its work directory for later janitor handling, and creates a normal timeout result bundle with the phase and diagnostic envelope. The controller then uploads that result while the immutable assignment deadline still permits it.
Normal cleanup remains cooperative. Cleanup that exceeds the deadline cannot keep the slot occupied; abandoned work is moved into a janitor-owned tree and reported in status.
Alternative considered: add deadline checks around existing Python calls. Rejected because blocking Python/filesystem/library calls cannot be hard-preempted reliably in the controller process.
### 4. Progress events are informational and never renew authority
The Worker API gains an authenticated progress endpoint accepting ordered phase events for the worker's current reservation. It validates reservation/device ownership and monotonically advances the latest accepted event sequence. Duplicate events are idempotent.
Progress updates do not change `remote_expires_at`, queue leases, bundle capacity, or receipt authority. Failure to send progress does not abort a healthy local scan; events remain in the local journal and retry with bounded backoff. The server can therefore display `last phase` and `last progress` without turning progress into a lease heartbeat.
Alternative considered: renewable leases. Rejected because a stuck client could retain work indefinitely and because the fixed-expiry fencing model is already proven.
### 5. Server-owned deadlines support global fallback and per-source policy
`supervisor.worker_api.assignment_ttl_seconds` remains the global fallback. A managed per-source override map adds GitLab, DockerHub, and HuggingFace assignment TTL values. The server selects and commits the immutable deadline when issuing an assignment and includes the effective scan and assignment deadlines in the response.
The editor labels these separately as `Target scan timeout`, `Assignment deadline (end-to-end)`, and `Result upload body deadline`. Validation retains absolute bounds and checks that each effective assignment deadline covers its source scan timeout, upload deadline, and handoff margin.
The server aggregates phase and end-to-end durations by source and outcome as p50, p95, and p99. Configuration remains explicit; metrics inform changes but do not silently rewrite policy.
Alternative considered: let each worker select or renew its TTL. Rejected because workload policy belongs to the assigning server and must be consistent for queue fencing.
### 6. One diagnostic envelope spans scan and prebundle failures
Diagnostics use one versioned envelope with indexed dimensions and exact optional payloads:
- diagnostic ID and schema;
- reservation, scan event, slot, attempt, and source;
- phase, kind, category, stable code, summary, and retryable disposition;
- provider operation and HTTP status/content type/request ID;
- process name, exit code, signal, and timeout state;
- exception type/message/fingerprint;
- raw body material;
- log stream head/tail material;
- occurred/captured/received timestamps;
- original byte count, stored byte count, content hash, and truncation state.
Categories are broad query dimensions such as `authorization`, `rate_limit`, `not_found`, `network`, `timeout`, `scanner`, `storage`, `protocol`, and `assignment_expired`. Stable codes express the concrete cause, such as `docker.manifest_http_403` or `trufflehog.exit_nonzero`. Phase is independent of category.
Text/bytes are preserved as captured. Non-text bodies use an explicit encoding field. Storage bounds are deterministic: body excerpt 16 KiB, combined log head/tail 32 KiB, one transmitted envelope 64 KiB, at most 32 diagnostics and 256 KiB per assignment. Metadata records every truncation; no truncation is presented as a complete body.
Accepted result bundles gain a diagnostic frame. Prebundle reports carry the same envelope under the existing JSON body limit. The old E-frame error string remains ingestible during rollout but is projected into the new model exactly once.
Alternative considered: keep scanner error strings, reservation error codes, and source metadata as separate taxonomies. Rejected because operators cannot correlate or filter them consistently.
### 7. Local diagnostics retain complete operator evidence when available
The supervisor writes:
```text
control/worker.instance.json
control/worker.status.json
events/worker-events.jsonl
history/worker-history.jsonl
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.json
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.body
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.log
logs/worker.log
```
The JSON envelope points to optional body/log files and records their hashes and sizes. Local retention is configurable by age and total bytes and is reported by `status` and `doctor`. Rotation never mutates terminal history entries; it changes attached-artifact availability explicitly.
Alternative considered: store every raw artifact directly in one JSONL. Rejected because large multiline/process output makes append recovery and bounded tailing expensive.
### 8. PostgreSQL stores diagnostics as first-class records
Add `worker_diagnostics` with an idempotent diagnostic UID, reservation FK, optional target-scan FK, indexed phase/category/code/kind/retryable columns, summary, canonical envelope JSON, received timestamp, and optional bounded body/log payload columns. Bundle ingestion writes diagnostics in the same transaction as the target scan and error projection. Prebundle diagnostics attach to the reservation before a target scan exists.
Existing `errors` rows remain the compatibility scan-error projection. Existing `last_error_code` is retained as assignment failure code but is no longer labeled as the complete error category.
Alternative considered: put all envelopes only into `metadata_json`. Rejected because filtering, detail lookup, idempotency, and prebundle diagnostics require first-class rows.
### 9. Admin views assignment outcome, scan outcome, progress, and diagnostics separately
The worker assignment table presents:
- assignment outcome: unfinished, accepted, prebundle failed, expired;
- scan outcome: clean, found, degraded, error, skipped, or not available;
- diagnostic count and highest-priority category/code;
- current/latest phase, phase age, assignment deadline, and last progress age;
- worker/device/package identity and slot where available.
Each row links to a detail page with an event timeline, duration breakdown, transport/receipt data, scan summary, diagnostics, raw body/log tabs, canonical JSON copy/download, and explicit truncation metadata. Filters operate independently on assignment outcome, scan outcome, source, phase, category, code, retryability, and time.
Alternative considered: make the existing `Error category` cell open a modal. Rejected because the list model itself conflates transport and scan semantics.
### 10. Human output and machine output are equal product contracts
Human status/attach uses a stable table and event stream. It displays elapsed time, configured scan deadline, assignment time remaining, last progress age, child state, and trustworthy counters. It never fabricates completion percentages.
`--json` commands emit one versioned JSON document. Follow modes emit NDJSON with one event per line and no decorative output. Tests treat both output forms as contracts.
## Risks / Trade-offs
- [Runner process split touches scanner lifecycle and recovery paths] -> Introduce it behind the same `execute_protocol2_remote_claim` contract, preserve deterministic bundle validation, and fault-test every boundary before replacing in-process execution.
- [Progress traffic increases database writes] -> Persist only monotonic phase transitions and coarse progress changes, deduplicate by reservation/sequence, and keep high-frequency local samples local.
- [Diagnostic payloads increase bundle and database volume] -> Enforce deterministic per-item/per-assignment byte and count limits, expose truncation metadata, and track storage usage.
- [One broad change can take too long] -> Implement as large vertical chunks that each finish a final architecture slice; do not ship throwaway status/error models.
- [Per-source policy adds configuration complexity] -> Keep one global fallback, explicit source overrides, editor-derived effective values, and validation based on existing scan/upload settings.
- [Local full-stage termination can leave work trees] -> Atomically detach them to janitor ownership and surface retained bytes/counts in status and doctor.
- [Old packages do not emit progress/diagnostics] -> Admin renders legacy records from existing fields and marks phase/diagnostic availability explicitly until packages are upgraded.
## Migration Plan
1. Add schema/event/diagnostic libraries, PostgreSQL tables, and read paths without changing current assignment execution.
2. Add the worker supervisor CLI and local projections while the current foreground invocation remains a migration alias.
3. Add assignment runner execution, full-stage watchdog, local phase events, and fault-injection tests.
4. Add Worker API progress and diagnostic transport, then enable server persistence and detail queries.
5. Replace the worker admin list/detail presentation and add percentile/deadline editor views.
6. Rebuild Windows/Linux packages, run protocol compatibility tests, then run bounded Windows and WSL production validation.
7. Update generated launchers and the from-zero operator guide; migrate the production worker launch definition to explicit `run`.
8. Remove the migration alias only in a separately declared breaking change after all known deployments use subcommands.
Rollback disables progress ingestion and runner selection while retaining additive event/diagnostic tables. Existing immutable assignment, bundle, receipt, and legacy E-frame paths remain authoritative throughout rollout.
## Open Questions
No blocking product questions remain. Exact local retention defaults and percentile windows can be selected from implementation benchmarks without changing the external contracts above.
@@ -0,0 +1,38 @@
## Why
Remote workers execute and settle production work correctly, but they behave as opaque background processes: an owner cannot tell whether a slot is downloading, scanning, cleaning up, bundling, uploading, stalled, or close to its deadline, while the admin console conflates assignment failures with scan errors. The next change should turn the proven worker engine into an understandable operator-facing product without building temporary UI around the current incomplete status and error fields.
## What Changes
- Add a worker-specific supervisor and CLI for install, foreground run, detached start, graceful stop, status, attach, live logs, local history, and diagnostics.
- Add one versioned assignment phase/event model shared by local status, attach, server progress, history, diagnostics, and admin rendering.
- Apply a hard local scan-stage deadline across permit acquisition, scanner execution, cleanup, filtering, and result staging instead of bounding only the scanner subprocess path.
- Make assignment deadlines observable and configurable as server-owned global and per-source policy, with explicit scan, upload, and assignment deadline semantics.
- Add a versioned diagnostic envelope for provider responses, process failures, exceptions, timeout state, structured categories/codes, and bounded raw body/log material.
- Persist local worker events and diagnostics as JSON/JSONL and transmit the same diagnostic model through terminal reports and accepted result bundles.
- Replace the ambiguous assignment-table `Error category` presentation with distinct assignment outcome, scan outcome, diagnostic count, and a clickable detail/timeline view.
- Add machine-readable `--json`/NDJSON output and human-readable status/attach views without inventing progress percentages.
- Add a from-zero operator guide covering installation, first start, attach/status, interpreting phases and errors, graceful drain/stop, recovery, update, and troubleshooting.
- Do not add a separate masking/redaction feature or silently generalize captured diagnostics. Storage bounds and any unavoidable transformation must be explicit in diagnostic metadata.
## Capabilities
### New Capabilities
- `worker-operator-supervisor`: Lifecycle CLI, detached supervision, attach/status/log/history/doctor commands, local projections, and operator documentation.
- `worker-progress-deadlines`: Canonical assignment phases, local and server progress events, full scan-stage watchdog behavior, deadline policy, and duration percentile observability.
- `worker-diagnostics`: Unified diagnostic envelope, local diagnostic archive, terminal/bundle transport, taxonomy, retention bounds, and exact transformation metadata.
- `worker-admin-experience`: Assignment and scan outcome separation, diagnostic persistence/querying, clickable timeline/detail UI, and human/machine-readable diagnostic views.
### Modified Capabilities
None. There is no synchronized main-spec directory in this workspace; the new capabilities define the operator-facing contract over the existing distributed worker implementation.
## Impact
- Worker client/bootstrap/package entrypoints and generated Windows/Linux launchers.
- Scanner subprocess lifecycle, scan-slot ownership, cleanup, bundle staging, and local runtime state.
- Worker API progress, prebundle report, result-bundle, and assignment-status contracts.
- PostgreSQL reservation, scan, error, diagnostic, and observability queries/schema.
- Server-rendered worker administration pages and runtime deadline editor fields.
- Windows process supervision and Linux/WSL/Docker foreground/detached operation.
- Worker package tests, protocol tests, fault-injection tests, admin UI tests, and operator documentation.
@@ -0,0 +1,207 @@
# Remote Worker Operator Experience Findings
Recorded: 2026-09-23.
This file preserves the investigation that led to `add-worker-operator-experience` so implementation does not have to rediscover current behavior and terminology.
## Production Validation Facts
- Native Windows protocol-2 worker reached exact concurrency 3 and never exceeded it.
- Its bounded production cohort accepted 180 assignments: DockerHub 63, GitLab 65, HuggingFace 52.
- All 180 accepted results ingested, queue-settled, projected, and reconciled without global lineage violations or quarantine.
- A later dual-worker run proved Windows max 1, WSL max 1, combined max 2.
- Two WSL DockerHub assignments in separate runs stayed unresolved until fixed two-hour lease expiry and produced no accepted bundle.
- The server expired/refunded those reservations correctly; the missing information is where the worker spent the time before expiry.
Durable validation report:
`docs/worker-parallelism-validation-2026-09-23.md`
## Timeout and Lease Map
### Assignment deadline
Production explicitly used:
```yaml
supervisor:
worker_api:
assignment_ttl_seconds: 7200
```
- Production value: 7,200 seconds.
- Code/template fallback: 86,400 seconds.
- Managed validation bounds: 60 through 604,800 seconds.
- Validation requires the effective TTL to cover the maximum configured source scan timeout, bundle upload deadline, and 60 seconds of handoff margin.
- The server commits one immutable expiry at assignment issuance.
- Ordinary worker contacts and progress do not renew it.
- Relevant code: `app/runtime_document.py`, `app/worker_api.py`, `app/worker_assignment.py`, `app/scanner_db.py`.
### Docker target timeout
Canonical managed full-runtime and production value:
```yaml
sources:
dockerhub:
timeout: 600
```
- Managed legacy Windows/full runtime: 600 seconds.
- Current production profile: 600 seconds.
- Direct CLI/function fallback without managed config: 1,800 seconds.
- Legacy Windows Job containment terminates the owned TruffleHog process tree at the subprocess deadline.
- Historical evidence: `tests/test_scanner_queue_high_fixes.py`, `tests/test_validated_high_scanner_fixes.py`, and `openspec/changes/stabilize-dockerhub-trufflehog-lifecycle/`.
### Other relevant timers
- Worker API client socket timeout: 120 seconds.
- Result upload absolute body deadline: 1,800 seconds.
- Result upload idle timeout: 30 seconds.
- Remote assignment expiry reaper interval: 60 seconds.
- Empty-claim `Retry-After`: normally 5 seconds.
- Result-ingester/projector leases: 300 seconds after upload, unrelated to pre-upload scanning.
### Why ten minutes became two hours
The 600-second budget hard-preempts the TruffleHog subprocess path, but it is not one hard preemptive boundary around every surrounding Python operation. Permit acquisition, process startup/termination recovery, cleanup, filtering, serialization, bundle staging/fsync, and handoff can outlive the subprocess deadline. The remote client checks assignment expiry before synchronous execution and then has no phase heartbeat or cancellation loop.
The evidence does not prove which phase stalled. Increasing the assignment TTL would only make the unknown stall retain ownership longer. The implementation needs phase instrumentation and a complete supervised execution-unit deadline.
## Current Worker Experience
Current supported CLI flags in `app/remote_worker_client.py`:
```text
--server
--token
--parallelism
```
There are no worker `start`, `stop`, `status`, `attach`, `logs`, `history`, `doctor`, or JSON-output commands.
Current behavior:
- one daemon thread per slot;
- slot recovery state in `slot-N.json`;
- ready bundles retained until an authoritative receipt;
- scanner call is synchronous from the slot controller;
- state remains broadly `assigned` until bundle readiness;
- client loop failures print generic lines without a structured timeline;
- successful claim/scan/upload/receipt is mostly silent;
- scanner stdout/stderr is captured in temporary files and returned only after process completion;
- exact Git execution can emit no useful start line;
- concurrent messages can interleave;
- server `last_contact_at` does not prove or disprove active scanner progress.
Default paths:
- Windows: `%LOCALAPPDATA%/TRUF/RemoteWorker`.
- Linux: `$XDG_STATE_HOME/truf/remote-worker` plus `$XDG_DATA_HOME/truf/remote-worker`.
- Docker production convention: persistent `/data` volume with separate state/data roots.
## Legacy Attach Pattern
The old full-runtime supervisor has an `--attach` implementation in `app/supervisor.py` and a wrapper `attach_runtime.ps1`. The wrapper is currently disabled and the supervisor is not in the remote-worker package.
Useful semantics to reuse:
- exact background-instance identity;
- startup and loopback control handshake;
- initial status table;
- interactive `attach>` prompt;
- alternate-screen `watch` table;
- bounded log tail;
- `q`/EOF/Ctrl-C detach without worker shutdown;
- explicit coordinated shutdown command.
The worker needs a smaller implementation over its own event/status model, not a copy of the complete server supervisor.
## Current Error Model
The admin `Error category` column is rendered in `app/admin_api.py` from:
```sql
result_reservations.last_error_code AS error_category
```
It therefore describes assignment/transport errors such as remote prebundle failure or assignment expiry. It is not `errors.category`, scanner `error_class`, or `source_failure_category`.
Consequences:
- an accepted bundle is shown as assignment `completed` even when scan status is `error`;
- accepted scan errors usually leave assignment `Error category` blank;
- process versus storage prebundle failure survives in resolution JSON but is not shown;
- permanent provider skips can exist in metadata/warnings without an `errors` row;
- provider response bodies are inconsistently reduced or discarded.
Useful data already persisted but not presented together:
- `result_reservations`: resolution kind/JSON, receipt, issue/expiry/resolve timestamps, last error code/detail;
- `target_scans`: status, error count, skipped reason, first error summary, start/end/duration;
- `errors`: category, summary, raw selected error line;
- `scan_result_compat.metadata_json`: error class, retryability, source failure category, warnings, degraded/skipped flags, process return/timeout/output metadata;
- `keycheck_results`: provider status group/message/metadata.
## Diagnostic Direction
Use separate dimensions rather than one overloaded category:
```text
assignment outcome: accepted | prebundle_failed | expired | unfinished
scan outcome: clean | found | degraded | error | skipped | unavailable
phase: scanning | cleaning | bundling | uploading | ...
kind: provider_http | scanner_process | exception | storage | protocol | ...
category: authorization | rate_limit | timeout | network | scanner | ...
code: stable concrete identifier
retryable: true | false
```
Candidate transmitted limits from the investigation:
- provider body material: 16 KiB;
- process log head/tail: 32 KiB combined;
- one diagnostic envelope: 64 KiB;
- at most 32 diagnostics and 256 KiB total per assignment;
- prebundle envelope profile sized to fit the existing Worker API JSON limit.
The local worker archive can retain larger/full artifacts under configurable age and byte rotation. Every transport transformation must state original size, stored size, hash, encoding, and truncation state.
## Progress and Duration Direction
Minimum phase transitions needed to explain the two-hour event:
```text
assignment_received
scan_permit_acquired
runner_started
source_prepare_started/completed
scanner_started/exited
filtering_started/completed
cleanup_started/completed
bundle_started/ready
upload_started/acknowledged
```
Progress is evidence only and does not renew the immutable lease.
Duration percentiles:
- p50: median duration;
- p95: 95 percent of observations finish at or below this duration;
- p99: 99 percent finish at or below it.
Compute them separately by source, phase, outcome, and time window, with sample counts. They describe observed behavior and inform policy; they do not silently set policy.
## Product Decisions
- Build one final operator architecture rather than a temporary admin patch.
- Implement it in large vertical chunks that each remain part of the final system.
- Use a worker-specific supervisor and local event stream.
- Split scanner execution into a supervised per-assignment runner process so the complete scan stage can be hard-preempted.
- Keep immutable server assignment ownership and make progress non-renewing.
- Add global fallback plus per-source assignment TTL policy.
- Use one diagnostic envelope across local files, terminal reports, bundles, PostgreSQL, API, and UI.
- Preserve diagnostic fidelity and expose all explicit transformations.
- Separate assignment outcome, scan outcome, and diagnostics in the admin UI.
- Include from-zero operator documentation and fault injection in the same change.
@@ -0,0 +1,71 @@
## ADDED Requirements
### Requirement: Separate assignment and scan outcomes
The worker administration list SHALL display assignment transport outcome, scan outcome, and diagnostic count as separate fields and SHALL not label `last_error_code` as the complete scan error category.
#### Scenario: Accepted scan has an error outcome
- **WHEN** a reservation has a durable accepted bundle whose target scan status is `error`
- **THEN** the list SHALL show assignment `accepted`, scan `error`, and the diagnostic count/categories
#### Scenario: Assignment expires before scan ingestion
- **WHEN** a reservation expires without an accepted bundle
- **THEN** the list SHALL show assignment `expired`, scan outcome unavailable, and the expiry diagnostic
### Requirement: Assignment detail timeline
Each worker assignment row SHALL link to a detail page that reconstructs issued, phase-progress, bundle/terminal report, receipt, ingestion, queue settlement, and projection timestamps that exist for that assignment.
#### Scenario: Administrator opens an active assignment
- **WHEN** progress events exist for an unresolved assignment
- **THEN** the page SHALL show current phase, phase age, last progress age, scan deadline, assignment deadline, and ordered prior phases
#### Scenario: Administrator opens a settled assignment
- **WHEN** the assignment has been accepted and projected
- **THEN** the timeline SHALL distinguish acceptance, ingestion, queue settlement, and projection completion rather than collapsing them into one completion time
### Requirement: Clickable diagnostic detail
The assignment detail page SHALL list diagnostics and SHALL provide human summary, canonical envelope JSON, raw body view, process log view, transformation metadata, and copy/download actions for each diagnostic.
#### Scenario: Diagnostic body is complete
- **WHEN** an HTTP diagnostic contains an untruncated body
- **THEN** the raw-body view SHALL identify it as complete and display the captured content and metadata
#### Scenario: Diagnostic material is truncated
- **WHEN** body or log material was size-truncated
- **THEN** the view SHALL prominently display original/stored sizes, hash, and truncation state
#### Scenario: Legacy scan error has no diagnostic envelope
- **WHEN** an older scan has only existing `errors.raw_error` or result metadata
- **THEN** the detail page SHALL display those fields as legacy evidence and SHALL not invent a new envelope
### Requirement: Diagnostic filtering and grouping
The admin UI SHALL filter independently by source, time, worker/device, assignment outcome, scan outcome, phase, category, stable code, and retryability and SHALL group repeated diagnostic fingerprints without hiding individual occurrences.
#### Scenario: Administrator filters rate-limit errors
- **WHEN** category `rate_limit` and a time window are selected
- **THEN** results SHALL include matching diagnostics regardless of whether their assignments were accepted or prebundle-failed
#### Scenario: Repeated diagnostics are grouped
- **WHEN** multiple diagnostics share a fingerprint
- **THEN** the UI SHALL show aggregate count and affected assignments while retaining links to each occurrence
### Requirement: Worker fleet status
The admin UI SHALL display each worker's configured cap, active slots, current phases, package identity, latest contact/progress ages, pending local-recovery indication when reported, and known idle/backoff reason.
#### Scenario: Worker is scanning without recent API contact
- **WHEN** a worker has an active assignment and recent progress events but its authentication contact timestamp is old
- **THEN** fleet status SHALL show active progress rather than classifying the worker as idle solely from contact age
#### Scenario: Worker cannot claim due to capacity
- **WHEN** the server rejects claims because a pipeline capacity axis is closed
- **THEN** fleet status SHALL show the capacity reason instead of a generic offline/idle state
### Requirement: Deadline and duration administration
The runtime editor and worker observability pages SHALL explain effective scan, upload, and assignment deadlines and SHALL show p50/p95/p99 duration metrics with sample counts by source and phase.
#### Scenario: Administrator edits a source assignment deadline
- **WHEN** a per-source TTL candidate is previewed
- **THEN** the editor SHALL show the effective policy, validation relationship to scan/upload bounds, and that only future assignments are affected
#### Scenario: Administrator compares policy to observations
- **WHEN** sufficient phase-duration samples exist
- **THEN** the page SHALL show the configured deadline alongside source-specific percentile values without automatically changing configuration
@@ -0,0 +1,86 @@
## ADDED Requirements
### Requirement: Unified versioned diagnostic envelope
Worker scan errors, provider failures, process failures, local exceptions, timeouts, prebundle failures, and assignment expiry context SHALL use one versioned diagnostic envelope with independent phase, kind, category, stable code, summary, retryability, attempt, timestamps, and optional HTTP/process/exception material.
#### Scenario: Provider returns an HTTP error
- **WHEN** a provider operation receives an unsuccessful HTTP response
- **THEN** the diagnostic SHALL identify the phase, provider operation, HTTP status/content type, stable category/code, retryability, and captured response material
#### Scenario: Scanner process fails
- **WHEN** a scanner process exits unsuccessfully or is terminated at its deadline
- **THEN** the diagnostic SHALL identify its process result, timeout/signal state, phase, stable category/code, and captured log material
#### Scenario: Assignment expires without a worker result
- **WHEN** the server expires an unresolved assignment
- **THEN** it SHALL create or expose a diagnostic describing assignment expiry and the last accepted progress phase without claiming a scanner error occurred
### Requirement: Diagnostic fidelity and explicit transformation
Captured body and log bytes SHALL be preserved without silent semantic rewriting. Every size limit, encoding conversion, or truncation SHALL record original bytes, stored bytes, content hash, encoding, and truncation state.
#### Scenario: Text body fits the bound
- **WHEN** a captured provider response body fits the configured diagnostic body bound
- **THEN** the transmitted diagnostic SHALL contain the complete captured text and SHALL mark it untruncated
#### Scenario: Body exceeds the bound
- **WHEN** captured body bytes exceed the transmitted bound
- **THEN** the diagnostic SHALL carry the bounded material plus original/stored sizes, full captured-content hash when available, and `truncated=true`
#### Scenario: Body is not text
- **WHEN** captured diagnostic body bytes are not valid text in the declared encoding
- **THEN** the envelope SHALL use an explicit binary encoding representation and SHALL preserve the same transformation metadata
### Requirement: Bounded diagnostic transport
The protocol SHALL enforce deterministic per-body, per-log, per-envelope, diagnostic-count, and aggregate diagnostic bounds while rejecting envelopes whose declared and actual sizes disagree.
#### Scenario: Accepted bundle contains diagnostics
- **WHEN** a worker uploads a result bundle with diagnostic frames within all bounds
- **THEN** bundle acceptance and ingestion SHALL validate and persist each diagnostic idempotently with the scan
#### Scenario: Diagnostic aggregate exceeds its limit
- **WHEN** a bundle or terminal report exceeds a diagnostic count or byte limit
- **THEN** the API SHALL reject it with a stable protocol error and SHALL NOT partially persist diagnostics
### Requirement: Prebundle and accepted-result parity
The same diagnostic envelope SHALL be usable in prebundle terminal reports and accepted scan-result bundles, with only transport-size profiles differing.
#### Scenario: Worker storage fails before bundle creation
- **WHEN** the worker cannot create a result bundle
- **THEN** its terminal report SHALL include a diagnostic envelope rather than replacing the exception with one generic fixed detail string
#### Scenario: Scan returns errors in a valid bundle
- **WHEN** scanning completes with structured errors and a valid bundle
- **THEN** those errors SHALL be represented as diagnostics attached to the ingested target scan and SHALL remain distinct from assignment transport outcome
### Requirement: Deterministic diagnostic identity
Each diagnostic SHALL have a deterministic UID derived from its canonical identity and content so retries and replay cannot create duplicates.
#### Scenario: Accepted upload is replayed
- **WHEN** an identical accepted result bundle is uploaded again
- **THEN** the server SHALL return the durable receipt and SHALL NOT insert duplicate diagnostic rows
#### Scenario: Same code occurs twice in one assignment
- **WHEN** two distinct occurrences share category and code but differ in occurrence identity or content
- **THEN** both SHALL be retained as distinct diagnostics with stable UIDs
### Requirement: Local diagnostic archive
The worker SHALL retain a queryable local JSON diagnostic envelope and optional body/log artifacts per assignment, with configurable age/byte rotation and explicit artifact-availability state in history.
#### Scenario: Operator opens a local diagnostic
- **WHEN** `truf-worker history` or `logs` selects a retained diagnostic
- **THEN** the worker SHALL present the canonical envelope and exact paths/availability of its body and log artifacts
#### Scenario: Artifact rotates out
- **WHEN** a body or log artifact is removed by configured local rotation
- **THEN** terminal history SHALL remain and SHALL state that the artifact is no longer locally retained
### Requirement: Orthogonal error taxonomy
The diagnostic model SHALL keep phase, kind, broad category, stable code, retryability, assignment outcome, and scan outcome as separate dimensions.
#### Scenario: Accepted scan has provider errors
- **WHEN** a result bundle is durably accepted but the scan outcome is `error`
- **THEN** the assignment outcome SHALL remain `accepted`, scan outcome SHALL be `error`, and provider diagnostics SHALL retain their own categories/codes
#### Scenario: Assignment expires
- **WHEN** an assignment expires before bundle acceptance
- **THEN** assignment outcome SHALL be `expired`, scan outcome SHALL be unavailable, and the expiry diagnostic SHALL not be categorized as a provider scan failure
@@ -0,0 +1,74 @@
## ADDED Requirements
### Requirement: Unified worker lifecycle CLI
The worker package SHALL provide `install`, `run`, `start`, `stop`, `status`, `attach`, `logs`, `history`, and `doctor` commands with equivalent lifecycle semantics on supported Windows and Linux/WSL platforms.
#### Scenario: Operator starts a detached worker
- **WHEN** an installed operator invokes `truf-worker start`
- **THEN** the command SHALL launch the worker supervisor, wait for its startup handshake, and return the verified instance identity and current state
#### Scenario: Operator runs in the foreground
- **WHEN** an operator invokes `truf-worker run`
- **THEN** the same supervisor implementation SHALL run in the foreground and SHALL begin graceful drain on the first interrupt
### Requirement: Verified detached supervisor lifecycle
The supervisor SHALL publish a versioned instance record, status projection, local control endpoint, startup result, and shutdown receipt tied to the exact running process identity.
#### Scenario: Status finds a stale instance record
- **WHEN** the recorded process no longer matches the recorded executable, creation identity, or live control handshake
- **THEN** `status` SHALL report the instance as stale and SHALL NOT represent it as a running worker
#### Scenario: Graceful stop has active slots
- **WHEN** `stop` is requested while one or more slots own assignments
- **THEN** the supervisor SHALL stop new claims, display the draining slots, and wait for terminal local reconciliation up to the requested stop deadline
### Requirement: Attachable live operator view
The supervisor SHALL provide an `attach` session that renders current worker and per-slot state and follows new events without making attachment own the worker lifetime.
#### Scenario: Operator detaches
- **WHEN** the operator presses `q`, sends EOF, or interrupts the attach client
- **THEN** only the attach session SHALL end and the worker supervisor SHALL continue running
#### Scenario: Concurrent slots update
- **WHEN** multiple slots emit interleaved phase events
- **THEN** attach SHALL render one coherent row per slot and SHALL preserve event ordering by local sequence
### Requirement: Honest per-slot status
Status and attach SHALL display source, phase, phase elapsed time, scan deadline, assignment time remaining, last progress age, and only counters measured by the execution path. They SHALL NOT synthesize percentage completion without a reliable denominator.
#### Scenario: Long scanner execution
- **WHEN** a slot remains in `scanning` with a live runner process
- **THEN** status SHALL continue updating elapsed time, deadline remaining, and last-progress age rather than appearing frozen
#### Scenario: Worker has no assignment
- **WHEN** a slot is idle because of server backoff, cap, paused dispatch, capacity, or an empty queue
- **THEN** status SHALL report the known idle/backoff reason and next claim time when supplied by the server
### Requirement: Human and machine output contracts
Every non-interactive inspection command SHALL support versioned JSON output, and every follow command SHALL support versioned NDJSON output containing no human decoration.
#### Scenario: Automation requests status
- **WHEN** `truf-worker status --json` is invoked
- **THEN** stdout SHALL contain exactly one parseable versioned status object representing the same state as the human view
#### Scenario: Automation follows events
- **WHEN** `truf-worker logs --follow --json` is invoked
- **THEN** stdout SHALL contain one complete versioned event object per line in sequence order
### Requirement: Local history and operational diagnosis
The supervisor SHALL retain terminal assignment history, rotating worker logs, event history, diagnostic references, and local storage usage, and SHALL expose them through `history`, `logs`, and `doctor`.
#### Scenario: Operator investigates a completed assignment
- **WHEN** the operator requests history for a terminal reservation
- **THEN** the worker SHALL show its terminal local/receipt outcome, durations, phase timeline, and available diagnostic artifact references
#### Scenario: Operator runs doctor
- **WHEN** `truf-worker doctor` is invoked
- **THEN** it SHALL inspect package identity, singleton/process state, local state readability, disk usage, server reachability, and retained work without claiming an assignment
### Requirement: From-zero operator documentation
The release SHALL include one canonical guide from package acquisition through installation, first start, attach/status interpretation, graceful stop, recovery, update, diagnostics, and removal.
#### Scenario: New operator follows the guide
- **WHEN** an operator starts with a supported worker package and issued server enrollment data
- **THEN** the documented commands SHALL lead to a running verified worker and explain every state visible before the first assignment
@@ -0,0 +1,79 @@
## ADDED Requirements
### Requirement: Canonical assignment phase model
The worker SHALL represent assignment execution with one versioned phase/event model shared by local status, local history, server progress, and administrative views.
#### Scenario: Assignment completes normally
- **WHEN** a slot claims, executes, stages, uploads, and receives acceptance for an assignment
- **THEN** it SHALL emit monotonic phase events sufficient to reconstruct the time spent from `assigned` through `awaiting_receipt`
#### Scenario: Process restarts during an assignment
- **WHEN** a worker restarts with persisted slot state
- **THEN** recovered events SHALL continue from the persisted sequence and SHALL record recovery without rewriting the prior timeline
### Requirement: Complete scan-stage watchdog
The worker SHALL enforce one hard scan-stage deadline across permit acquisition, local preparation/resolution, data acquisition, scanner execution, filtering, cleanup, and result staging by supervising the complete execution unit outside the controller process.
#### Scenario: Scanner child exceeds the deadline
- **WHEN** the assignment runner remains active at the scan-stage deadline
- **THEN** the controller SHALL terminate its complete process tree, release/detach local resources, and produce a timeout result identifying the final phase
#### Scenario: Cleanup blocks after scanner exit
- **WHEN** scanner execution has ended but cleanup or staging remains blocked at the deadline
- **THEN** the same hard deadline SHALL terminate the runner and SHALL prevent the slot from remaining occupied until assignment expiry
#### Scenario: Permit acquisition consumes the budget
- **WHEN** no scan permit is acquired before the scan-stage deadline
- **THEN** the worker SHALL produce a phase-specific timeout result without starting the scanner
### Requirement: Non-renewing server progress
The Worker API SHALL accept idempotent monotonic progress events for the current reservation while preserving the original immutable assignment and queue deadlines.
#### Scenario: Progress is accepted
- **WHEN** the assigned device submits the next valid event sequence for its unresolved reservation
- **THEN** the server SHALL persist the event/latest phase and SHALL NOT alter assignment expiry or ownership
#### Scenario: Duplicate progress is retried
- **WHEN** an already accepted event sequence is submitted again
- **THEN** the server SHALL return the prior acceptance without creating a duplicate timeline event
#### Scenario: Progress cannot reach the server
- **WHEN** local phase transitions occur during a temporary connection failure
- **THEN** execution SHALL continue under the fixed deadline and events SHALL remain available locally for ordered retry
### Requirement: Observable deadline semantics
Assignments SHALL carry distinct effective target-scan, result-upload, and end-to-end assignment deadlines, and every operator/admin view SHALL label them by those meanings.
#### Scenario: Operator inspects active work
- **WHEN** status or admin renders an active reservation
- **THEN** it SHALL show the effective scan deadline, assignment deadline, time remaining, and current phase without conflating them
#### Scenario: Assignment expires
- **WHEN** the immutable assignment deadline passes without an accepted terminal result
- **THEN** expiry evidence SHALL include the last accepted phase and last-progress timestamp when available
### Requirement: Global and per-source assignment policy
The managed runtime configuration SHALL provide a global assignment TTL fallback and optional explicit overrides for GitLab, DockerHub, and HuggingFace, selected by the server at issuance.
#### Scenario: Source override exists
- **WHEN** a DockerHub assignment is issued and a DockerHub assignment TTL override is configured
- **THEN** its immutable expiry SHALL use the override and the assignment SHALL report that effective policy
#### Scenario: Source override is absent
- **WHEN** an assignment is issued for a source without an override
- **THEN** the global assignment TTL SHALL be used
#### Scenario: Invalid deadline policy is previewed
- **WHEN** an effective assignment deadline cannot cover its source scan timeout, upload deadline, and required handoff margin
- **THEN** managed configuration preview SHALL reject the candidate with a field-specific explanation
### Requirement: Phase duration percentiles
The server SHALL expose p50, p95, and p99 durations by source, phase, outcome, and selected time window, based only on completed observations appropriate to that metric.
#### Scenario: Administrator reviews DockerHub latency
- **WHEN** duration metrics are requested for DockerHub
- **THEN** the result SHALL separate end-to-end, scanning, cleanup, bundling, and upload percentiles and SHALL report sample counts
#### Scenario: Insufficient samples exist
- **WHEN** a percentile does not have the configured minimum sample count
- **THEN** the UI/API SHALL label it insufficient rather than presenting it as a stable policy recommendation
@@ -0,0 +1,44 @@
## 1. Canonical Contracts and Persistence
- [x] 1.1 Implement the versioned worker phase/event model, canonical phase transitions, monotonic sequence validation, JSON/NDJSON serialization, and contract tests shared by worker, API, and admin code.
- [x] 1.2 Implement the unified diagnostic envelope, orthogonal taxonomy, deterministic diagnostic UID, exact body/log representation, explicit truncation metadata, aggregate limits, and serialization/validation tests.
- [x] 1.3 Add PostgreSQL progress-event and worker-diagnostic persistence, indexes, idempotent writes, reservation/scan joins, migration coverage, and authoritative query methods.
- [x] 1.4 Add managed global/per-source assignment deadline policy, effective-value validation against scan/upload bounds, assignment serialization, and runtime-document/editor tests.
## 2. Complete Worker Supervisor
- [x] 2.1 Build the `truf-worker` command surface (`install`, `run`, `start`, `stop`, `status`, `attach`, `logs`, `history`, `doctor`) over one supervisor implementation, including the existing foreground invocation migration alias.
- [x] 2.2 Implement verified Windows and Linux/WSL supervisor instance lifecycle, startup handshake, local control endpoint, graceful drain/stop, shutdown receipt, stale-instance handling, and lifecycle tests.
- [x] 2.3 Implement append-only local events, rebuildable status projection, terminal history, rotating logs, per-assignment diagnostic/body/log files, retention accounting, and crash/restart recovery tests.
- [x] 2.4 Implement human status/attach views and versioned JSON/NDJSON modes with coherent concurrent-slot rendering, honest phase/deadline/progress fields, bounded follow/tail behavior, and command-level tests.
- [x] 2.5 Update Windows portable and Linux/Docker package entrypoints, manifests, launchers, and package self-tests so the supervisor is the supported runtime on every platform.
## 3. Assignment Runner and Full-Stage Watchdog
- [x] 3.1 Introduce the contained per-assignment runner process and controller protocol while preserving existing claim state, deterministic bundle authority, source capabilities, and receipt/recovery behavior.
- [x] 3.2 Instrument permit wait, preparation, source resolution, download/clone, scanner execution, filtering, cleanup, bundle staging, upload, and receipt transitions with the canonical phase events and measured durations.
- [x] 3.3 Enforce one hard scan-stage deadline across the runner process tree, produce a normal phase-specific timeout result, detach abandoned work to janitor ownership, and return the slot without waiting for assignment expiry.
- [x] 3.4 Implement controller and runner crash recovery for persisted assignments, incomplete runner outputs, ready bundles, stale results, lowered parallelism, and supervisor restart.
- [x] 3.5 Add deterministic fault-injection tests for blocking/failure in every phase, including permit starvation, child non-exit, cleanup stall, staging/fsync failure, upload retry, deadline crossing, and process restart.
## 4. Worker API and Result Pipeline
- [x] 4.1 Add the authenticated non-renewing progress endpoint with ownership checks, monotonic/idempotent sequencing, latest-phase projection, bounded retry behavior, and API/database tests.
- [x] 4.2 Add diagnostic frames to protocol-2 bundles and the same diagnostic envelope to prebundle terminal reports, including exact size accounting, deterministic replay, and protocol compatibility tests.
- [x] 4.3 Ingest diagnostics transactionally with target scans/errors and attach prebundle diagnostics to reservations, while preserving durable receipt, queue settlement, projection, and replay invariants.
- [x] 4.4 Extend assignment/status responses with effective deadlines, latest phase/progress, known idle/backoff reason, and diagnostic availability, and cover old-package records explicitly in compatibility tests.
## 5. Final Administration Experience
- [x] 5.1 Replace the worker-list query/view model with separate assignment outcome, scan outcome, diagnostic summary, active phase, phase/progress age, effective deadlines, slot/cap, and package fields.
- [x] 5.2 Build the assignment detail page with ordered phase/receipt/ingestion/settlement/projection timeline, duration breakdown, scan summary, diagnostic list, exact body/log views, canonical JSON copy/download, and explicit transformation metadata.
- [x] 5.3 Add independent filters and repeated-diagnostic grouping for source, worker/device, assignment outcome, scan outcome, phase, category, stable code, retryability, and time window without hiding individual occurrences.
- [x] 5.4 Add source/phase/outcome p50, p95, and p99 duration queries and admin views with sample counts, and present them beside effective scan/upload/assignment deadline policy in the runtime editor.
- [x] 5.5 Add end-to-end admin tests for active progress, accepted scan errors, prebundle failures, assignment expiry, legacy records, complete/truncated bodies, repeated fingerprints, and machine-readable detail output.
## 6. Operator Release and Production Proof
- [x] 6.1 Write and validate the canonical from-zero operator guide covering package acquisition, install, first run, start/status/attach/logs/history, phase/deadline interpretation, diagnostics, graceful stop/drain, recovery, update, and removal.
- [x] 6.2 Run the complete unit/integration/protocol/package test matrix and build reproducible Windows and Linux worker artifacts with registered manifests and documented identities.
- [x] 6.3 Perform bounded production validation on native Windows and WSL/Docker covering multi-slot progress, attach while active, a forced full-stage timeout, diagnostic body/log inspection, restart recovery, accepted/ingested/projected reconciliation, and final production restoration.
- [x] 6.4 Record final duration percentiles, watchdog evidence, diagnostic/admin screenshots or snapshots, operator command transcript, known limits, and rollout/rollback results in a durable dated report.