148 lines
20 KiB
Markdown
148 lines
20 KiB
Markdown
## Context
|
|
|
|
The current PostgreSQL pipeline already has single-target admission, capacity reservation, immutable Git/Docker plans, scan/error classification, canonical v2 `.trb` staging, transactional ingestion, projection, and detailed keychecks. The execution seam is the claim-to-`stage_claim` path in `app/console_runner.py`, not a replacement scheduler.
|
|
|
|
`scan_target_result()` in `app/scanner.py` dispatches existing source implementations. Bundle staging extracts candidates, including structured Postman evidence; ingestion persists findings/candidates and separately schedules projection and keycheck. JSONL projection is not a prerequisite for keycheck, and a TruffleHog `Verified` field is not a completed detailed keycheck.
|
|
|
|
Current execution is not automatically portable to a DB-free client: process launch authority, exact Git/Docker plans, reservation accounting, and producer recovery depend on the local runtime. Recovery can inspect local PID/executable identity and local files, which cannot establish remote worker liveness. The design adapts these boundaries while keeping scanner/provider/error behavior intact.
|
|
|
|
Development takes place in `D:\truf-workers`, a source-only snapshot of the current `D:\truf-docker` working tree, including uncommitted fixes. Production data and Git history were not copied. Existing unarchived specifications are references; `openspec/specs/` has no canonical baseline. Old three-global-permit/`S:` handoff assumptions apply to their historical local deployment, not this remote boundary. PostgreSQL remains authoritative despite older status-file accounting descriptions.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals:**
|
|
- Move expensive download/scan work to trusted Windows/Linux clients with minimum new code and persistent state.
|
|
- Preserve scanner results, attribution, exact-plan coverage, error classification, retry decisions, candidate routing, and detailed server-side keychecks.
|
|
- Keep source/provider/scanner configuration centralized; support client-selected N slots capped by the server across a user's devices.
|
|
- Recover through a fixed 24-hour default assignment deadline, durable retries, and existing identity fencing, without heartbeat traffic.
|
|
- Secure the small public surface and test locally against empty storage and synthetic inputs.
|
|
|
|
**Non-Goals:**
|
|
- Client detailed keychecks, a second broker/queue/result format, generic job infrastructure, batch claims, worker affinity, or worker blacklists.
|
|
- New scan retry counts, a three-attempt/dead-letter rule, exactly-once physical execution, or heartbeat/lease renewal protocols.
|
|
- Client disk encryption, mTLS, public/untrusted workers, auto-update, complex roles, multiple server replicas, or PostgreSQL container extraction.
|
|
- Production database import, production deployment, changes to the active runtime, or archiving unrelated OpenSpec changes.
|
|
|
|
## Decisions
|
|
|
|
### 1. Keep one runtime and add a thin remote boundary
|
|
|
|
```text
|
|
Internet -> Caddy :443
|
|
|-- /api/v1/worker/* [device token] -> Worker API
|
|
`-- /<random-admin>/* [login/password] -> micro-admin
|
|
|
|
Server runtime: PostgreSQL + existing queue/reservations
|
|
discovery/scheduling -> admission
|
|
receiver -> existing ingester -> projection
|
|
`-> detailed keycheck
|
|
|
|
Client slot: claim -> existing download/scan -> canonical .trb -> upload/ack
|
|
```
|
|
|
|
Run Worker API and micro-admin as runtime-managed processes, not independent database-owning services. Keep the current dashboard backend private; any dashboard information in the admin area shares its authentication boundary. PostgreSQL, supervisor control, Caddy's control API, and raw backend ports have no public host bindings.
|
|
|
|
Reuse `reserve_and_claim_target`, ambiguous-admission reconciliation, `scan_target_result`, bundle staging/validation, `mark_result_bundle_ready`, and ingester/projector/keycheck paths. Extract only the code needed to run one already-planned scan and build its complete bundle without PostgreSQL. Adapt DB-bound plan inputs and node-local process authority rather than giving a client a DSN or a supervisor credential. Preserve `OwnedProcess` containment and cleanup; the client must launch only the known scanner tools, not arbitrary server-supplied commands.
|
|
|
|
Alternative rejected: copying `run_cycle_v2` unchanged or implementing a second scanner/queue. Both either retain server authority on clients or create divergent behavior.
|
|
|
|
### 2. Central configuration, minimal device identity, compatible jobs
|
|
|
|
Client-authored operational configuration has only server URL, opaque device token, and desired positive slot count N. Generated local pending-work state and workspace paths are not independent source configuration. The server resolves source/scanner settings, immutable plan, limits, and only the credentials required for this task. It must not send the whole config, discovery credential pool, database credentials, admin secrets, or control authority.
|
|
|
|
Bind each device token to its owning user and a server-issued/stable worker identity; store token hashes and support revocation. Keep only the identity/quota metadata required for administration, bound to current reservations rather than creating `remote_jobs` or another scheduler. Server configuration supplies default limits; admin changes the per-user cap across all devices.
|
|
|
|
Send protocol/build/policy compatibility metadata with normal claim traffic, not a background registration/liveness service. Validate the pinned scanner/custom detector policy and supported source execution on the client's OS/architecture before consuming a target. Pass a config snapshot or stable effective-config identity so the result stays attributable even if the server config changes mid-task. A client rejects an incompatible job instead of silently using its own settings.
|
|
|
|
Alternative rejected: worker-owned provider settings/tokens or another config management system. Trusted clients receive the task-specific secrets they need over HTTPS; this is not a sandbox against a compromised client.
|
|
|
|
### 3. One claim per free slot with existing backpressure
|
|
|
|
Each free client slot requests one assignment. Effective admission is bounded by N locally, the atomic per-user server cap across devices, and existing source/global/output-capacity rules. A client advertising N=8 does not override an admin cap of 3. Concurrent claims from different devices must not overshoot a shared cap. Reducing a cap stops new admissions until usage falls below it; it does not invent cancellation semantics for already-issued work.
|
|
|
|
Keep the current node-local scan-permit mechanics. A client work slot holds its assignment through durable result acknowledgement, which bounds pending uploads when the server is unavailable. Release the native scan permit at its existing handoff point; do not conflate that permit with remote user quota. A durably accepted bundle, acknowledged pre-bundle terminal disposition, or completed expiry recovery releases remote assignment quota exactly once; server spool credits remain governed by existing ingestion accounting.
|
|
|
|
Persist the existing admission/request identity before the first claim request. Retry ambiguous requests with the same identity and reconcile the same reservation; do not allocate another task because a claim response was lost. Scope recovery to the authenticated device. Empty queues or denied capacity return a bounded polling delay, not a scan error or a new queue item.
|
|
|
|
Alternative rejected: batches and client backlogs. Per-free-slot claims reuse current scheduling and avoid another recovery structure.
|
|
|
|
### 4. Fixed expiry, no remote heartbeat
|
|
|
|
Set `expires_at` from the server clock at assignment commit, with a configurable duration defaulting to 24 hours. The deadline includes download, scan, and upload. Ordinary API contact, retries, and client restarts do not renew it. Internal scanner timeouts and server discovery leases keep their existing meanings; local slot bookkeeping is not a worker heartbeat.
|
|
|
|
Use one periodic runtime recovery pass for expired remote assignments, initially once per minute and configurable. Extend existing recovery ownership/state checks instead of inventing a remote liveness monitor. Reconcile reservations, target claims, output credits, and dependent exact-plan/blob leases together. Local producer-PID recovery must not release remote claims. No associated lease may silently expire earlier and allow a second owner while the remote assignment remains valid.
|
|
|
|
Requeue unfinished expired work using the existing pre-handoff infrastructure-loss/refund path where applicable, not a fabricated scanner/provider failure or a new retry limit. The same worker may claim the task again. New issuance has a distinct current reservation/attempt identity. Both periodic recovery and result acceptance use the same atomic ownership/deadline boundary.
|
|
|
|
Reject an unaccepted old result once its assignment expires or is superseded; it must not finish newer work, update plan coverage, or return somebody else's credit. A retry of an already accepted identical result still receives its original acknowledgement even after the deadline. Ready/accepted bundles awaiting ingestion are not unfinished worker assignments and must not be expired back into the queue.
|
|
|
|
Trade-off accepted: a crashed client can delay work for about a day, and a partitioned client can keep physically scanning after reissue. Fencing gives one authoritative acceptance, not exactly-once physical execution.
|
|
|
|
### 5. Reuse the full bundle and durable handoff
|
|
|
|
The API needs only claim/reconcile, upload/acknowledge, and the existing pre-bundle failure/release outcome where necessary. Exact URL naming is implementation detail under `/api/v1/worker/`; do not add a heartbeat endpoint, remote shell, separate keycheck jobs, or a general command API.
|
|
|
|
Pre-bundle terminal reports use the same issuance fencing and replay rules as results. Persist/retry their identity until acknowledged; a lost reply resolves to the original disposition without charging retries, updating counters, or releasing quota/credits again. A delayed report from an expired/superseded issuance cannot mutate a new issuance even when the same device owns both. Authoritative acknowledgement resolves that client slot exactly once.
|
|
|
|
Send canonical v2 `.trb` bytes containing the existing findings, detector identities, source/origin/context, errors, metadata, candidate evidence, and exact-plan identity. Preserve structured Postman candidate extraction before ingestion. Do not send only provider names or reconstruct a reduced result JSON. Keep detailed keychecks entirely server-side after ingestion through current candidate/dedup/cache/routing logic. Retained `runtime/keychecks` outputs keep their existing location and retention rules.
|
|
|
|
Upload a binary stream, not base64 JSON, into a bounded server-owned partial file. Enforce the existing bundle limit/capacity reservation (currently 192 MiB where configured), time limits, identity, version, and hash. Validate all paths/archive members with the existing codec; clients cannot choose arbitrary server paths. Publish durably using the existing same-filesystem atomic handoff and mark the exact reservation ready before acknowledging server custody.
|
|
|
|
Acknowledgement means the validated bundle and its ready/recovery state survive a server restart; it does not mean projection or detailed keycheck has finished. An identical retry returns the same receipt without double ingestion/accounting. A conflicting body for an accepted identity is rejected. Interrupted uploads never count as successful scans; full-bundle retry is sufficient for the MVP, with no resumable-upload protocol.
|
|
|
|
Keep the accepted identity, digest, and reconciliation outcome in existing authoritative records independently of the spool file. Ordinary ingestion and bundle cleanup must not erase the information needed to recover a receipt: accept-with-lost-reply, ingest, clean up, restart, and retry after the deadline must still return the original acceptance for identical bytes and reject conflicting bytes. No second receipt queue or result store is required.
|
|
|
|
Keep the pending bundle and its assignment identity locally until durable acceptance is confirmed. After acknowledgement, existing cleanup may delete local task artifacts. A definitive fenced rejection transitions local work to an explicit stale/discard outcome with bounded cleanup, not an infinite retry or a false success. Scan errors still use the existing disposition path; HTTP retry/backoff does not consume scanner/provider retry budgets.
|
|
|
|
Crash recovery must cover publication-before-ready and ready-before-response windows using exact current ownership and existing spool reconciliation. Never infer success from the mere presence of an unvalidated file or expire already accepted work because a producer PID is absent.
|
|
|
|
### 6. Restricted edge and separate admin login bans
|
|
|
|
Use direct DNS to Caddy for the initial deployment and HTTPS for all worker/admin transport. The public admin prefix contains at least 128 bits of randomness, but login/password remains mandatory for every administrative route and asset. Caddy password authentication with a supported password hash is the minimal starting point; no plaintext password configuration or public signup. Protect typed mutating actions against CSRF with validated Origin/CSRF handling. No generic supervisor-command passthrough.
|
|
|
|
Only authenticated admin pages show operational summaries or typed user/device-token/quota and existing queue controls. Keep the standalone `/dashboard` route closed. Unknown paths return 404. Add no-store/same-origin-referrer, a restrictive compatible CSP, and production HSTS; avoid third-party assets and secrets/raw findings in diagnostics.
|
|
|
|
Run fail2ban on the host, reading redacted Caddy authentication-failure events. Two actual bad-credential submissions to the admin boundary within ten minutes produce a 24-hour IP ban. Do not count an ordinary initial Basic-auth challenge without credentials, unrelated 404s, or Worker API authentication failures. With shared HTTPS ingress, the ban action must update an ADMIN-ONLY Caddy IP deny matcher, with validated configuration reload, not globally drop that address on port 443. Persist ban state and provide SSH unban/recovery.
|
|
|
|
Use the direct connection address as client IP. Ignore arbitrary forwarded IP headers; a future trusted proxy requires explicit trust configuration and new tests. If network-level fail2ban actions are ever chosen instead, prove Docker forwarding-chain enforcement and preserve worker availability; a blanket host INPUT rule does not satisfy the contract.
|
|
|
|
Alternative rejected: a secret URL as sole protection, separate dashboard exposure, or a shared admin/worker IP jail. No extra auth service, mTLS, Cloudflare dependency, or client-disk encryption is required.
|
|
|
|
### 7. Observability from existing state
|
|
|
|
Correlate source/target, user/worker, reservation/issuance, issue/finish time, duration, outcome, and safe error category. Expose per-worker unfinished, completed, failed, and expired counts and last authenticated API contact. Distinguish bundle acceptance from later ingestion/keycheck completion and define counters from existing authoritative transitions so duplicate uploads do not inflate them.
|
|
|
|
Record issuance, acceptance, expiry/requeue, duplicate upload, and stale-result rejection events using existing logs/records. Never label a worker online/offline merely from silence: there is no heartbeat. Do not introduce a telemetry database or monitoring stack. Raw findings belong in the existing protected result storage, not transport/auth/debug logs.
|
|
|
|
### 8. Empty and isolated local validation first
|
|
|
|
Keep production checkout, processes, Docker containers, and volumes untouched. Reuse the existing stdlib checks, reviewed unit harness, and synthetic E2E driver. Before any runtime test, make test project names, image tags, volumes, ports, and config independent of inherited deployment defaults. The copied `compose.yaml`/legacy import overrides are not safe test launch instructions.
|
|
|
|
Initialize PostgreSQL empty in new test-owned storage; never restore a production dump or mount production PGDATA. Scrub inherited DSNs, provider secrets, runtime overrides, and proxies before application imports. Disable live discovery/keycheck autostart; explicitly feed synthetic tasks and mock detailed provider transports. No production credentials, paid calls, discovered public targets, or uncontrolled TruffleHog verification.
|
|
|
|
Exercise existing local and new remote scan execution on the same fixtures and compare normalized bundle/DB results while ignoring only transport identity/timing. Include Git/Docker immutable-plan and Postman-candidate parity, both custom detector directions, and all current source error dispositions. Test Windows and Linux clients, N>1 and multi-device caps, fixed-clock expiry/reissue races, lost claim/result replies, restarts, malformed/conflicting bundles, and auth isolation. Use controlled clocks instead of waiting a real day.
|
|
|
|
Use an isolated local TLS/internal network for Caddy/API tests and deny external egress. Validate admin bans through the actual edge, including a worker sharing the banned admin IP. Package a pinned Windows portable client and Linux client/container using current tool dependencies; no installer service or updater is necessary. Do not claim cross-platform readiness from mocked scans alone.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
- [Long assignment lifetime] Slow recovery and duplicate physical work are deliberate trade-offs; fencing and idempotent acceptance protect authoritative state.
|
|
- [DB-bound execution seams] Git/Docker plans and process authority require adaptation; parity tests are a release gate, not permission to rewrite provider logic.
|
|
- [Trusted client access] Clients can see task credentials and findings; issue least-needed credentials, redact logs, revoke device tokens, and use HTTPS.
|
|
- [Shared NAT and aggressive admin bans] Two mistakes can lock out an operator; scope bans to admin and retain tested SSH recovery.
|
|
- [Existing active OpenSpec deltas] Avoid mixing unrelated changes or reviving superseded local-only assumptions; archive/sync is a separate request.
|
|
- [Unsafe inherited launch defaults] Static checks can run now; runtime/E2E execution waits for explicit test isolation and image/config provenance checks.
|
|
|
|
## Migration Plan
|
|
|
|
1. Prepare the separate source-only Git workspace and complete this plan; do not start runtime services.
|
|
2. Establish neutral isolated tests, then implement the narrow execution/admission/transport boundary against empty PostgreSQL with synthetic fixtures.
|
|
3. Run local parity, failure/recovery, cross-platform, and edge-auth gates before enabling any real remote claims.
|
|
4. Keep remote admission disabled by default until explicitly configured. Preserve the local execution path using shared scanner logic; do not require a remote worker for existing local operation.
|
|
5. Production deployment and any additive identity/reservation schema migration need a separate reviewed rollout. No production data migration or rebuild is performed during this planning task.
|
|
6. To roll back a later rollout, stop new remote claims, retain/reconcile accepted bundles, and drain or expire outstanding assignments through current recovery before returning to local-only execution. Do not blindly drop reservation metadata or switch binaries underneath unfinished remote work.
|
|
|
|
## Open Questions
|
|
|
|
No blocking product decisions remain. Implementation must verify the smallest DB-free Git/Docker execution seam, the exact existing reservation fields to extend, and the host-specific fail2ban/Caddy reload mechanism through tests before release. Deployment-specific hostname, generated admin prefix/password, device tokens, and quota values are supplied during provisioning, not embedded in this plan.
|