Initial server source import
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-02
|
||||
@@ -0,0 +1,352 @@
|
||||
## Context
|
||||
|
||||
Docker scanning currently has two materially different paths. The full-image TruffleHog source is
|
||||
authoritative for new and normally completing images, but it treats an immutable image as one
|
||||
indivisible operation. The bounded layer path can resume and globally reuse content digests, but its
|
||||
fixed highest-eight-layer selector was admitted only for prior full-image timeouts. In the completed
|
||||
control cohort it retained 21.1% of routed identities and 5.8% of detector identities, so enabling
|
||||
that selector broadly would trade too much useful coverage for speed.
|
||||
|
||||
The layer path already provides the expensive safety primitives this change needs: exact manifest
|
||||
resolution, authenticated bounded Registry transfer, digest verification, private artifacts,
|
||||
contained TruffleHog filesystem execution, canonical reservation-bound plans, policy-scoped global
|
||||
blob leases, fenced result ingestion, and explicit per-image coverage. The missing pieces are a
|
||||
content-aware selector, a reusable execution-policy identity that is independent of selector
|
||||
budgets, and trustworthy evidence for deciding whether the adaptive path is safe to broaden.
|
||||
|
||||
Docker/OCI configuration contains an ordered `history` list that can help identify `COPY`, `ADD`,
|
||||
application setup, package installation, generic `RUN`, and likely bulk-data layers. That history is
|
||||
untrusted and may itself contain secret material. It is therefore only a bounded selection hint; it
|
||||
must never become an authority for content identity or successful coverage and its raw commands
|
||||
must never be persisted or logged.
|
||||
|
||||
The production database already contains durable version-one layer plans and covered blob rows.
|
||||
Changing their interpretation in place would invalidate audit and lease fences. The migration must
|
||||
be additive, retain version-one validation, and allow only explicitly proven compatible successful
|
||||
coverage to enter the new execution namespace.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Scan all supported unique content for images that fit conservative bounds.
|
||||
- Prioritize likely application, configuration, and source-bearing layers when a large image cannot
|
||||
fit those bounds.
|
||||
- Reuse successful immutable blob coverage across images and selector revisions when execution
|
||||
semantics are unchanged.
|
||||
- Freeze the first deterministic selection for an image and selector policy across every checkpoint,
|
||||
retry, and execution-policy transition.
|
||||
- Reduce per-image checkpoint overhead with bounded multi-blob leases without weakening per-blob
|
||||
execution and ingestion fences.
|
||||
- Preserve exact selected, reused, skipped, failed, and partial coverage semantics.
|
||||
- Produce private, aggregate, non-authoritative shadow evidence against 50-100 completed full-image
|
||||
controls before any broad adaptive rollout.
|
||||
- Retain full-image scanning and the existing timeout-only layer canary as immediate rollback paths.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Infer arbitrary file contents without downloading a compressed layer.
|
||||
- Claim complete coverage for an image with unsupported, oversized, failed, or budget-excluded
|
||||
descriptors.
|
||||
- Persist raw Docker config history, shadow findings, provider keys, target names, or Registry bearer
|
||||
tokens in rollout evidence.
|
||||
- Change detector classification, keycheck routing, Docker account ownership, repository resolver
|
||||
scheduling, guaranteed scan-slot capacity, worker count, or non-Docker scanners.
|
||||
- Reconstruct a merged container filesystem or remove historical whiteout content from individual
|
||||
layer evidence.
|
||||
- Destructively rewrite or delete version-one plans and coverage rows.
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. Introduce a version-two immutable content plan
|
||||
|
||||
Adaptive execution uses a canonical version-two plan. It retains the version-one image, repository,
|
||||
manifest, platform, media, limits, scan-policy, descriptor, reservation, and plan-hash fences and
|
||||
adds:
|
||||
|
||||
- a versioned selector algorithm and `selection_policy_sha256`;
|
||||
- an `execution_policy_sha256` for reusable successful blob evidence;
|
||||
- one bounded classifier value per descriptor;
|
||||
- exact selection and omission reasons; and
|
||||
- bounded checkpoint lease limits that do not alter the frozen selection set.
|
||||
|
||||
Version-one plan and execution validators remain exact because their JSON is already durable.
|
||||
Version-two validation dispatches by the exact integer version and rejects unknown fields, unknown
|
||||
classes, malformed hashes, descriptor reordering, and oversized canonical JSON. A reservation can
|
||||
bind one canonical plan only; idempotent replay must be byte-identical.
|
||||
|
||||
After exact manifest resolution and before plan binding, the scanner fetches the configuration blob
|
||||
through the existing bounded, authenticated, digest-verifying Registry path. It parses the JSON in
|
||||
bounded private memory/storage and maps non-`empty_layer` history entries from base to top onto the
|
||||
ordered manifest layers. The optional `rootfs.diff_ids` count and the non-empty history count must
|
||||
agree with the layer count. A missing field, malformed value, excess entry count, or alignment
|
||||
mismatch classifies every layer as `unknown`; it does not change descriptor identity or fail a
|
||||
valid immutable manifest.
|
||||
|
||||
The classifier uses a small versioned allow-list and emits only these bounded classes:
|
||||
`config`, `copy_add`, `app_config_run`, `package_run`, `other_run`, `bulk_data`, and `unknown`.
|
||||
Raw `created_by` values are discarded before plan construction and are excluded from logs, result
|
||||
metadata, errors, and database rows.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- Persisting normalized command text would retain unnecessary secret-bearing input.
|
||||
- Treating history as authoritative would let malformed or adversarial metadata hide content.
|
||||
- Mutating version-one plans would break exact replay and auditability.
|
||||
|
||||
### 2. Separate scan, execution, and selection identities
|
||||
|
||||
Three hashes have distinct responsibilities:
|
||||
|
||||
- `scan_policy_sha256` remains the installed scanner/detector/config fingerprint.
|
||||
- `execution_policy_sha256` hashes the scan policy plus versioned content validation and all archive
|
||||
semantics that can change which bytes a successful command examines. It excludes image/layer
|
||||
selection budgets, classifier weights, retry counts, lease duration, checkpoint size, and delay.
|
||||
- `selection_policy_sha256` hashes the selector version, class ordering, deterministic tie-breaks,
|
||||
supported descriptor classes, and all limits that determine the initial selected set. It excludes
|
||||
mutable global coverage and execution scheduling.
|
||||
|
||||
For version-two rows, the existing `docker_content_blobs.coverage_policy_sha256` key stores the
|
||||
execution-policy hash. Selector changes therefore do not force an identical successfully scanned
|
||||
digest through TruffleHog again, while archive or detector semantic changes still create a separate
|
||||
coverage namespace.
|
||||
|
||||
`docker_image_blob_coverage` receives a non-null `selection_policy_sha256`. Existing rows are
|
||||
backfilled from the exact bound plan on their linked reservation and indexed by
|
||||
`(queue_id, manifest_digest, selection_policy_sha256, position, reservation_id)`. The earliest full
|
||||
position map under that key is the immutable selection baseline. Coverage-policy changes may require
|
||||
new execution but cannot expand or contract that baseline silently.
|
||||
|
||||
Successfully covered version-one evidence may be copied lazily into the version-two execution
|
||||
namespace only in the same transaction that validates all of the following:
|
||||
|
||||
- the old row is durably `covered`, not pending, leased, submitted, failed, or ambiguous;
|
||||
- a linked covered image row and reservation contain an exact valid version-one plan;
|
||||
- the descriptor digest, kind, declared bytes, and semantic media class match;
|
||||
- the old coverage key is exactly the legacy hash derived from that plan; and
|
||||
- the scan fingerprint and every scan-affecting archive semantic equal the requested version-two
|
||||
execution policy.
|
||||
|
||||
The alias keeps the original successful reservation, plan, byte count, and completion provenance.
|
||||
Legacy rows are never rekeyed or deleted. If any compatibility proof is absent, the new namespace
|
||||
starts uncovered and normal fenced execution is required.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- Keeping selector limits in the coverage hash defeats global reuse whenever budgets are tuned.
|
||||
- Reusing every covered digest across scanner versions can suppress required rescans.
|
||||
- Bulk-rekeying legacy rows destroys provenance and races active leases.
|
||||
|
||||
### 3. Select all-fit images and rank large-image payload deterministically
|
||||
|
||||
Configuration remains independently eligible under its hard configuration-byte bound. Supported
|
||||
unique layer digests are evaluated under the hard per-layer bound. Covered digests in the matching
|
||||
execution namespace and duplicate positions in the same image are selected at zero new transfer
|
||||
bytes and zero new execution count.
|
||||
|
||||
If every supported unique descriptor fits the configured aggregate bytes and unique-layer count,
|
||||
the selector selects all of them regardless of history class. This is the complete bounded path for
|
||||
small images and avoids reducing their coverage merely because history hints are absent.
|
||||
|
||||
When the complete set does not fit, new unique layer candidates are sorted by:
|
||||
|
||||
1. class priority: `copy_add`, `app_config_run`, `package_run`, `unknown`, `other_run`, `bulk_data`;
|
||||
2. highest manifest position first;
|
||||
3. smallest compressed descriptor first; and
|
||||
4. lexical digest as the final stable tie-break.
|
||||
|
||||
The selector greedily admits candidates while both aggregate compressed-byte and unique-layer-count
|
||||
limits permit them. Hard per-descriptor limits are never exceeded. Each descriptor records one exact
|
||||
reason, including selected class, `already_covered`, `duplicate_digest`, `unsupported_media_type`,
|
||||
`config_too_large`, `layer_too_large`, `image_budget_exhausted`, or `layer_limit_exhausted`.
|
||||
Changing any class order, classifier rule, supported-media rule, or selection bound changes the
|
||||
selector hash.
|
||||
|
||||
The selection algorithm receives a transactionally consistent coverage snapshot, but mutable
|
||||
coverage is not part of its identity. The first complete descriptor-position map is written before
|
||||
any new lease and reused exactly on later checkpoints. Consequently a layer skipped by the original
|
||||
budget never becomes newly selected merely because an earlier selected layer became globally
|
||||
covered.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- Fixed highest-first selection has already failed the completed-control recall gate.
|
||||
- A whole-image byte cutoff loses small application layers above giant data layers.
|
||||
- Selecting only recognized commands lets missing or unusual history hide useful payload.
|
||||
- Selecting globally covered content only when it still fits the current budget wastes verified
|
||||
immutable evidence.
|
||||
|
||||
### 4. Lease bounded multi-blob checkpoints
|
||||
|
||||
The current executor and ingestion format already support more than one leased descriptor, but the
|
||||
binder leases one new digest and then defers the parent for 60 seconds. Adaptive execution leases a
|
||||
deterministic bounded batch from the frozen selected set under configurable maximum blob count and
|
||||
compressed bytes. A first eligible blob larger than the checkpoint-byte target but within its hard
|
||||
per-layer bound may be leased alone so it cannot starve indefinitely.
|
||||
|
||||
Every digest still has its own advisory lock, lease token, attempt count, execution record, digest
|
||||
verification, and final state. The executor processes the batch sequentially inside the same owned
|
||||
slot and bundle. Ingestion may cover successful earlier blobs while returning a later retryable blob
|
||||
to pending. A crash before durable handoff covers none of the un-ingested batch and normal exact
|
||||
lease expiry/recovery applies.
|
||||
|
||||
Checkpoint count, byte target, retry count, lease duration, and continuation delay are scheduling
|
||||
controls. They do not enter execution or selector hashes because they cannot turn an incomplete blob
|
||||
into successful coverage or change the frozen selected set.
|
||||
|
||||
### 5. Keep image coverage explicit and policy-specific
|
||||
|
||||
An image is complete only when its configuration and every manifest layer position are successfully
|
||||
covered under the requested execution policy. Reused and duplicate digests count as covered only
|
||||
after exact policy-compatible evidence exists. Any unsupported, oversized, budget-excluded, failed,
|
||||
or otherwise unselected descriptor makes the image bounded partial coverage.
|
||||
|
||||
Selected retryable work keeps the parent deferred. Shared active work does not consume another blob
|
||||
attempt. Exhausted selected work produces terminal incomplete disposition. Findings from completed
|
||||
selected blobs retain image, digest, kind, class, and position provenance and use the existing
|
||||
authoritative ingestion, projection, and keycheck paths.
|
||||
|
||||
Config history classification affects only selection order. It never changes detector output,
|
||||
finding authority, digest identity, or completion criteria.
|
||||
|
||||
### 6. Add adaptive modes without changing legacy rollout semantics
|
||||
|
||||
Existing `full`, timeout-only `canary`, and legacy `layer` meanings remain available for durable
|
||||
version-one work and rollback. Two explicit version-two modes are added:
|
||||
|
||||
- `adaptive-canary` assigns a configured basis-point cohort across all immutable Docker manifests by
|
||||
a stable versioned hash; cohort members use adaptive plans and non-members use full-image scanning.
|
||||
- `adaptive` uses adaptive plans for every eligible immutable Docker claim.
|
||||
|
||||
Neither mode depends on a previous full-image timeout. Retry and checkpoint attempts for the same
|
||||
manifest and selector retain the same assignment. Invalid mode, policy, migration, or gate state
|
||||
fails closed to full-image execution before any adaptive plan is bound. Returning configuration to
|
||||
`full` changes only new claims and leaves adaptive plans and audit rows intact.
|
||||
|
||||
The existing timeout-only canary remains independent and may continue while adaptive shadow evidence
|
||||
is gathered. The broad legacy `layer` mode remains operationally disabled because its selector did
|
||||
not pass recall gates.
|
||||
|
||||
### 7. Gate rollout with non-authoritative aggregate shadow evidence
|
||||
|
||||
An operator-invoked bounded shadow evaluator selects 50-100 exact immutable images whose authoritative
|
||||
full-image scans completed successfully under one scan fingerprint. It executes the candidate
|
||||
adaptive policy using the same downloader, validators, process containment, deadlines, and scanner
|
||||
fingerprint, but under a shadow authority that cannot call normal result ingestion or mutate target
|
||||
status, result reservations, global blob coverage, findings, keycheck candidates, projections, or
|
||||
source counters.
|
||||
|
||||
For each paired control, full routed identities are read from the protected database as
|
||||
`(service, provider_key_hash)` and adaptive routed identities are derived in private memory through
|
||||
the same candidate normalization. Detector identities use `detector_secret_hash`. Identity sets,
|
||||
raw findings, commands, provider material, image names, and bearer tokens are discarded after
|
||||
intersection counts are computed.
|
||||
|
||||
The durable report contains only policy hashes, cohort and completion counts, aggregate full,
|
||||
adaptive, and intersection counts, aggregate slot milliseconds, bounded failure counts, threshold
|
||||
results, and timestamps. It records no per-image row or identity. Slot timing uses the same outer
|
||||
monotonic boundary from admitted work through durable shadow sink completion for both paths; scan
|
||||
subprocess duration remains a diagnostic, not the gate denominator.
|
||||
|
||||
A report passes only when:
|
||||
|
||||
- 50-100 controls completed both paths without integrity, containment, or fence failure;
|
||||
- aggregate routed-identity recall, `intersection / full`, is at least 85%;
|
||||
- aggregate adaptive/full slot-time ratio is at most 40%;
|
||||
- every adaptive omission is represented in coverage counts; and
|
||||
- no credential persistence, quarantine, projection, source-failure, or resource-bound regression
|
||||
is observed.
|
||||
|
||||
Reports are bound to exact selector, execution, and scan policy hashes. Stale or incomplete reports
|
||||
cannot authorize another policy. Adaptive canary remains fail-closed until a matching report passes.
|
||||
Broad `adaptive` enablement additionally requires a stable low-percentage production canary over at
|
||||
least one repository-refresh interval. Operators change rollout configuration explicitly; shadow
|
||||
evidence never changes execution mode by itself.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- Routing shadow candidates through keycheck would make the experiment authoritative and consume
|
||||
external capacity.
|
||||
- Persisting per-image shadow identities creates unnecessary sensitive correlation data.
|
||||
- Comparing only detector counts does not measure the routed identities the scanner is intended to
|
||||
produce.
|
||||
- Automatically enabling adaptive mode from a report removes the operational rollback checkpoint.
|
||||
|
||||
### 8. Preserve incomplete warning semantics without redundant retries
|
||||
|
||||
The first completed 50-control production shadow report failed closed. Routed recall was 12 of 18
|
||||
identities (66.7%), adaptive/full slot time was 46.6%, and the report recorded 43 aggregate
|
||||
failures. Its selection evidence showed 230 descriptors omitted by the eight-layer limit and 14
|
||||
oversized descriptors, so selector recall remains the primary rollout blocker.
|
||||
|
||||
The same evidence exposed a separate execution defect. TruffleHog diagnostics such as
|
||||
`chunk_processing` and `detector_timeout` are explicitly deterministic, non-retryable warnings.
|
||||
Their findings must be retained, but the affected blob cannot establish complete coverage. The
|
||||
diagnostic adapter previously discarded the non-retryable bit, causing the layer executor to
|
||||
download and scan the same incomplete blob up to three times before reaching the same terminal
|
||||
state. The adapter now preserves aggregate warning retryability and the layer executor terminates
|
||||
that blob after the first deterministic warning. It does not mark the blob covered or remove the
|
||||
report failure.
|
||||
|
||||
The historical report schema retained only a total failure count, so its 43 failures cannot be
|
||||
decomposed exactly after the fact. Future shadow runs keep a fixed allow-list of aggregate-only
|
||||
failure categories in protected memory and print their totals in the final operator summary without
|
||||
changing report authority or persisting target-level evidence.
|
||||
|
||||
Shadow execution also suppresses target labels in finding-filter logs. Production scans retain their
|
||||
existing target logging, while both private full and private layer paths emit only aggregate filter
|
||||
counts. This closes a privacy gap found in the first report log without weakening normal operational
|
||||
diagnostics.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [History is malformed, misleading, or secret-bearing] -> Bound and validate it, persist only an
|
||||
enum, fall back to `unknown`, and keep identity/coverage independent of classification.
|
||||
- [Application secrets exist in a low-priority or giant layer] -> Scan all-fit images, keep unknown
|
||||
ahead of generic/bulk classes, record partial scope, retain full controls, and enforce the 85%
|
||||
routed-recall gate.
|
||||
- [Unsafe policy reuse suppresses a required rescan] -> Separate hashes and permit legacy aliasing
|
||||
only from exact successful compatible evidence in one fenced transaction.
|
||||
- [Selector changes expand coverage during retry] -> Freeze the earliest full position map under the
|
||||
selector hash before leases are issued.
|
||||
- [Multi-blob checkpoints increase work lost on crash] -> Bound count/bytes and retain independent
|
||||
per-blob leases and ingestion records; no pre-handoff result becomes covered.
|
||||
- [Config prefetch adds Registry traffic] -> Reuse the already bounded authenticated downloader and
|
||||
avoid a second fetch when the config is leased in the same plan.
|
||||
- [Individual layer scans expose whiteouted historical files] -> Preserve position provenance and
|
||||
describe evidence as image-content coverage, not merged-root state.
|
||||
- [Shadow evaluation consumes slots] -> Keep it operator-invoked, bounded, deterministic, and
|
||||
subject to the existing slot/resource controls.
|
||||
- [Aggregate reports hide individual anomalies] -> Fail the whole report on incomplete paired work
|
||||
and retain bounded failure counts without persisting target identity.
|
||||
- [Deterministic scanner warnings consume repeated transfer and slot time] -> Preserve their
|
||||
non-retryable policy, keep findings and incomplete coverage, and terminate the blob on its first
|
||||
bounded attempt.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Ship version-one and version-two validators, new modes, and migration code while production stays
|
||||
on its existing timeout-only canary.
|
||||
2. Stop authoritative runtime and verify no active result, queue, or blob leases remain.
|
||||
3. Add the selector-policy coverage column, aggregate shadow-report state, required timing/fence
|
||||
columns, indexes, and a new migration marker.
|
||||
4. Backfill every historical image-coverage row from its exact valid reservation plan. Abort and
|
||||
roll back the migration transaction on an orphan, malformed plan, or invalid hash; then make the
|
||||
column non-null and run exact schema validation.
|
||||
5. Restart with unchanged mode and verify legacy claims, projection, keycheck, resolver, quarantine,
|
||||
and source health before creating version-two work.
|
||||
6. Run the private shadow evaluator for 50-100 completed controls. Keep adaptive modes fail-closed if
|
||||
the matching report misses recall, timing, safety, or completion gates.
|
||||
7. Enable a low deterministic `adaptive-canary`, monitor at least one repository-refresh interval,
|
||||
and compare source failures, coverage reasons, routed yield, slot time, and quarantine.
|
||||
8. Increase canary basis points and finally enable `adaptive` only after every gate remains satisfied.
|
||||
9. Roll back immediately by setting mode to `full` or the existing timeout-only `canary`. Keep all
|
||||
version-two plans, reports, and coverage rows for audit and exact future resume.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Which initial bounded checkpoint count and byte target provide the best reduction in continuation
|
||||
delay without increasing crash rework materially?
|
||||
- Which classifier allow-list revisions improve routed recall in the first 50-100 controls? Every
|
||||
revision will receive a new selector-policy hash rather than changing an existing policy.
|
||||
- What adaptive-canary basis-point sequence should operators use after the shadow gate passes?
|
||||
@@ -0,0 +1,29 @@
|
||||
## Why
|
||||
|
||||
The timeout-only Docker layer fallback cuts heavy-image work to roughly 5-6% of full-image time, but its fixed highest-eight-layer policy retained only 21.1% of routed identities on completed controls. Broad Docker throughput therefore remains limited by indivisible full-image scans, while enabling the existing bounded layer selector globally would lose too much useful coverage.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a deterministic adaptive Docker payload policy that scans all eligible content for small images and prioritizes application-bearing content for large images using bounded, untrusted config-history hints.
|
||||
- Reuse successfully covered immutable blobs across images under a scan-execution policy independent of selector budgets, without weakening digest, reservation, lease, plan, or ingestion fences.
|
||||
- Give every adaptive plan a versioned selector identity and freeze its first selection across checkpoints and retries.
|
||||
- Record explicit reasons and bytes for every selected, reused, unsupported, oversized, and budget-excluded descriptor; adaptive completion remains partial whenever any content is omitted.
|
||||
- Add non-authoritative shadow evaluation and aggregate privacy-preserving evidence for completed full-image controls.
|
||||
- Keep full-image scanning as the default and rollback path; permit broad adaptive rollout only after a 50-100 image control cohort retains at least 85% routed-identity recall while using at most 40% of full-image slot time.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `docker-layer-content-scanning`: Replace fixed highest-first broad selection with versioned adaptive payload selection, separate selector identity from reusable execution evidence, and require non-authoritative shadow gates before broad rollout.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects Docker Registry config access, post-claim mode assignment, layer-plan binding, policy hashing, image-to-blob coverage state, aggregate rollout evidence, Docker configuration, and related PostgreSQL migration/runtime validation.
|
||||
- Reuses the existing immutable digest identities, account pool, bounded Registry downloader, global blob leases, scan slots, Windows Job containment, result bundles, findings projection, and keycheck pipeline.
|
||||
- Does not change detector classification, credential persistence, Git or Hugging Face scanning, guaranteed scan-slot capacity, Docker worker count, or repository resolver scheduling.
|
||||
- Requires an additive stopped-runtime migration before adaptive execution can be enabled; rollback leaves durable adaptive audit state intact.
|
||||
+218
@@ -0,0 +1,218 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Immutable image content plans are bound after claim
|
||||
The system SHALL resolve a claimed Docker image's exact platform manifest into a canonical bounded plan containing its configuration and ordered layer descriptors, and SHALL bind that plan to the active result reservation before content execution. Adaptive plans SHALL use an exactly validated version-two schema that binds versioned selector and execution-policy identities plus one bounded payload class per descriptor while retaining exact validation of durable version-one plans.
|
||||
|
||||
#### Scenario: Valid immutable manifest
|
||||
- **WHEN** the claimed `repository@sha256:<digest>` resolves to a valid requested-platform manifest
|
||||
- **THEN** the bound plan identifies the same manifest digest and contains only bounded valid SHA-256 content descriptors, sizes, media types, order, policy hashes, classes, and selection reasons
|
||||
|
||||
#### Scenario: Changed replay
|
||||
- **WHEN** the same reservation attempts to bind a different content plan
|
||||
- **THEN** the system rejects the replay as a fenced conflict and executes neither plan
|
||||
|
||||
#### Scenario: Invalid or oversized manifest
|
||||
- **WHEN** the Registry manifest is malformed, exceeds descriptor bounds, or disagrees with the immutable target
|
||||
- **THEN** the system fails closed without claiming complete content coverage
|
||||
|
||||
#### Scenario: Valid aligned configuration history
|
||||
- **WHEN** a bounded digest-verified configuration has non-empty history entries aligned base-to-top with every manifest layer
|
||||
- **THEN** the selector classifies each layer through the versioned bounded classifier and persists only its class enum
|
||||
|
||||
#### Scenario: Untrusted or misaligned configuration history
|
||||
- **WHEN** configuration history is absent, malformed, oversized, or inconsistent with layer or rootfs counts
|
||||
- **THEN** every affected layer is deterministically classified `unknown` and no raw history command is persisted or logged
|
||||
|
||||
#### Scenario: Durable version-one replay
|
||||
- **WHEN** recovery loads a previously bound valid version-one plan
|
||||
- **THEN** the system validates and executes it under its original exact schema without rewriting it as version two
|
||||
|
||||
### Requirement: Content selection is bounded and application-first
|
||||
The system SHALL select bounded image configuration and SHALL deterministically select supported unique layers under configured per-layer, aggregate compressed-byte, and unique-layer-count limits. It SHALL select every eligible unique descriptor when the complete set fits, and for larger images SHALL prioritize versioned payload classes before manifest position and stable tie-breaks.
|
||||
|
||||
#### Scenario: Complete bounded image fits
|
||||
- **WHEN** configuration and every supported unique layer fit all configured descriptor, aggregate-byte, and unique-layer-count bounds
|
||||
- **THEN** the system selects every descriptor regardless of history class
|
||||
|
||||
#### Scenario: Large image requires prioritization
|
||||
- **WHEN** all new unique layers cannot fit the configured aggregate bounds
|
||||
- **THEN** the system greedily considers `copy_add`, `app_config_run`, `package_run`, `unknown`, `other_run`, and `bulk_data` in that order, then highest position, smallest compressed size, and lexical digest
|
||||
|
||||
#### Scenario: Giant base or model layer
|
||||
- **WHEN** a layer exceeds the configured per-layer limit
|
||||
- **THEN** the layer is not downloaded by the normal layer scanner and coverage records `layer_too_large`
|
||||
|
||||
#### Scenario: Image byte budget is exhausted
|
||||
- **WHEN** another unscanned layer would exceed the remaining per-image budget
|
||||
- **THEN** the layer remains unselected and coverage records `image_budget_exhausted`
|
||||
|
||||
#### Scenario: Unique layer count is exhausted
|
||||
- **WHEN** another unscanned unique layer would exceed the configured selected-layer count
|
||||
- **THEN** the layer remains unselected and coverage records `layer_limit_exhausted`
|
||||
|
||||
#### Scenario: Shared layer is already covered
|
||||
- **WHEN** a layer digest has successful coverage under the matching execution policy
|
||||
- **THEN** the image reuses that coverage without consuming its transfer-byte or new-execution-count budget
|
||||
|
||||
#### Scenario: Duplicate positions share one digest
|
||||
- **WHEN** multiple positions in one manifest reference the same eligible immutable digest
|
||||
- **THEN** the system selects all matching positions but budgets and executes that digest at most once
|
||||
|
||||
#### Scenario: Selector input changes
|
||||
- **WHEN** a classifier rule, class order, supported-media rule, or selection bound changes
|
||||
- **THEN** the system derives a different versioned `selection_policy_sha256`
|
||||
|
||||
### Requirement: Layer coverage is globally deduplicated and fenced
|
||||
The system SHALL maintain one authoritative content-scan state per immutable digest and scan-execution policy and SHALL change successful coverage only through matching reservation, lease, plan, and ingestion fences. Selection budgets, classifier weights, retries, lease timing, and checkpoint scheduling MUST NOT partition otherwise identical successful execution evidence.
|
||||
|
||||
#### Scenario: Concurrent images share a layer
|
||||
- **WHEN** two image plans reference the same unscanned digest under the same execution policy concurrently
|
||||
- **THEN** at most one reservation owns its active scan and the other image records shared pending work without duplicate execution
|
||||
|
||||
#### Scenario: Selector policy changes only
|
||||
- **WHEN** an immutable digest was covered successfully and a later plan changes only selector or scheduling policy
|
||||
- **THEN** the later plan reuses the existing execution-policy coverage without launching another scan
|
||||
|
||||
#### Scenario: Execution semantics change
|
||||
- **WHEN** the scanner fingerprint, content validator, or scan-affecting archive semantics differ
|
||||
- **THEN** the system uses a distinct execution-policy namespace and does not reuse incompatible successful coverage
|
||||
|
||||
#### Scenario: Compatible legacy coverage exists
|
||||
- **WHEN** an exact valid version-one plan proves a linked digest is durably covered with execution semantics identical to the requested version-two policy
|
||||
- **THEN** the binding transaction may create a covered version-two alias that retains the original successful reservation, plan, byte count, and completion provenance
|
||||
|
||||
#### Scenario: Legacy evidence is incomplete or ambiguous
|
||||
- **WHEN** legacy evidence is pending, leased, submitted, failed, orphaned, malformed, or not exactly execution-compatible
|
||||
- **THEN** the system does not alias it and requires normal fenced execution under the new policy
|
||||
|
||||
#### Scenario: Successful matching ingestion
|
||||
- **WHEN** a result bundle contains a successful execution for a blob leased by its exact bound plan
|
||||
- **THEN** ingestion marks that digest globally covered for the matching execution policy in the same durable transaction
|
||||
|
||||
#### Scenario: Stale completion
|
||||
- **WHEN** a bundle or worker presents an expired, refunded, or mismatched blob lease
|
||||
- **THEN** it cannot mark the digest covered or advance image coverage
|
||||
|
||||
#### Scenario: Reservation is refunded
|
||||
- **WHEN** an image reservation is durably refunded before handoff
|
||||
- **THEN** only blob leases owned by that reservation are released for bounded reclamation
|
||||
|
||||
### Requirement: Image coverage is explicit and honest
|
||||
The system SHALL persist and expose selected, covered, shared-pending, failed, and intentionally skipped content for each immutable image plan, including bounded payload class, exact reason, and execution and selector policy identities. Configuration history classification SHALL influence priority only and SHALL NOT establish content identity or successful coverage.
|
||||
|
||||
#### Scenario: Every descriptor is covered
|
||||
- **WHEN** configuration and all image layer positions have successful compatible execution-policy coverage
|
||||
- **THEN** the image records complete content coverage
|
||||
|
||||
#### Scenario: Bounds skip content
|
||||
- **WHEN** one or more descriptors are excluded by configured size, budget, count, or format bounds
|
||||
- **THEN** the image may finish as bounded partial coverage but SHALL NOT report complete content coverage
|
||||
|
||||
#### Scenario: Adaptive priority skips content
|
||||
- **WHEN** a large-image descriptor loses deterministic selection to a higher-priority candidate
|
||||
- **THEN** its class, declared bytes, omission reason, and partial image scope remain durable without persisting raw history
|
||||
|
||||
#### Scenario: Selected content remains retryable
|
||||
- **WHEN** at least one selected blob failed retryably or is actively covered by another reservation
|
||||
- **THEN** the image remains deferred without claiming complete coverage
|
||||
|
||||
#### Scenario: Selected content exhausts retries
|
||||
- **WHEN** required selected content reaches its terminal attempt limit
|
||||
- **THEN** the image receives terminal incomplete disposition with durable coverage detail
|
||||
|
||||
#### Scenario: Covered duplicate appears at multiple positions
|
||||
- **WHEN** one successfully covered digest backs multiple positions in an image
|
||||
- **THEN** every matching position records compatible covered scope without duplicate execution
|
||||
|
||||
### Requirement: Layer work resumes without repeating completed content
|
||||
The system SHALL resume an incomplete image from its earliest complete descriptor-position selection map for the same manifest and selector policy, SHALL NOT expand or contract that selection because mutable coverage changed, and SHALL NOT relaunch content with compatible successful execution-policy coverage. It MAY lease a bounded deterministic batch while retaining independent per-blob fences.
|
||||
|
||||
#### Scenario: Parent image retries
|
||||
- **WHEN** an image retry follows partial layer completion
|
||||
- **THEN** the new plan reuses the exact frozen selection, reuses compatible covered digests, and leases only remaining eligible content
|
||||
|
||||
#### Scenario: Covered content changes the available budget
|
||||
- **WHEN** a selected digest becomes globally covered after the first plan was bound
|
||||
- **THEN** a descriptor skipped by the original byte or count budget remains skipped on every later checkpoint under that selector policy
|
||||
|
||||
#### Scenario: Execution policy changes
|
||||
- **WHEN** the same frozen selector baseline runs under a new incompatible execution policy
|
||||
- **THEN** its selected positions remain unchanged while only content lacking compatible coverage becomes executable
|
||||
|
||||
#### Scenario: Bounded multi-blob checkpoint
|
||||
- **WHEN** multiple remaining selected digests fit the configured checkpoint count and byte targets
|
||||
- **THEN** one reservation may lease that deterministic batch while each digest retains an independent lease token, attempt, execution record, and ingestion transition
|
||||
|
||||
#### Scenario: Process crashes after one layer
|
||||
- **WHEN** one layer was durably ingested before a later checkpoint or parent process failed
|
||||
- **THEN** recovery preserves the completed layer and reclaims only unfinished leased content
|
||||
|
||||
#### Scenario: Batch fails before durable handoff
|
||||
- **WHEN** a process scans one or more leased blobs but crashes before their result bundle is durably accepted
|
||||
- **THEN** none of those un-ingested blobs becomes covered and exact lease recovery remains required
|
||||
|
||||
### Requirement: Full-image compatibility and deterministic canary are retained
|
||||
The system SHALL retain the existing full-image scanner, timeout-only canary, and legacy layer path behind configuration. It SHALL additionally provide deterministic `adaptive-canary` and `adaptive` version-two modes, fail closed to full execution without a matching passed rollout gate, and preserve full-image rollback without deleting durable layer state.
|
||||
|
||||
#### Scenario: Full mode
|
||||
- **WHEN** Docker layer mode is disabled or set to `full`
|
||||
- **THEN** the existing immutable full-image execution path remains authoritative
|
||||
|
||||
#### Scenario: Legacy canary retry
|
||||
- **WHEN** an image with a durable previous full-image timeout is retried under the existing canary
|
||||
- **THEN** its immutable manifest digest selects the same legacy scanner mode as its previous attempt
|
||||
|
||||
#### Scenario: Non-timeout image during legacy canary rollout
|
||||
- **WHEN** an image has no durable previous full-image command timeout
|
||||
- **THEN** existing canary configuration keeps that image on the full-image execution path
|
||||
|
||||
#### Scenario: Adaptive canary assignment
|
||||
- **WHEN** a matching passed shadow gate permits `adaptive-canary`
|
||||
- **THEN** a versioned stable hash of every immutable manifest identity selects the same configured adaptive cohort across retries while non-members remain full-image controls
|
||||
|
||||
#### Scenario: Adaptive gate is absent or stale
|
||||
- **WHEN** adaptive configuration lacks a completed passing report for its exact scan, execution, and selector policy hashes
|
||||
- **THEN** no adaptive plan is bound and the claim uses full-image execution
|
||||
|
||||
#### Scenario: Broad adaptive mode
|
||||
- **WHEN** matching shadow evidence passes and the low-percentage production canary remains within safety gates for at least one repository-refresh interval
|
||||
- **THEN** operators may explicitly configure `adaptive` for all eligible immutable Docker claims
|
||||
|
||||
#### Scenario: Layer mode rollback
|
||||
- **WHEN** operators return configuration from `layer`, `canary`, `adaptive-canary`, or `adaptive` to `full`
|
||||
- **THEN** new claims use full-image execution without deleting durable layer plans, coverage, or rollout evidence
|
||||
|
||||
### Requirement: Controlled evidence gates production rollout
|
||||
The system SHALL compare adaptive scanning with completed full-image controls through a bounded non-authoritative evaluator and SHALL keep adaptive production modes fail-closed until a policy-exact aggregate report passes security, coverage, completion, and throughput gates. Shadow execution MUST NOT mutate authoritative queue, reservation, coverage, finding, candidate, keycheck, projection, or source-counter state.
|
||||
|
||||
#### Scenario: Controlled adaptive shadow cohort
|
||||
- **WHEN** the evaluator runs against 50-100 exact immutable images with completed full-image controls under one scan fingerprint
|
||||
- **THEN** it executes the candidate adaptive policy under the same bounded scanner semantics and records only aggregate policy hashes, counts, slot milliseconds, failures, thresholds, and timestamps
|
||||
|
||||
#### Scenario: Shadow identity comparison
|
||||
- **WHEN** routed and detector recall are calculated
|
||||
- **THEN** `(service, provider_key_hash)` and `detector_secret_hash` sets exist only in protected memory long enough to calculate aggregate full, adaptive, and intersection counts
|
||||
|
||||
#### Scenario: Shadow privacy
|
||||
- **WHEN** a shadow run completes, fails, or logs diagnostics
|
||||
- **THEN** no target name, raw config command, finding, secret, provider material, identity set, or Registry bearer is persisted in rollout evidence or ordinary logs
|
||||
|
||||
#### Scenario: Shadow non-authority
|
||||
- **WHEN** shadow execution emits candidate material or completes a blob
|
||||
- **THEN** it cannot create authoritative findings or candidates, route keychecks, mark global blob coverage, alter image disposition, or change production mode
|
||||
|
||||
#### Scenario: Routed recall or slot-time gate fails
|
||||
- **WHEN** paired routed-identity recall is below 85% or aggregate adaptive slot time exceeds 40% of aggregate full slot time
|
||||
- **THEN** the report does not authorize adaptive production execution
|
||||
|
||||
#### Scenario: Safety or completion gate fails
|
||||
- **WHEN** fewer than 50 paired controls complete, any digest/fence/containment requirement fails, omissions are unaccounted, or resource, quarantine, projection, credential, or source health regresses
|
||||
- **THEN** the report does not authorize adaptive production execution
|
||||
|
||||
#### Scenario: Policy changes after a passing report
|
||||
- **WHEN** scan, execution, or selector policy identity changes
|
||||
- **THEN** the earlier report is stale and adaptive execution remains fail-closed until the new exact policy passes another shadow cohort
|
||||
|
||||
#### Scenario: Acceptance criteria pass
|
||||
- **WHEN** 50-100 paired controls retain at least 85% routed identities at no more than 40% slot time with every safety gate satisfied
|
||||
- **THEN** operators may start a low deterministic adaptive canary but SHALL NOT enable broad adaptive mode until that canary remains stable for at least one repository-refresh interval
|
||||
@@ -0,0 +1,46 @@
|
||||
## 1. Add Versioned Adaptive Policy Inputs
|
||||
|
||||
- [x] 1.1 Add fail-closed `adaptive-canary` and `adaptive` configuration/CLI parsing, bounded selector and checkpoint settings, and stable immutable-manifest assignment
|
||||
- [x] 1.2 Separate scanner, execution, and selector policy hashes so selection and scheduling changes do not partition compatible successful blob coverage
|
||||
- [x] 1.3 Fetch and verify bounded image configuration, map aligned non-empty history to ordered layers, and emit only versioned bounded payload classes with deterministic `unknown` fallback
|
||||
- [x] 1.4 Add exact version-two plan validation and canonical hashing while retaining unchanged version-one plan and execution validation
|
||||
|
||||
## 2. Migrate Durable Policy And Evidence State
|
||||
|
||||
- [x] 2.1 Add and exactly validate the image-coverage selector-policy column, policy-specific baseline index, aggregate shadow-report state, timing/fence fields, and migration marker
|
||||
- [x] 2.2 Backfill selector identities transactionally from exact linked version-one plans and fail the migration on orphaned, malformed, or ambiguous coverage rows
|
||||
- [x] 2.3 Implement transactionally fenced legacy-to-execution-policy coverage aliasing only for exact successful semantically compatible evidence
|
||||
- [x] 2.4 Add stopped-runtime migration preconditions and prove rollback leaves all legacy plans and coverage rows intact
|
||||
|
||||
## 3. Implement Adaptive Selection And Resume
|
||||
|
||||
- [x] 3.1 Select all supported unique descriptors for complete bounded images and implement deterministic class, position, size, and digest ordering for larger images
|
||||
- [x] 3.2 Persist exact classes and selected, reused, duplicate, unsupported, oversized, byte-budget, and count-budget reasons with honest partial coverage
|
||||
- [x] 3.3 Freeze the earliest complete descriptor-position map by queue, manifest, and selector policy across checkpoints, retries, and execution-policy changes
|
||||
- [x] 3.4 Lease deterministic bounded multi-blob checkpoints while retaining independent digest locks, lease tokens, attempts, execution records, and reclaim behavior
|
||||
- [x] 3.5 Execute and ingest version-two plans with mixed per-blob outcomes, immutable class/provenance metadata, and no repeat work for compatible covered digests
|
||||
|
||||
## 4. Add Non-Authoritative Shadow Gates
|
||||
|
||||
- [x] 4.1 Implement a bounded operator-invoked paired evaluator for 50-100 completed full-image controls using the exact candidate scan, execution, and selector policies
|
||||
- [x] 4.2 Derive routed and detector identity intersections only in protected memory and persist only aggregate counts, slot timing, failures, thresholds, policy hashes, and timestamps
|
||||
- [x] 4.3 Fence shadow execution from queue disposition, reservations, global coverage, findings, candidates, keychecks, projections, source counters, and automatic mode changes
|
||||
- [x] 4.4 Require a matching completed report with routed recall at least 85% and adaptive/full slot ratio at most 40% before adaptive canary assignment
|
||||
- [x] 4.5 Expose aggregate adaptive selection, reuse, omission, checkpoint, completion, slot-time, and gate metrics without target or secret material
|
||||
|
||||
## 5. Verify Safety And Behavior
|
||||
|
||||
- [x] 5.1 Add unit tests for bounded history parsing, secret-free class persistence, all-fit selection, large-image ranking, stable hashes, exact reasons, and malformed-plan rejection
|
||||
- [x] 5.2 Add PostgreSQL tests for atomic migration/backfill, selector-frozen resume, compatible coverage aliasing, incompatible policy partitioning, concurrent deduplication, and stale fences
|
||||
- [x] 5.3 Add checkpoint tests proving bounded multi-blob mixed outcomes, crash recovery, independent attempts, and no post-coverage selection expansion
|
||||
- [x] 5.4 Add rollout tests for stable adaptive canary assignment, stale/missing gate fallback, shadow non-authority/privacy, aggregate recall, and claim-through-handoff slot timing
|
||||
- [x] 5.5 Run targeted unit, runtime-safety, migration, query-shape, and real PostgreSQL integration suites plus strict OpenSpec validation
|
||||
- [x] 5.6 Preserve deterministic warning retryability so incomplete findings remain visible without repeated blob downloads or false successful coverage
|
||||
- [x] 5.7 Suppress target labels in private shadow filter logs and expose only fixed aggregate failure categories for future evidence
|
||||
|
||||
## 6. Migrate And Roll Out Conservatively
|
||||
|
||||
- [x] 6.1 Stop runtime, verify lease quiescence, apply the additive migration, restart in unchanged timeout-only canary mode, and verify source/keycheck/projection/quarantine health
|
||||
- [ ] 6.2 Run the aggregate-only shadow evaluator on 50-100 completed controls and keep adaptive execution disabled unless every recall, timing, completion, privacy, and safety gate passes
|
||||
- [ ] 6.3 Enable a low deterministic adaptive canary only after a matching passing report and monitor it for at least one repository-refresh interval
|
||||
- [ ] 6.4 Expand canary or enable broad adaptive mode only if runtime gates remain satisfied; otherwise return new claims to full or timeout-only canary without deleting audit state
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-01
|
||||
@@ -0,0 +1,254 @@
|
||||
## Context
|
||||
|
||||
DockerHub currently queues immutable `repository@sha256:<platform-manifest>` targets and gives each target to TruffleHog's Docker source as one indivisible operation. The process must fetch and inspect every layer before the external 600-second deadline. Large model images contain tens of compressed GiB, commonly dominated by one binary/model-weight layer, so two such targets can occupy both Docker workers while still producing only partial findings.
|
||||
|
||||
Measured over 48 hours, 196 deadline terminations consumed about 32.7 Docker worker-hours. Every timeout emitted findings, but all timeout findings routed to only 14 provider keys and none was usable. The queue currently resets target attempts after a timeout and the timeout disposition bypasses the maximum-attempt check, so the same immutable target can restart from byte zero indefinitely.
|
||||
|
||||
Registry manifest resolution already obtains the exact platform child manifest and ordered layer digests. The missing information is each descriptor's compressed size, media type, image configuration descriptor, durable per-layer coverage, and a scanner capable of processing one bounded content blob independently.
|
||||
|
||||
The runtime must retain its authenticated Docker pool, immutable digest identities, slot-first PostgreSQL admission, result-bundle fencing, Windows Job containment, output bounds, private temporary storage, and fail-closed behavior. Bearer tokens and provider keys must remain memory-only or in their existing protected stores.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- End unbounded timeout retries while preserving findings emitted before a deadline.
|
||||
- Scan image configuration and useful application layers without downloading giant model/data layers.
|
||||
- Resume image coverage after failure at layer granularity rather than restarting completed work.
|
||||
- Scan each immutable content digest once globally and reuse its durable coverage across images.
|
||||
- Keep download, disk, archive expansion, command execution, and database state bounded and fenced.
|
||||
- Report exact selected, covered, skipped, failed, and shared-pending scope for every image.
|
||||
- Prove the strategy against full-image control results before broad production enablement.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Reconstruct a runnable merged container root filesystem.
|
||||
- Claim complete image coverage when configured size or format bounds skip content.
|
||||
- Download giant layers in a separate long-running production lane in the initial change.
|
||||
- Implement arbitrary Registry hosts, private registries, cross-host credential forwarding, or resumable CDN downloads.
|
||||
- Change global guaranteed scan slots, Docker worker count, keycheck classification, or non-Docker scanners.
|
||||
- Delete the existing full-image implementation or schema on rollback.
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. Keep the image target, bind a fenced layer plan after claim
|
||||
|
||||
The existing immutable image target remains the queue and reporting identity. After slot-first admission claims an image, a dedicated PostgreSQL connection resolves its exact manifest and binds a canonical `docker_layer_plan` to that result reservation under the queue lease/event fences, analogous to exact Git plan binding.
|
||||
|
||||
The plan contains a version, repository, platform manifest digest, configuration descriptor, ordered layer descriptors, configured byte limits, selection reason for each descriptor, and a canonical SHA-256. The reservation stores the canonical JSON and hash. Replay is idempotent; changed replay or manifest mismatch is a conflict.
|
||||
|
||||
This avoids multiplying normal target-queue rows and keeps one authoritative parent disposition while still permitting durable child coverage.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- A separate full-image heavy lane still redownloads all content and cannot resume.
|
||||
- One target-queue row per layer complicates parent completion, finding attribution, and alternate repository fetch sources.
|
||||
- Encoding mutable coverage metadata into queue identity would break immutable image deduplication.
|
||||
|
||||
### 2. Add durable content and image-coverage tables
|
||||
|
||||
Add PostgreSQL-compatible tables through the additive runtime-safety migration:
|
||||
|
||||
- `docker_content_blobs`: one row per valid SHA-256 content digest, descriptor kind (`config` or `layer`), declared compressed bytes, media type, state, bounded attempts, active reservation/lease fence, completion metadata, and last bounded error.
|
||||
- `docker_image_blob_coverage`: one row per platform manifest and content digest with ordered position, selected/skipped reason, plan hash, and covered timestamp.
|
||||
|
||||
Successful content state is global because the digest authenticates the bytes. The current image repository is retained only as the bounded fetch source in the reservation plan. If one reservation owns a blob, another image records it as shared-pending and defers without duplicating the download. Expired/refunded reservation leases are reclaimable.
|
||||
|
||||
Blob completion changes only during ingestion of a matching reservation plan and matching per-blob execution record. Refund/recovery releases matching blob leases. A stale token cannot mark coverage or insert findings.
|
||||
|
||||
### 3. Select content by configurable byte budget, top layers first
|
||||
|
||||
The image configuration is always selected within a small independent hard bound because `Env`, labels and history are high-value and cheap.
|
||||
|
||||
Layer selection walks the manifest from highest layer to base layer. Already-covered digests require no byte budget. New layers are selected while both the per-layer compressed-byte cap and remaining per-image compressed-byte budget permit them. Non-selected descriptors are recorded with explicit reasons such as `layer_too_large`, `image_budget_exhausted`, or `unsupported_media_type`.
|
||||
|
||||
Initial numeric defaults are chosen only after the controlled spike. Configuration always has hard upper bounds; invalid values fail closed to full-image mode while the production gate is disabled.
|
||||
|
||||
This is a deliberate partial-coverage policy. It targets application/config layers and globally amortizes common base layers without pretending that skipped model weights were scanned.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- A whole-image size cutoff can discard a small valuable application layer sitting above giant weights.
|
||||
- Selecting oldest/base layers first spends budget on widely shared dependencies before application content.
|
||||
- Inferring binary content without downloading a compressed tar stream is not reliable.
|
||||
|
||||
### 4. Stream, verify, scan, and delete one blob at a time
|
||||
|
||||
For each blob leased by the reservation:
|
||||
|
||||
1. Obtain an account-scoped Registry bearer through the existing Docker account manager.
|
||||
2. Request the exact Registry blob with a bounded custom redirect policy. Redirects must remain HTTPS, must not contain userinfo, must reject local/private destinations, and must never receive the Registry Authorization header on another host.
|
||||
3. Stream into a private bounded work file while computing SHA-256. Reject excess bytes, short bodies, digest mismatch, unsupported media types, disk-floor violations, and response deadline exhaustion.
|
||||
4. Scan configuration JSON directly. Scan supported layer archives with TruffleHog `filesystem` under the existing OwnedProcess Job, timeout, output cap, configured detector policy, archive size/depth/time bounds, and normal process priority.
|
||||
5. Delete the blob work file before releasing the scan slot. No layer bearer, provider key, or raw result is written outside existing protected result artifacts.
|
||||
|
||||
Each layer runs as its own bounded TruffleHog command. This adds small startup overhead but gives exact completion, global deduplication, and restart from the first unfinished layer. Findings are enriched with immutable image, blob digest, kind, and layer position before normal bundle staging.
|
||||
|
||||
The controlled spike must prove that the installed TruffleHog build correctly scans supported real layer archive media types. Unsupported formats remain explicit uncovered scope rather than being silently accepted.
|
||||
|
||||
### 5. Parent disposition derives from explicit child execution
|
||||
|
||||
The result bundle carries the exact bound plan plus one execution record per claimed blob. Ingestion verifies plan hash and lease ownership before updating blob state.
|
||||
|
||||
- Successful blob command and digest verification marks that blob covered globally.
|
||||
- Incomplete blob execution preserves its findings but does not mark it covered.
|
||||
- Retryable blob failures release it to bounded retry; terminal failures remain explicit uncovered scope.
|
||||
- A blob active under another valid reservation causes a short parent deferral without charging a content attempt.
|
||||
- An image whose selected blobs are covered and whose remaining blobs are intentionally skipped completes with `coverage_complete=false` and detailed reasons.
|
||||
- An image with retryable selected work remains deferred; exhausted selected work becomes terminal failed/degraded according to the bound policy.
|
||||
|
||||
The existing full-image path remains available as a rollback/control path.
|
||||
|
||||
### 6. Fix timeout accounting before layer rollout
|
||||
|
||||
Both production-v2 and legacy completion paths stop resetting attempts for target-scoped timeouts. Timeout disposition checks the configured maximum before returning deferred. Source-wide infrastructure failures may retain their existing attempt-refund semantics.
|
||||
|
||||
A stopped-runtime repair reconciles only unfenced Docker queue rows whose latest durable result is a command timeout. It derives prior immutable-target attempt count from durable scans, sets the queue attempt count up to the configured maximum, and terminally closes already exhausted rows. It does not requeue or mutate active leases, reservations, findings, or successful targets.
|
||||
|
||||
### 7. Roll out through deterministic modes and evidence gates
|
||||
|
||||
Configuration exposes `full`, `canary`, and `layer` modes. Canary eligibility requires a durable previous full-image command timeout, and membership within that eligible set is a stable hash of immutable manifest digest. New images and normally completed controls therefore remain on the full path during canary rollout. Defaults remain `full` until migration and spike criteria pass.
|
||||
|
||||
The offline spike compares 20-50 timeout-heavy images and completed controls without printing findings or keys. It records bytes transferred, wall/slot time, peak resource use, distinct detector identities, and routed key identity recall. Production canary additionally tracks coverage reasons, blob reuse, timeout rate, keycheck candidate yield, and strict usable yield.
|
||||
|
||||
Broad enablement requires:
|
||||
|
||||
- no credential/token persistence regression;
|
||||
- no stale-fence or duplicate-blob completion;
|
||||
- exact digest verification for every covered blob;
|
||||
- material byte and slot-hour reduction on heavy images;
|
||||
- all routed key identities from completed control images retained, unless an explicitly reviewed coverage bound explains the difference;
|
||||
- no quarantine growth, projection regression, or source failure increase.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Secrets can exist in skipped giant layers] -> Record explicit incomplete scope, keep configurable budgets, compare controls, and retain full mode for targeted replay.
|
||||
- [Scanning individual layers can report files deleted by later whiteouts] -> Preserve layer provenance and treat this as historical image-content evidence rather than merged-root truth.
|
||||
- [Archive support differs by media type] -> Verify installed binary in the spike and mark unsupported formats uncovered.
|
||||
- [Global deduplication can be poisoned by stale completion] -> Require streamed digest verification plus reservation/lease/plan fences in the same ingestion transaction.
|
||||
- [Registry blob redirects introduce SSRF or credential-forwarding risk] -> Use a bounded validated redirect implementation and strip authorization across hosts.
|
||||
- [Per-layer process startup adds overhead for tiny layers] -> Batch measurement first; skip already-covered layers and permit a bounded future batching optimization only if needed.
|
||||
- [A crash can strand blob leases] -> Tie leases to result reservations, release on refund, and permit exact expiry reclamation.
|
||||
- [Database state grows with layer relationships] -> Enforce descriptor count bounds, compact metadata, indexed identities, and retention metrics.
|
||||
- [Layer mode can reduce broad secret coverage] -> Report coverage honestly and retain deterministic full-mode controls and rollback.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Ship and test bounded timeout accounting independently; repair exhausted historical timeout rows with sources stopped.
|
||||
2. Add nullable reservation plan columns, content/coverage tables, indexes, runtime validation, and additive migration marker. Keep mode `full`.
|
||||
3. Run a local controlled spike against retained immutable targets and select conservative byte/archive defaults from evidence.
|
||||
4. Deploy code with `full` mode, migrate offline, restart, and verify no behavior change.
|
||||
5. Enable deterministic low-percentage canary only for DockerHub images whose latest durable full-image result timed out. Monitor at least one full repository-refresh interval and sufficient heavy-image samples.
|
||||
6. Increase the timeout-fallback canary only after acceptance gates pass. Keep broad `layer` mode disabled until completed-control routed recall becomes adequate under revised bounds.
|
||||
7. Roll back by returning mode to `full`. Durable layer tables and nullable columns remain for audit and future resume; no destructive migration is required.
|
||||
|
||||
## Capability Spike Evidence
|
||||
|
||||
The installed `C:\Tools\trufflehog.exe` development build was exercised through 39 bounded
|
||||
`filesystem` invocations over synthetic direct tar, gzip-tar, zstd-tar, Docker outer-tar, and OCI
|
||||
outer-tar fixtures. All invocations exited successfully and all 24 expected synthetic findings were
|
||||
preserved. Extensionless gzip and zstd blobs were content-sniffed successfully.
|
||||
|
||||
The minimum archive depth was two for a direct tar, three for direct compressed layers, and four
|
||||
for an OCI/Docker archive containing a compressed layer. Shallower bounds produced an explicit
|
||||
non-fatal `max archive depth reached` diagnostic. `--archive-max-size` was proven to be a per-member
|
||||
bound rather than a cumulative compressed-image bound, so application-side per-blob and aggregate
|
||||
byte limits remain mandatory.
|
||||
|
||||
Initial conservative implementation bounds are therefore archive depth four, archive member size
|
||||
256 MiB, archive timeout 30 seconds, one-GiB aggregate selected compressed bytes per image, 256 MiB
|
||||
per selected layer, and filesystem concurrency two. These are implementation starting points, not
|
||||
broad-rollout acceptance evidence. The required aggregate timeout-heavy/completed-control image
|
||||
comparison remains an explicit gate before canary expansion.
|
||||
|
||||
The controlled aggregate comparison then ran against ten repeatedly timed-out immutable images and
|
||||
ten longest completed controls. No target names, findings, keys, or credentials were printed or
|
||||
persisted. Exact manifest resolution succeeded for all 20 images. The timeout-heavy manifests
|
||||
contained 341.68 GB of compressed descriptors; the bounded layer policy selected 944.02 MB (0.28%)
|
||||
and processed 84 blobs in 313.03 worker-seconds versus 6,015.31 historical full-image seconds
|
||||
(5.2%). All selected timeout-heavy blobs completed. The bounded path produced three routed
|
||||
credential identities, none overlapping the two identities in the historical incomplete full
|
||||
results, so it added useful scope while avoiding another byte-zero full-image retry.
|
||||
|
||||
The completed controls contained 98.05 GB; the policy selected 1.61 GB (1.64%) and processed 77
|
||||
blobs in 357.83 worker-seconds versus 5,956.38 historical seconds (6.0%). Five blobs reported
|
||||
bounded incomplete chunk processing. Only four of 19 historical routed identities were retained
|
||||
(21.1% recall), and distinct detector-identity recall was 19 of 326 (5.8%). System sampling across
|
||||
the run observed average CPU 20.27%, peak CPU 44.48%, minimum available physical memory 17.38 GB,
|
||||
and peak committed memory 30.00 GB; no resource-limit failure occurred.
|
||||
|
||||
This evidence accepts the existing 1 MiB config, 256 MiB per-layer, 1 GiB aggregate, eight-layer,
|
||||
archive-depth-four, 256 MiB archive-member, 30-second archive, 600-second blob, and filesystem
|
||||
concurrency-two bounds only for a deterministic timeout-fallback canary. It rejects broad random
|
||||
canary or broad layer mode because completed-control routed recall failed the acceptance gate. Full
|
||||
mode remains authoritative for new and normally completed images.
|
||||
|
||||
The post-refinement regression gate passed 329 related tests, including bounded transfer and content
|
||||
validation, parent/blob timeout accounting, PostgreSQL migration atomicity, policy-scoped global
|
||||
deduplication, stale fences, reclaim/quarantine, multi-checkpoint resume, and the durable full-timeout
|
||||
canary-eligibility transition. Strict OpenSpec validation passed, and no application bytecode was
|
||||
present after the run.
|
||||
|
||||
Rejected approaches from the capability spike are `--force-skip-archives` (it suppresses expected
|
||||
archive findings), relying on media-type labels without content verification, and relying on
|
||||
TruffleHog's `--archive-max-size` as an outer download or aggregate expansion bound. The exact
|
||||
development binary must remain fingerprinted by a capability contract because its reported version
|
||||
does not identify a stable release.
|
||||
|
||||
## Initial Production Canary Evidence
|
||||
|
||||
The additive migration and guarded historical timeout repair were applied offline before the
|
||||
timeout-only canary. Runtime first restarted in full mode with no layer rows, and the bounded canary
|
||||
was then enabled at 2,500 basis points only for images whose latest durable full-image result was a
|
||||
command timeout. New and normally completed images remained on the full scanner.
|
||||
|
||||
The first selected production image completed nine unique content checkpoints plus one final
|
||||
no-work completion checkpoint. The nine bounded blobs transferred 327,512 bytes and completed in
|
||||
33.176 seconds of aggregate parent duration, including 9.358 seconds of transfer and 23.496 seconds
|
||||
of contained filesystem scanning. All nine blobs reached policy-scoped global coverage, the parent
|
||||
finished `done`, no selected continuation remained, and quarantine stayed unchanged at 192 items /
|
||||
337,349,428 bytes. The prior full-image path for this eligibility class reached the 600-second
|
||||
deadline.
|
||||
|
||||
The initial run exposed and then verified a selection-continuity invariant: per-image byte/layer
|
||||
bounds must apply to the first immutable selection set, not be recomputed after each covered blob.
|
||||
The binder now reuses the earliest exact `(queue, manifest, coverage policy, position)` selection map
|
||||
on every later reservation. A PostgreSQL regression proves that layers skipped by the original
|
||||
count budget remain skipped after selected layers become globally covered. Production replay then
|
||||
completed without leasing content outside the original config-plus-eight-layer selection.
|
||||
|
||||
This single successful image proves the end-to-end checkpoint, resume, bounded selection and final
|
||||
completion paths, but is not enough evidence to increase the 25% timeout-only canary. Expansion
|
||||
still requires a longer observation window and more naturally eligible timeout samples.
|
||||
|
||||
The following overnight window added a second timeout-only image before a host reboot. Across both
|
||||
images, twelve unique blobs reached policy-scoped coverage with no retryable or terminal blob
|
||||
failure. The layer path used 53.192 seconds and transferred 79,614,382 bytes, compared with 1,201.330
|
||||
seconds consumed by the immediately preceding full-image timeout attempts. One parent completed;
|
||||
the second retained six exact selected checkpoints for durable resume. No layer findings or routed
|
||||
candidate identities were produced in this small sample, and quarantine remained unchanged.
|
||||
|
||||
Because this installation is an experimental rather than production service, the operator approved
|
||||
expanding the stable timeout-only cohort from 2,500 to 10,000 basis points. This does not enable broad
|
||||
layer mode: every new or normally completing image still uses the full scanner, and only an image
|
||||
with a durable prior full-image timeout may enter the bounded layer fallback. Broad layer mode
|
||||
remains rejected by the completed-control recall result. Task 7.5 remains open until the expanded
|
||||
cohort produces additional completed parents and a stable runtime observation window.
|
||||
|
||||
The expanded cohort exposed one additional metadata-normalization defect: identical content digests
|
||||
and sizes can be referenced through equivalent Docker and OCI media-type labels. Treating the label
|
||||
text itself as immutable metadata caused the Docker source to stop fail-closed before handoff. The
|
||||
binder now compares the validated semantic content class (`config-json`, `layer-tar`, `layer-gzip`,
|
||||
or `layer-zstd`) while still rejecting kind, byte-size, and compression-class conflicts. A real
|
||||
PostgreSQL concurrency regression and the related layer/runtime suites passed 126 tests. After the
|
||||
restart, 116 layer reservations were acknowledged in the first ten minutes, global covered blobs
|
||||
grew from 9 to 113, no blob entered a failed state, the source remained running, and quarantine was
|
||||
unchanged. The immediate digest queue remained intentionally thin (six due and eighteen delayed),
|
||||
while 12,679 repository anchors remained available to refill it after claimable digest work drains.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Which revised per-layer and per-image bounds can improve completed-control routed recall beyond the measured 21.1% without losing the measured slot-hour advantage?
|
||||
- Which OCI/Docker layer compression media types does the installed TruffleHog filesystem source handle reliably?
|
||||
- Is one TruffleHog process per selected layer sufficiently efficient, or is a later bounded multi-layer archive batch warranted?
|
||||
- What short deferral is appropriate when all remaining selected blobs are actively leased by other reservations?
|
||||
@@ -0,0 +1,31 @@
|
||||
## Why
|
||||
|
||||
Large Docker images currently monopolize both Docker scan workers until the 600-second deadline and can retry indefinitely because timeout completion resets the target attempt counter. Over the measured 48-hour window, hard timeouts consumed about 32.7 worker-hours while repeated partial scans produced no usable LLM access, so full-image retries are reducing useful throughput without providing proportional coverage.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Enforce the existing bounded target-attempt policy for Docker timeouts while preserving findings emitted before termination.
|
||||
- Resolve immutable image manifests into image configuration and ordered content-addressed layers with bounded size metadata.
|
||||
- Scan image configuration and selected layer content under an explicit per-image byte budget instead of treating every image as an indivisible download.
|
||||
- Deduplicate successful layer scans globally by immutable layer digest so shared base layers are not downloaded and scanned repeatedly.
|
||||
- Prioritize upper application layers and small layers; record oversized or out-of-budget layers as explicit uncovered scope rather than silently claiming complete image coverage.
|
||||
- Preserve the existing full-image path behind a rollout gate for controlled comparison and rollback.
|
||||
- Repair currently deferred Docker targets whose timeout attempts were incorrectly reset.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `docker-layer-content-scanning`: Bounded, content-addressed Docker config and layer scanning with global deduplication, explicit coverage, safe retry limits, and controlled rollout against the existing full-image scanner.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects Docker Registry manifest/blob access, immutable Docker target planning, scan queue state, result metadata, and Docker source configuration.
|
||||
- Adds durable PostgreSQL state for layer identities, leases, coverage, attempts, and image-to-layer plans.
|
||||
- Reuses the existing authenticated Docker account pool, scan-slot limiter, Windows Job containment, bundle ingestion, findings projection, and keycheck pipeline.
|
||||
- Requires an offline additive runtime-safety migration before enabling production layer scanning.
|
||||
- Does not change Git, Hugging Face, keycheck classification, global guaranteed scan-slot capacity, or secret persistence boundaries.
|
||||
+190
@@ -0,0 +1,190 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Docker timeout retries are bounded
|
||||
The system SHALL count target-scoped Docker command timeouts against the configured target-attempt maximum and SHALL preserve partial findings without creating an unbounded retry loop.
|
||||
|
||||
#### Scenario: Timeout before attempt limit
|
||||
- **WHEN** a Docker command times out before the configured maximum attempt
|
||||
- **THEN** its emitted findings remain durable and the immutable target is deferred using the configured timeout delay without resetting its attempt count
|
||||
|
||||
#### Scenario: Timeout reaches attempt limit
|
||||
- **WHEN** a Docker command times out at the configured maximum attempt
|
||||
- **THEN** its emitted findings remain durable and the queue records a terminal target disposition
|
||||
|
||||
#### Scenario: Source-wide outage
|
||||
- **WHEN** Docker execution is prevented by a source-wide infrastructure failure rather than target-scoped work
|
||||
- **THEN** the existing fenced source-failure recovery policy remains applicable
|
||||
|
||||
### Requirement: Immutable image content plans are bound after claim
|
||||
The system SHALL resolve a claimed Docker image's exact platform manifest into a canonical bounded plan containing its configuration and ordered layer descriptors, and SHALL bind that plan to the active result reservation before content execution.
|
||||
|
||||
#### Scenario: Valid immutable manifest
|
||||
- **WHEN** the claimed `repository@sha256:<digest>` resolves to a valid requested-platform manifest
|
||||
- **THEN** the bound plan identifies the same manifest digest and contains only bounded valid SHA-256 content descriptors, sizes, media types, and order
|
||||
|
||||
#### Scenario: Changed replay
|
||||
- **WHEN** the same reservation attempts to bind a different content plan
|
||||
- **THEN** the system rejects the replay as a fenced conflict and executes neither plan
|
||||
|
||||
#### Scenario: Invalid or oversized manifest
|
||||
- **WHEN** the Registry manifest is malformed, exceeds descriptor bounds, or disagrees with the immutable target
|
||||
- **THEN** the system fails closed without claiming complete content coverage
|
||||
|
||||
### Requirement: Content selection is bounded and application-first
|
||||
The system SHALL always select bounded image configuration and SHALL select new layers from highest to lowest under configured per-layer and per-image compressed-byte limits.
|
||||
|
||||
#### Scenario: Giant base or model layer
|
||||
- **WHEN** a layer exceeds the configured per-layer limit
|
||||
- **THEN** the layer is not downloaded by the normal layer scanner and coverage records `layer_too_large`
|
||||
|
||||
#### Scenario: Image byte budget is exhausted
|
||||
- **WHEN** another unscanned layer would exceed the remaining per-image budget
|
||||
- **THEN** the layer remains unselected and coverage records `image_budget_exhausted`
|
||||
|
||||
#### Scenario: Shared layer is already covered
|
||||
- **WHEN** a layer digest has successful global coverage
|
||||
- **THEN** the image reuses that coverage without consuming its transfer budget or launching another scan
|
||||
|
||||
#### Scenario: Upper and base layers both fit
|
||||
- **WHEN** multiple unscanned layers fit within all configured bounds
|
||||
- **THEN** the system selects them in highest-to-lowest manifest order
|
||||
|
||||
### Requirement: Layer coverage is globally deduplicated and fenced
|
||||
The system SHALL maintain one authoritative content-scan state per immutable digest and SHALL change successful coverage only through matching reservation, lease, plan, and ingestion fences.
|
||||
|
||||
#### Scenario: Concurrent images share a layer
|
||||
- **WHEN** two image plans reference the same unscanned digest concurrently
|
||||
- **THEN** at most one reservation owns its active scan and the other image records shared pending work without duplicate execution
|
||||
|
||||
#### Scenario: Successful matching ingestion
|
||||
- **WHEN** a result bundle contains a successful execution for a blob leased by its exact bound plan
|
||||
- **THEN** ingestion marks that digest globally covered in the same durable transaction
|
||||
|
||||
#### Scenario: Stale completion
|
||||
- **WHEN** a bundle or worker presents an expired, refunded, or mismatched blob lease
|
||||
- **THEN** it cannot mark the digest covered or advance image coverage
|
||||
|
||||
#### Scenario: Reservation is refunded
|
||||
- **WHEN** an image reservation is durably refunded before handoff
|
||||
- **THEN** only blob leases owned by that reservation are released for bounded reclamation
|
||||
|
||||
### Requirement: Registry blob transfer is authenticated, bounded, and verified
|
||||
The system SHALL fetch selected content from the trusted Docker Registry using the existing account pool, bounded streaming, private storage, safe redirect handling, and exact digest verification.
|
||||
|
||||
#### Scenario: Valid content download
|
||||
- **WHEN** the Registry returns exactly the declared bounded blob bytes whose SHA-256 matches the descriptor
|
||||
- **THEN** the private work artifact becomes eligible for scanning
|
||||
|
||||
#### Scenario: Cross-host redirect
|
||||
- **WHEN** the trusted Registry redirects a blob request to an allowed public HTTPS content host
|
||||
- **THEN** the system follows only the bounded validated redirect and does not forward Registry authorization to the other host
|
||||
|
||||
#### Scenario: Unsafe redirect
|
||||
- **WHEN** a blob redirect uses HTTP, userinfo, a local/private destination, or exceeds redirect bounds
|
||||
- **THEN** the transfer fails closed without exposing authentication material
|
||||
|
||||
#### Scenario: Size or digest mismatch
|
||||
- **WHEN** streamed bytes exceed bounds, end short, or do not match the expected digest
|
||||
- **THEN** the system deletes the work artifact and does not record successful coverage
|
||||
|
||||
#### Scenario: Insufficient private storage
|
||||
- **WHEN** the configured private work volume cannot retain its required free-space floor
|
||||
- **THEN** no blob download begins and the failure receives bounded retry disposition
|
||||
|
||||
### Requirement: Configuration and layers are scanned independently
|
||||
The system SHALL scan bounded image configuration and each newly leased supported layer as independent contained commands while preserving image and layer provenance on findings.
|
||||
|
||||
#### Scenario: Configuration contains candidate material
|
||||
- **WHEN** bounded configuration JSON contains detector-matching data
|
||||
- **THEN** findings identify the image and configuration digest and enter the normal result and keycheck pipeline
|
||||
|
||||
#### Scenario: Supported layer completes
|
||||
- **WHEN** TruffleHog filesystem scanning of a verified layer archive completes successfully
|
||||
- **THEN** its findings retain image, layer digest, kind, and position provenance and the layer becomes globally covered after fenced ingestion
|
||||
|
||||
#### Scenario: Layer scan is incomplete
|
||||
- **WHEN** a layer command times out or exits without confirmed completion
|
||||
- **THEN** emitted findings remain durable but that digest does not become covered
|
||||
|
||||
#### Scenario: Unsupported media type
|
||||
- **WHEN** a layer compression or media type is not supported by the validated scanner path
|
||||
- **THEN** no unsafe fallback executes and image coverage records `unsupported_media_type`
|
||||
|
||||
### Requirement: Image coverage is explicit and honest
|
||||
The system SHALL persist and expose selected, covered, shared-pending, failed, and intentionally skipped content for each immutable image plan.
|
||||
|
||||
#### Scenario: Every descriptor is covered
|
||||
- **WHEN** configuration and all image layers have successful global coverage
|
||||
- **THEN** the image records complete content coverage
|
||||
|
||||
#### Scenario: Bounds skip content
|
||||
- **WHEN** one or more descriptors are excluded by configured size, budget, or format bounds
|
||||
- **THEN** the image may finish as bounded partial coverage but SHALL NOT report complete content coverage
|
||||
|
||||
#### Scenario: Selected content remains retryable
|
||||
- **WHEN** at least one selected blob failed retryably or is actively covered by another reservation
|
||||
- **THEN** the image remains deferred without claiming complete coverage
|
||||
|
||||
#### Scenario: Selected content exhausts retries
|
||||
- **WHEN** required selected content reaches its terminal attempt limit
|
||||
- **THEN** the image receives terminal incomplete disposition with durable coverage detail
|
||||
|
||||
### Requirement: Layer work resumes without repeating completed content
|
||||
The system SHALL resume an incomplete image from selected content that lacks successful global coverage and SHALL NOT relaunch completed content digests.
|
||||
|
||||
#### Scenario: Parent image retries
|
||||
- **WHEN** an image retry follows partial layer completion
|
||||
- **THEN** the new plan reuses covered digests and leases only remaining eligible content
|
||||
|
||||
#### Scenario: Process crashes after one layer
|
||||
- **WHEN** one layer was durably ingested before a later layer or parent process failed
|
||||
- **THEN** recovery preserves the completed layer and reclaims only unfinished leased content
|
||||
|
||||
### Requirement: Full-image compatibility and deterministic canary are retained
|
||||
The system SHALL retain the existing full-image scanner behind configuration. Canary layer execution SHALL require a durable previous full-image command timeout and SHALL be selected deterministically from immutable manifest identity within that eligible set.
|
||||
|
||||
#### Scenario: Full mode
|
||||
- **WHEN** Docker layer mode is disabled or set to `full`
|
||||
- **THEN** the existing immutable full-image execution path remains authoritative
|
||||
|
||||
#### Scenario: Canary retry
|
||||
- **WHEN** a canary image is retried
|
||||
- **THEN** its immutable manifest digest selects the same scanner mode as its previous attempt
|
||||
|
||||
#### Scenario: Non-timeout image during canary rollout
|
||||
- **WHEN** an image has no durable previous full-image command timeout
|
||||
- **THEN** canary configuration keeps that image on the full-image execution path
|
||||
|
||||
#### Scenario: Layer mode rollback
|
||||
- **WHEN** operators return configuration from `layer` or `canary` to `full`
|
||||
- **THEN** new claims use full-image execution without deleting durable layer audit state
|
||||
|
||||
### Requirement: Controlled evidence gates production rollout
|
||||
The system SHALL compare layer scanning with completed full-image controls and SHALL keep broad production layer mode disabled until security, coverage, and throughput gates pass.
|
||||
|
||||
#### Scenario: Controlled spike
|
||||
- **WHEN** the offline spike runs against bounded heavy and completed-control samples
|
||||
- **THEN** it records aggregate bytes, wall time, slot time, coverage, distinct detector identities, and routed key recall without printing findings or keys
|
||||
|
||||
#### Scenario: Acceptance criteria fail
|
||||
- **WHEN** digest integrity, fence safety, routed-key recall, resource bounds, or throughput criteria fail
|
||||
- **THEN** production remains in full mode
|
||||
|
||||
#### Scenario: Completed-control recall fails but timeout fallback passes
|
||||
- **WHEN** bounded layer scanning materially reduces timeout-heavy work but does not retain completed-control routed-key recall
|
||||
- **THEN** operators may canary only prior full-image timeout retries and SHALL NOT enable broad layer mode
|
||||
|
||||
#### Scenario: Acceptance criteria pass
|
||||
- **WHEN** the controlled spike and deterministic production canary satisfy all defined gates
|
||||
- **THEN** operators may increase canary coverage or enable layer mode through configuration
|
||||
|
||||
### Requirement: Historical timeout state is repaired safely
|
||||
The system SHALL reconcile incorrectly reset Docker timeout attempts only while sources are stopped and only for unfenced immutable targets backed by durable timeout results.
|
||||
|
||||
#### Scenario: Exhausted historical timeout target
|
||||
- **WHEN** an unfenced deferred Docker target has durable timeout executions at or above the configured maximum
|
||||
- **THEN** repair marks it terminal without deleting its existing findings
|
||||
|
||||
#### Scenario: Active or ambiguous target
|
||||
- **WHEN** a Docker target has an active lease, reservation, event fence, or ambiguous latest result
|
||||
- **THEN** repair leaves it unchanged
|
||||
@@ -0,0 +1,45 @@
|
||||
## 1. Restore Bounded Timeout Semantics
|
||||
|
||||
- [x] 1.1 Make Docker timeout disposition terminal at the configured target-attempt maximum in production-v2 and legacy paths without resetting attempts
|
||||
- [x] 1.2 Add regression tests proving partial findings survive and timeout attempts stop at the configured limit
|
||||
- [x] 1.3 Add a stopped-runtime guarded repair for unfenced historical Docker targets whose durable timeout attempts were reset
|
||||
|
||||
## 2. Validate Layer Scanning
|
||||
|
||||
- [x] 2.1 Prove the installed TruffleHog filesystem source safely scans bounded Docker gzip and supported OCI layer archives with preserved findings
|
||||
- [x] 2.2 Run an aggregate-only spike across timeout-heavy and completed-control images and select conservative config, layer, image, archive, and deadline bounds
|
||||
- [x] 2.3 Record spike acceptance evidence and rejected formats/approaches in the design
|
||||
|
||||
## 3. Add Durable Layer State
|
||||
|
||||
- [x] 3.1 Add reservation plan columns plus Docker content-blob and image-coverage tables, indexes, migration marker, and exact runtime validation
|
||||
- [x] 3.2 Implement bounded canonical Docker layer-plan validation, hashing, idempotent fenced binding, and deterministic canary selection
|
||||
- [x] 3.3 Implement content lease claim, expiry, refund, retry, terminal failure, and globally successful coverage transitions
|
||||
- [x] 3.4 Wire matching blob execution and image coverage updates into the authoritative fenced result-ingestion transaction
|
||||
|
||||
## 4. Implement Bounded Registry Content Access
|
||||
|
||||
- [x] 4.1 Extend exact platform manifest resolution with bounded configuration and ordered layer size/media descriptors
|
||||
- [x] 4.2 Implement authenticated Registry blob streaming with safe redirects, byte/disk/deadline bounds, private artifacts, and SHA-256 verification
|
||||
- [x] 4.3 Implement contained configuration and layer archive scans with immutable image/blob provenance and deterministic cleanup
|
||||
|
||||
## 5. Integrate Layer-Aware Execution
|
||||
|
||||
- [x] 5.1 Select configuration and highest-first layers under per-layer and per-image byte budgets while reusing globally covered digests
|
||||
- [x] 5.2 Resolve and bind the Docker layer plan after the fenced parent claim using a dedicated database connection
|
||||
- [x] 5.3 Execute only leased blobs, preserve partial findings, and emit exact plan/execution/coverage metadata through result bundles
|
||||
- [x] 5.4 Resume deferred images without rerunning covered blobs and defer shared active content without charging duplicate attempts
|
||||
|
||||
## 6. Add Controlled Rollout and Observability
|
||||
|
||||
- [x] 6.1 Add bounded `full`, deterministic `canary`, and `layer` configuration with full-image rollback
|
||||
- [x] 6.2 Expose aggregate selected/covered/skipped/shared/failed bytes, blob reuse, transfer duration, timeout, and image coverage metrics without secret material
|
||||
- [x] 6.3 Keep existing scan-slot, Windows Job, output, keycheck, and credential-isolation invariants unchanged
|
||||
|
||||
## 7. Verify and Deploy
|
||||
|
||||
- [x] 7.1 Add unit tests for plan bounds, selection order, downloader security, digest verification, provenance, timeout policy, and canary stability
|
||||
- [x] 7.2 Add PostgreSQL integration tests for concurrent global deduplication, stale fences, refund/reclaim, partial ingestion, resume, and image coverage
|
||||
- [x] 7.3 Run related regression suites, strict OpenSpec validation, and aggregate control comparison; document evidence
|
||||
- [x] 7.4 Apply the additive migration offline, repair historical attempts, restart in full mode, and verify no behavior regression
|
||||
- [ ] 7.5 Enable a bounded deterministic canary, monitor throughput/coverage/keycheck/quarantine gates, and expand only if acceptance criteria pass
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-17
|
||||
@@ -0,0 +1,128 @@
|
||||
# Task 1.3 Baseline Evidence
|
||||
|
||||
## Status
|
||||
|
||||
This file records the best available historical baseline for task 1.3 and its
|
||||
reproducibility limits. It does not claim that a pre-change image was rerun during
|
||||
this change, and it does not convert current tests into pre-change evidence.
|
||||
|
||||
An immutable rerun of the exact pre-change source and image is unavailable. On
|
||||
2026-09-18, the reviewer explicitly accepted this documentary baseline and waived
|
||||
that rerun requirement. Task 1.3 therefore relies on the historical record below;
|
||||
current regression results remain separately identified as post-change evidence.
|
||||
|
||||
## Historical pre-change record
|
||||
|
||||
The repository's [`DOCKER_MIGRATION.md`](../../../DOCKER_MIGRATION.md), under
|
||||
"Current Verified Evidence," records a final run on 2026-09-15. It reports:
|
||||
|
||||
- The selected container regression suite completed with **578 passed** and
|
||||
**7 expected Windows-only tests skipped on Linux**, with no failures recorded.
|
||||
- The selected suite used test image
|
||||
`sha256:4ce11325728ba3e58e6643c1c8e800f317179d5c7c50e7e80568b58f62dbdfd0`.
|
||||
- The separate six-test `docker/test_verify.py` suite passed on Windows and WSL
|
||||
Linux; those six tests were not included in the 578 count.
|
||||
- The recorded offline E2E passed its result/check gates and removed its owned
|
||||
resources. This is contextual historical evidence, not a rerun for task 1.3.
|
||||
|
||||
The `add-minimal-remote-scan-workers` OpenSpec change was created on 2026-09-17,
|
||||
after that recorded run. The preserved records contain no pre-existing failure
|
||||
for the cited 578-test selection to list separately. "No recorded failure" means
|
||||
only that no failure record was found; it is not proof that no unrecorded attempt
|
||||
failed.
|
||||
|
||||
## Reproducibility limits
|
||||
|
||||
- The old test image is unavailable and was not retrieved, rebuilt, or executed.
|
||||
Current Docker runs use new, dedicated test image identities.
|
||||
- This workspace has no commit history. On 2026-09-18, `git status` reported
|
||||
`No commits yet on master` and every repository path as untracked; `git log`
|
||||
failed because the branch has no commits.
|
||||
- `WORKSPACE.md` names a source-side provenance commit, but also states that this
|
||||
source-only snapshot includes modified and selected untracked files and did not
|
||||
copy source Git history. It is not an immutable tree for either 2026-09-15 or
|
||||
the moment immediately before the 2026-09-17 OpenSpec change.
|
||||
- The image digest and prose result are therefore useful historical evidence but
|
||||
cannot independently reconstruct or rerun the claimed chronology from this
|
||||
repository.
|
||||
|
||||
## Current expanded selection
|
||||
|
||||
The current reviewed allowlist is the `SELECTION` mapping in
|
||||
`tests/container_unit.py`. On 2026-09-18, its stdlib-only selection check reported
|
||||
**33 modules and 649 test definitions**, with no runner skips at declaration time;
|
||||
pytest parametrization may expand the executed count.
|
||||
|
||||
The current mapping selects:
|
||||
|
||||
```text
|
||||
test_docker_foundation.py
|
||||
test_container_runtime.py
|
||||
test_owned_process.py
|
||||
test_owned_process_linux.py
|
||||
test_owned_process_boundary.py
|
||||
test_runtime_bootstrap_authority.py
|
||||
test_supervisor_foreground_shutdown.py
|
||||
test_supervisor_startup_rollback.py
|
||||
test_supervisor_managed_postgres_gate.py
|
||||
test_observer_only_coordinated_shutdown.py
|
||||
test_supervisor_safety.py
|
||||
test_postgres_runtime.py
|
||||
test_container_security.py
|
||||
test_runtime_security.py
|
||||
test_postgres_empty_initialization.py
|
||||
test_container_migration_paths.py
|
||||
test_container_provider_portability.py
|
||||
test_container_e2e_helpers.py
|
||||
test_container_import.py
|
||||
test_container_import_config.py
|
||||
test_container_projection_recovery.py
|
||||
test_result_bundle_v2.py
|
||||
test_pipeline_cutover_invariants.py
|
||||
test_custom_provider_detector_compatibility.py
|
||||
test_scan_execution.py
|
||||
test_synthetic_llm_pipeline.py
|
||||
test_worker_api.py
|
||||
test_worker_api_runtime.py
|
||||
test_worker_assignment.py
|
||||
test_worker_package.py
|
||||
test_remote_worker_db.py
|
||||
test_admin_api.py
|
||||
test_edge_deployment.py
|
||||
```
|
||||
|
||||
The remote-worker, package, administration, edge, synthetic-pipeline, and bundle
|
||||
modules in this current selection postdate the historical baseline. Their
|
||||
presence demonstrates current review scope, not pre-change execution.
|
||||
|
||||
## Current post-change evidence
|
||||
|
||||
The expanded selection and isolated end-to-end gates were rerun on 2026-09-19.
|
||||
They validate the completed implementation but do not replace the historical
|
||||
pre-change baseline:
|
||||
|
||||
```text
|
||||
python -I -S -B tests/container_unit.py --check-selection
|
||||
container-unit: AST OK; 33 modules, 650 test definitions, no runner skips (parametrizations expand in pytest)
|
||||
|
||||
isolated Linux container selection
|
||||
867 passed, 7 expected platform skips
|
||||
|
||||
wsl.exe -d Ubuntu-24.04 -- python3 docker/verify.py
|
||||
E2E passed; project truf-worker-test-99b193a64f46a0002e34da9c5ff8029a; artifacts removed
|
||||
|
||||
python -B docker/verify_packaged_workers.py ...
|
||||
Windows/Linux packaged-worker E2E passed; run 348d24046fb7b8d4; cleanup complete; foreign Docker state unchanged
|
||||
Windows package manifest sha256 f0bbbbf79d9f5f3f17440a60561b307fa7bafd312ab6110b14098ff90755f8e6
|
||||
Linux worker image manifest list sha256 888bffc5519fd3a27cb52d20eaea1143f89ac9779bb2e4b0d998e450583c57a0
|
||||
|
||||
python -B docker/verify_edge_e2e.py --timeout-seconds 1200
|
||||
Edge/fail2ban E2E passed; run 597b3ff7826055e1; cleanup complete; foreign Docker state unchanged
|
||||
|
||||
openspec validate add-minimal-remote-scan-workers --strict --no-interactive
|
||||
Change 'add-minimal-remote-scan-workers' is valid
|
||||
```
|
||||
|
||||
The current runs used only random or dedicated test-owned identities and verified
|
||||
that foreign container and volume metadata remained unchanged. No current result
|
||||
is represented as historical or pre-change evidence.
|
||||
@@ -0,0 +1,147 @@
|
||||
## Context
|
||||
|
||||
The current PostgreSQL pipeline already has single-target admission, capacity reservation, immutable Git/Docker plans, scan/error classification, canonical v2 `.trb` staging, transactional ingestion, projection, and detailed keychecks. The execution seam is the claim-to-`stage_claim` path in `app/console_runner.py`, not a replacement scheduler.
|
||||
|
||||
`scan_target_result()` in `app/scanner.py` dispatches existing source implementations. Bundle staging extracts candidates, including structured Postman evidence; ingestion persists findings/candidates and separately schedules projection and keycheck. JSONL projection is not a prerequisite for keycheck, and a TruffleHog `Verified` field is not a completed detailed keycheck.
|
||||
|
||||
Current execution is not automatically portable to a DB-free client: process launch authority, exact Git/Docker plans, reservation accounting, and producer recovery depend on the local runtime. Recovery can inspect local PID/executable identity and local files, which cannot establish remote worker liveness. The design adapts these boundaries while keeping scanner/provider/error behavior intact.
|
||||
|
||||
Development takes place in `D:\truf-workers`, a source-only snapshot of the current `D:\truf-docker` working tree, including uncommitted fixes. Production data and Git history were not copied. Existing unarchived specifications are references; `openspec/specs/` has no canonical baseline. Old three-global-permit/`S:` handoff assumptions apply to their historical local deployment, not this remote boundary. PostgreSQL remains authoritative despite older status-file accounting descriptions.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Move expensive download/scan work to trusted Windows/Linux clients with minimum new code and persistent state.
|
||||
- Preserve scanner results, attribution, exact-plan coverage, error classification, retry decisions, candidate routing, and detailed server-side keychecks.
|
||||
- Keep source/provider/scanner configuration centralized; support client-selected N slots capped by the server across a user's devices.
|
||||
- Recover through a fixed 24-hour default assignment deadline, durable retries, and existing identity fencing, without heartbeat traffic.
|
||||
- Secure the small public surface and test locally against empty storage and synthetic inputs.
|
||||
|
||||
**Non-Goals:**
|
||||
- Client detailed keychecks, a second broker/queue/result format, generic job infrastructure, batch claims, worker affinity, or worker blacklists.
|
||||
- New scan retry counts, a three-attempt/dead-letter rule, exactly-once physical execution, or heartbeat/lease renewal protocols.
|
||||
- Client disk encryption, mTLS, public/untrusted workers, auto-update, complex roles, multiple server replicas, or PostgreSQL container extraction.
|
||||
- Production database import, production deployment, changes to the active runtime, or archiving unrelated OpenSpec changes.
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. Keep one runtime and add a thin remote boundary
|
||||
|
||||
```text
|
||||
Internet -> Caddy :443
|
||||
|-- /api/v1/worker/* [device token] -> Worker API
|
||||
`-- /<random-admin>/* [login/password] -> micro-admin
|
||||
|
||||
Server runtime: PostgreSQL + existing queue/reservations
|
||||
discovery/scheduling -> admission
|
||||
receiver -> existing ingester -> projection
|
||||
`-> detailed keycheck
|
||||
|
||||
Client slot: claim -> existing download/scan -> canonical .trb -> upload/ack
|
||||
```
|
||||
|
||||
Run Worker API and micro-admin as runtime-managed processes, not independent database-owning services. Keep the current dashboard backend private; any dashboard information in the admin area shares its authentication boundary. PostgreSQL, supervisor control, Caddy's control API, and raw backend ports have no public host bindings.
|
||||
|
||||
Reuse `reserve_and_claim_target`, ambiguous-admission reconciliation, `scan_target_result`, bundle staging/validation, `mark_result_bundle_ready`, and ingester/projector/keycheck paths. Extract only the code needed to run one already-planned scan and build its complete bundle without PostgreSQL. Adapt DB-bound plan inputs and node-local process authority rather than giving a client a DSN or a supervisor credential. Preserve `OwnedProcess` containment and cleanup; the client must launch only the known scanner tools, not arbitrary server-supplied commands.
|
||||
|
||||
Alternative rejected: copying `run_cycle_v2` unchanged or implementing a second scanner/queue. Both either retain server authority on clients or create divergent behavior.
|
||||
|
||||
### 2. Central configuration, minimal device identity, compatible jobs
|
||||
|
||||
Client-authored operational configuration has only server URL, opaque device token, and desired positive slot count N. Generated local pending-work state and workspace paths are not independent source configuration. The server resolves source/scanner settings, immutable plan, limits, and only the credentials required for this task. It must not send the whole config, discovery credential pool, database credentials, admin secrets, or control authority.
|
||||
|
||||
Bind each device token to its owning user and a server-issued/stable worker identity; store token hashes and support revocation. Keep only the identity/quota metadata required for administration, bound to current reservations rather than creating `remote_jobs` or another scheduler. Server configuration supplies default limits; admin changes the per-user cap across all devices.
|
||||
|
||||
Send protocol/build/policy compatibility metadata with normal claim traffic, not a background registration/liveness service. Validate the pinned scanner/custom detector policy and supported source execution on the client's OS/architecture before consuming a target. Pass a config snapshot or stable effective-config identity so the result stays attributable even if the server config changes mid-task. A client rejects an incompatible job instead of silently using its own settings.
|
||||
|
||||
Alternative rejected: worker-owned provider settings/tokens or another config management system. Trusted clients receive the task-specific secrets they need over HTTPS; this is not a sandbox against a compromised client.
|
||||
|
||||
### 3. One claim per free slot with existing backpressure
|
||||
|
||||
Each free client slot requests one assignment. Effective admission is bounded by N locally, the atomic per-user server cap across devices, and existing source/global/output-capacity rules. A client advertising N=8 does not override an admin cap of 3. Concurrent claims from different devices must not overshoot a shared cap. Reducing a cap stops new admissions until usage falls below it; it does not invent cancellation semantics for already-issued work.
|
||||
|
||||
Keep the current node-local scan-permit mechanics. A client work slot holds its assignment through durable result acknowledgement, which bounds pending uploads when the server is unavailable. Release the native scan permit at its existing handoff point; do not conflate that permit with remote user quota. A durably accepted bundle, acknowledged pre-bundle terminal disposition, or completed expiry recovery releases remote assignment quota exactly once; server spool credits remain governed by existing ingestion accounting.
|
||||
|
||||
Persist the existing admission/request identity before the first claim request. Retry ambiguous requests with the same identity and reconcile the same reservation; do not allocate another task because a claim response was lost. Scope recovery to the authenticated device. Empty queues or denied capacity return a bounded polling delay, not a scan error or a new queue item.
|
||||
|
||||
Alternative rejected: batches and client backlogs. Per-free-slot claims reuse current scheduling and avoid another recovery structure.
|
||||
|
||||
### 4. Fixed expiry, no remote heartbeat
|
||||
|
||||
Set `expires_at` from the server clock at assignment commit, with a configurable duration defaulting to 24 hours. The deadline includes download, scan, and upload. Ordinary API contact, retries, and client restarts do not renew it. Internal scanner timeouts and server discovery leases keep their existing meanings; local slot bookkeeping is not a worker heartbeat.
|
||||
|
||||
Use one periodic runtime recovery pass for expired remote assignments, initially once per minute and configurable. Extend existing recovery ownership/state checks instead of inventing a remote liveness monitor. Reconcile reservations, target claims, output credits, and dependent exact-plan/blob leases together. Local producer-PID recovery must not release remote claims. No associated lease may silently expire earlier and allow a second owner while the remote assignment remains valid.
|
||||
|
||||
Requeue unfinished expired work using the existing pre-handoff infrastructure-loss/refund path where applicable, not a fabricated scanner/provider failure or a new retry limit. The same worker may claim the task again. New issuance has a distinct current reservation/attempt identity. Both periodic recovery and result acceptance use the same atomic ownership/deadline boundary.
|
||||
|
||||
Reject an unaccepted old result once its assignment expires or is superseded; it must not finish newer work, update plan coverage, or return somebody else's credit. A retry of an already accepted identical result still receives its original acknowledgement even after the deadline. Ready/accepted bundles awaiting ingestion are not unfinished worker assignments and must not be expired back into the queue.
|
||||
|
||||
Trade-off accepted: a crashed client can delay work for about a day, and a partitioned client can keep physically scanning after reissue. Fencing gives one authoritative acceptance, not exactly-once physical execution.
|
||||
|
||||
### 5. Reuse the full bundle and durable handoff
|
||||
|
||||
The API needs only claim/reconcile, upload/acknowledge, and the existing pre-bundle failure/release outcome where necessary. Exact URL naming is implementation detail under `/api/v1/worker/`; do not add a heartbeat endpoint, remote shell, separate keycheck jobs, or a general command API.
|
||||
|
||||
Pre-bundle terminal reports use the same issuance fencing and replay rules as results. Persist/retry their identity until acknowledged; a lost reply resolves to the original disposition without charging retries, updating counters, or releasing quota/credits again. A delayed report from an expired/superseded issuance cannot mutate a new issuance even when the same device owns both. Authoritative acknowledgement resolves that client slot exactly once.
|
||||
|
||||
Send canonical v2 `.trb` bytes containing the existing findings, detector identities, source/origin/context, errors, metadata, candidate evidence, and exact-plan identity. Preserve structured Postman candidate extraction before ingestion. Do not send only provider names or reconstruct a reduced result JSON. Keep detailed keychecks entirely server-side after ingestion through current candidate/dedup/cache/routing logic. Retained `runtime/keychecks` outputs keep their existing location and retention rules.
|
||||
|
||||
Upload a binary stream, not base64 JSON, into a bounded server-owned partial file. Enforce the existing bundle limit/capacity reservation (currently 192 MiB where configured), time limits, identity, version, and hash. Validate all paths/archive members with the existing codec; clients cannot choose arbitrary server paths. Publish durably using the existing same-filesystem atomic handoff and mark the exact reservation ready before acknowledging server custody.
|
||||
|
||||
Acknowledgement means the validated bundle and its ready/recovery state survive a server restart; it does not mean projection or detailed keycheck has finished. An identical retry returns the same receipt without double ingestion/accounting. A conflicting body for an accepted identity is rejected. Interrupted uploads never count as successful scans; full-bundle retry is sufficient for the MVP, with no resumable-upload protocol.
|
||||
|
||||
Keep the accepted identity, digest, and reconciliation outcome in existing authoritative records independently of the spool file. Ordinary ingestion and bundle cleanup must not erase the information needed to recover a receipt: accept-with-lost-reply, ingest, clean up, restart, and retry after the deadline must still return the original acceptance for identical bytes and reject conflicting bytes. No second receipt queue or result store is required.
|
||||
|
||||
Keep the pending bundle and its assignment identity locally until durable acceptance is confirmed. After acknowledgement, existing cleanup may delete local task artifacts. A definitive fenced rejection transitions local work to an explicit stale/discard outcome with bounded cleanup, not an infinite retry or a false success. Scan errors still use the existing disposition path; HTTP retry/backoff does not consume scanner/provider retry budgets.
|
||||
|
||||
Crash recovery must cover publication-before-ready and ready-before-response windows using exact current ownership and existing spool reconciliation. Never infer success from the mere presence of an unvalidated file or expire already accepted work because a producer PID is absent.
|
||||
|
||||
### 6. Restricted edge and separate admin login bans
|
||||
|
||||
Use direct DNS to Caddy for the initial deployment and HTTPS for all worker/admin transport. The public admin prefix contains at least 128 bits of randomness, but login/password remains mandatory for every administrative route and asset. Caddy password authentication with a supported password hash is the minimal starting point; no plaintext password configuration or public signup. Protect typed mutating actions against CSRF with validated Origin/CSRF handling. No generic supervisor-command passthrough.
|
||||
|
||||
Only authenticated admin pages show operational summaries or typed user/device-token/quota and existing queue controls. Keep the standalone `/dashboard` route closed. Unknown paths return 404. Add no-store/same-origin-referrer, a restrictive compatible CSP, and production HSTS; avoid third-party assets and secrets/raw findings in diagnostics.
|
||||
|
||||
Run fail2ban on the host, reading redacted Caddy authentication-failure events. Two actual bad-credential submissions to the admin boundary within ten minutes produce a 24-hour IP ban. Do not count an ordinary initial Basic-auth challenge without credentials, unrelated 404s, or Worker API authentication failures. With shared HTTPS ingress, the ban action must update an ADMIN-ONLY Caddy IP deny matcher, with validated configuration reload, not globally drop that address on port 443. Persist ban state and provide SSH unban/recovery.
|
||||
|
||||
Use the direct connection address as client IP. Ignore arbitrary forwarded IP headers; a future trusted proxy requires explicit trust configuration and new tests. If network-level fail2ban actions are ever chosen instead, prove Docker forwarding-chain enforcement and preserve worker availability; a blanket host INPUT rule does not satisfy the contract.
|
||||
|
||||
Alternative rejected: a secret URL as sole protection, separate dashboard exposure, or a shared admin/worker IP jail. No extra auth service, mTLS, Cloudflare dependency, or client-disk encryption is required.
|
||||
|
||||
### 7. Observability from existing state
|
||||
|
||||
Correlate source/target, user/worker, reservation/issuance, issue/finish time, duration, outcome, and safe error category. Expose per-worker unfinished, completed, failed, and expired counts and last authenticated API contact. Distinguish bundle acceptance from later ingestion/keycheck completion and define counters from existing authoritative transitions so duplicate uploads do not inflate them.
|
||||
|
||||
Record issuance, acceptance, expiry/requeue, duplicate upload, and stale-result rejection events using existing logs/records. Never label a worker online/offline merely from silence: there is no heartbeat. Do not introduce a telemetry database or monitoring stack. Raw findings belong in the existing protected result storage, not transport/auth/debug logs.
|
||||
|
||||
### 8. Empty and isolated local validation first
|
||||
|
||||
Keep production checkout, processes, Docker containers, and volumes untouched. Reuse the existing stdlib checks, reviewed unit harness, and synthetic E2E driver. Before any runtime test, make test project names, image tags, volumes, ports, and config independent of inherited deployment defaults. The copied `compose.yaml`/legacy import overrides are not safe test launch instructions.
|
||||
|
||||
Initialize PostgreSQL empty in new test-owned storage; never restore a production dump or mount production PGDATA. Scrub inherited DSNs, provider secrets, runtime overrides, and proxies before application imports. Disable live discovery/keycheck autostart; explicitly feed synthetic tasks and mock detailed provider transports. No production credentials, paid calls, discovered public targets, or uncontrolled TruffleHog verification.
|
||||
|
||||
Exercise existing local and new remote scan execution on the same fixtures and compare normalized bundle/DB results while ignoring only transport identity/timing. Include Git/Docker immutable-plan and Postman-candidate parity, both custom detector directions, and all current source error dispositions. Test Windows and Linux clients, N>1 and multi-device caps, fixed-clock expiry/reissue races, lost claim/result replies, restarts, malformed/conflicting bundles, and auth isolation. Use controlled clocks instead of waiting a real day.
|
||||
|
||||
Use an isolated local TLS/internal network for Caddy/API tests and deny external egress. Validate admin bans through the actual edge, including a worker sharing the banned admin IP. Package a pinned Windows portable client and Linux client/container using current tool dependencies; no installer service or updater is necessary. Do not claim cross-platform readiness from mocked scans alone.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Long assignment lifetime] Slow recovery and duplicate physical work are deliberate trade-offs; fencing and idempotent acceptance protect authoritative state.
|
||||
- [DB-bound execution seams] Git/Docker plans and process authority require adaptation; parity tests are a release gate, not permission to rewrite provider logic.
|
||||
- [Trusted client access] Clients can see task credentials and findings; issue least-needed credentials, redact logs, revoke device tokens, and use HTTPS.
|
||||
- [Shared NAT and aggressive admin bans] Two mistakes can lock out an operator; scope bans to admin and retain tested SSH recovery.
|
||||
- [Existing active OpenSpec deltas] Avoid mixing unrelated changes or reviving superseded local-only assumptions; archive/sync is a separate request.
|
||||
- [Unsafe inherited launch defaults] Static checks can run now; runtime/E2E execution waits for explicit test isolation and image/config provenance checks.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Prepare the separate source-only Git workspace and complete this plan; do not start runtime services.
|
||||
2. Establish neutral isolated tests, then implement the narrow execution/admission/transport boundary against empty PostgreSQL with synthetic fixtures.
|
||||
3. Run local parity, failure/recovery, cross-platform, and edge-auth gates before enabling any real remote claims.
|
||||
4. Keep remote admission disabled by default until explicitly configured. Preserve the local execution path using shared scanner logic; do not require a remote worker for existing local operation.
|
||||
5. Production deployment and any additive identity/reservation schema migration need a separate reviewed rollout. No production data migration or rebuild is performed during this planning task.
|
||||
6. To roll back a later rollout, stop new remote claims, retain/reconcile accepted bundles, and drain or expire outstanding assignments through current recovery before returning to local-only execution. Do not blindly drop reservation metadata or switch binaries underneath unfinished remote work.
|
||||
|
||||
## Open Questions
|
||||
|
||||
No blocking product decisions remain. Implementation must verify the smallest DB-free Git/Docker execution seam, the exact existing reservation fields to extend, and the host-specific fail2ban/Caddy reload mechanism through tests before release. Deployment-specific hostname, generated admin prefix/password, device tokens, and quota values are supplied during provisioning, not embedded in this plan.
|
||||
@@ -0,0 +1,29 @@
|
||||
## Why
|
||||
|
||||
Truf needs to use trusted Windows and Linux machines for download/scan work without rewriting its already-debugged parser, queue, error policy, ingestion, or detailed keychecks. A separate source-only workspace and an empty, isolated local test environment let us develop this boundary without copying the large production database or touching the active runtime.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a thin authenticated HTTPS adapter around existing target reservation and canonical `.trb` result handoff, not a second task system.
|
||||
- Keep PostgreSQL, discovery, scheduling, ingestion, projection, and detailed keycheck in one server Docker runtime; use a separate Caddy edge.
|
||||
- Run existing download/scan logic on trusted clients. The server supplies the task, immutable plan, and required source/scanner settings; the client config contains server URL, device token, and desired execution slots.
|
||||
- Claim one task per free client slot, subject to an atomic server-side per-user cap across devices and existing admission/backpressure rules.
|
||||
- Use a configurable fixed assignment lifetime, default 24 hours, with periodic expiry recovery and no worker heartbeat. Preserve existing retry/error decisions, allow the same worker to reclaim work, fence stale assignments, and acknowledge result retries idempotently.
|
||||
- Expose only the authenticated Worker API and a long-random-path authenticated admin area. Keep the standalone dashboard and backend/control/database ports private; ban admin IPs for 24 hours after two actual failed login attempts within ten minutes without banning Worker API traffic.
|
||||
- Reuse existing records/logs for worker counts, durations, assignment outcomes, and last API contact.
|
||||
- Validate with empty local storage and synthetic fixtures, including crash/retry/expiry scenarios, without production data or real provider requests.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `distributed-scan-workers`: Minimal remote scan execution, centralized settings, admission, fixed expiry, durable result handoff, observability, and isolated local verification.
|
||||
- `restricted-public-access`: Private backend/dashboard topology, authenticated worker/admin access, admin-only login bans, and safe diagnostics.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None. `openspec/specs/` is empty in this snapshot. Existing unarchived deltas are design references, not canonical specifications to modify or archive as part of this work.
|
||||
|
||||
## Impact
|
||||
|
||||
The change touches the execution boundary in `app/console_runner.py` and `app/scanner.py`, reservation/recovery in `app/scanner_db.py`, the existing bundle/ingester pipeline, runtime lifecycle wiring, packaging, Docker/Caddy deployment, and regression tests. New persistent state is limited to necessary worker identity/token/quota bindings and metadata attached to existing reservations; PostgreSQL remains the only server queue authority. Client detailed keychecks, another broker, a parallel result format, worker heartbeats, automatic updates, and production migration are out of scope.
|
||||
+145
@@ -0,0 +1,145 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Existing pipeline semantics remain authoritative
|
||||
The system SHALL execute existing download/scan logic on trusted Windows/Linux clients and SHALL retain discovery, queue/reservation authority, ingestion, projection, candidate routing, and detailed keycheck on the server. It MUST NOT create a second queue, reduced result format, client detailed-keycheck flow, or new scanner retry/dead-letter policy.
|
||||
|
||||
#### Scenario: Remote scan matches current local behavior
|
||||
- **WHEN** local and remote execution process the same synthetic target, immutable plan, and effective scanner settings
|
||||
- **THEN** normalized findings, detector/provider identities, origin/context, errors, candidate evidence, coverage decisions, and queue dispositions match apart from transport identities and timing
|
||||
- **AND** detailed keychecks run through the existing server candidate pipeline after ingestion, independently of JSONL projection completion
|
||||
|
||||
### Requirement: Centralized task settings and bounded client authority
|
||||
The server SHALL supply the assigned target, immutable plan, effective source/scanner configuration, compatibility identity, and only task-required credentials. Client-authored operational settings SHALL be server URL, device token, and desired slot count. Clients MUST NOT require PostgreSQL, server supervisor authority, independent provider configuration, or arbitrary remote-command execution.
|
||||
|
||||
#### Scenario: A client has no provider configuration
|
||||
- **WHEN** an authorized compatible client with only its bootstrap settings claims work
|
||||
- **THEN** it receives enough task-specific input to use the existing scanner without database access or worker-maintained provider settings
|
||||
- **AND** it receives no database/admin secrets or unrelated discovery credential pool
|
||||
|
||||
#### Scenario: Incompatible client cannot silently change scan policy
|
||||
- **WHEN** a client's scanner build, detector policy, or source/OS support is incompatible
|
||||
- **THEN** admission refuses incompatible work without consuming a target or falling back to different scan settings
|
||||
|
||||
### Requirement: Per-slot claims respect atomic shared quotas
|
||||
Each free client work slot SHALL claim at most one task. The server SHALL atomically enforce the owner's active-assignment cap across devices together with existing admission/capacity restrictions. Client slots SHALL bound pending unacknowledged work, and quota accounting MUST NOT be confused with node-local scan permits or server spool credits.
|
||||
|
||||
#### Scenario: Multiple devices race for the last slots
|
||||
- **WHEN** a user capped at three active assignments runs two clients configured for eight slots each and they claim concurrently
|
||||
- **THEN** at most three assignments are admitted across both clients and no unused batch/backlog is handed out
|
||||
|
||||
#### Scenario: An administrator lowers a running user's cap
|
||||
- **WHEN** the new cap is below the user's current active-assignment count
|
||||
- **THEN** new claims wait until usage permits admission without inventing cancellation or error outcomes for current work
|
||||
|
||||
### Requirement: Ambiguous claim delivery is recoverable
|
||||
The client SHALL persist a stable admission/request identity before sending a claim, and retries SHALL reconcile the same existing reservation scoped to the authenticated device rather than allocating another target.
|
||||
|
||||
#### Scenario: A claim commits but its reply is lost
|
||||
- **WHEN** a client retries the original claim identity after a network failure or restart
|
||||
- **THEN** the server returns that assignment or its authoritative terminal state without creating an additional reservation or consuming extra quota
|
||||
|
||||
### Requirement: Assignment expiry is fixed and server-owned
|
||||
Remote assignments SHALL expire after a configurable fixed interval, default 24 hours from server-side issuance, including upload. API contact SHALL NOT renew this deadline. No worker heartbeat or liveness probe SHALL be required. A periodic server recovery pass SHALL process expired unfinished assignments using existing infrastructure-loss recovery/accounting, preserving existing scan retry policy.
|
||||
|
||||
#### Scenario: Worker disappears without reporting a scan outcome
|
||||
- **WHEN** its assignment deadline passes and the periodic recovery pass runs
|
||||
- **THEN** the unfinished target becomes claimable again with correctly reconciled credits/quota and dependent plan leases
|
||||
- **AND** the loss does not become a fabricated scanner/provider error, arbitrary retry limit, or worker blacklist
|
||||
|
||||
#### Scenario: The same worker returns after expiry
|
||||
- **WHEN** the previous worker requests available work after its expired assignment is recovered
|
||||
- **THEN** it is eligible to claim that target again under a new issuance identity
|
||||
|
||||
#### Scenario: Client contact does not extend work
|
||||
- **WHEN** the client retries an upload or another API request before the deadline
|
||||
- **THEN** the original server expiry remains unchanged
|
||||
|
||||
### Requirement: Remote ownership fences stale results
|
||||
Result acceptance and expiry/reissue SHALL serialize on current reservation ownership and deadline. Local producer PID checks MUST NOT reclaim remote assignments. Dependent plan/blob/target leases SHALL remain consistent with the remote ownership interval. Only the current unexpired assignment can first publish an authoritative result.
|
||||
|
||||
#### Scenario: Old worker uploads after another issuance
|
||||
- **WHEN** an old worker uploads a previously unaccepted result after expiry or reissue
|
||||
- **THEN** the server rejects it as stale without completing the newer assignment, advancing plan coverage, or releasing its credits
|
||||
|
||||
#### Scenario: Result races with periodic recovery
|
||||
- **WHEN** upload acceptance and expiry recovery race for the same assignment
|
||||
- **THEN** exactly one authoritative transition wins and neither duplicate ingestion nor double capacity release occurs
|
||||
|
||||
#### Scenario: Server restarts while a remote client is still working
|
||||
- **WHEN** runtime recovery cannot find a local producer PID for a valid remote assignment
|
||||
- **THEN** it retains remote ownership until the fixed deadline rather than treating the absent local process as worker death
|
||||
|
||||
### Requirement: Full canonical results have durable idempotent acceptance
|
||||
Clients SHALL send the existing canonical v2 `.trb` as a bounded binary stream, preserving findings, errors, metadata, attribution, exact-plan identity, and candidate evidence. The server SHALL validate identity, format, size, paths, and hash and durably publish the bundle plus ready/recovery state before acknowledging custody. Existing transactional ingestion SHALL remain authoritative. Accepted identity/digest/receipt information SHALL survive ordinary ingestion and spool cleanup in authoritative records independently of the bundle file.
|
||||
|
||||
#### Scenario: Acceptance reply is lost
|
||||
- **WHEN** the server accepts a bundle but its acknowledgement is lost and the client uploads the identical bundle again
|
||||
- **THEN** the server returns the original acceptance without duplicate ingestion, candidates, quota release, or counters
|
||||
- **AND** this acknowledgement remains recoverable after the assignment deadline because acceptance already occurred
|
||||
|
||||
#### Scenario: Accepted identity receives a different body
|
||||
- **WHEN** another bundle with conflicting content is submitted for an accepted identity
|
||||
- **THEN** the server rejects the conflict without replacing the accepted result
|
||||
|
||||
#### Scenario: Client retries after ingestion and normal cleanup
|
||||
- **WHEN** an accepted bundle's reply is lost, ingestion and ordinary cleanup finish, the server restarts, and the client retries after the original deadline
|
||||
- **THEN** identical bytes recover the original receipt despite the absence of the spool file
|
||||
- **AND** conflicting bytes are rejected without duplicate results, candidates, accounting, or counters
|
||||
|
||||
#### Scenario: Upload is truncated or invalid
|
||||
- **WHEN** an upload exceeds bounds, fails validation, disconnects, or crosses the deadline before first acceptance
|
||||
- **THEN** it produces no successful acknowledgement or authoritative findings and cannot mutate another assignment
|
||||
- **AND** bounded partial-file cleanup and existing transport/recovery handling apply
|
||||
|
||||
#### Scenario: Server crashes around bundle publication
|
||||
- **WHEN** the receiver crashes after durable publication or ready-state recording but before replying
|
||||
- **THEN** restart reconciliation and a same-identity client retry recover a single consistent acceptance or authoritative rejection without adopting an unvalidated/stale file
|
||||
|
||||
#### Scenario: Ingestion is delayed past the worker deadline
|
||||
- **WHEN** an accepted ready bundle awaits server ingestion after its former assignment deadline
|
||||
- **THEN** worker expiry recovery does not requeue it as unfinished remote work
|
||||
|
||||
### Requirement: Client recovery separates transport from scan outcomes
|
||||
The client SHALL persist pending bundle and assignment identity until durable acknowledgement and SHALL retry transport using that identity without consuming scanner/provider retry budgets. Existing scan errors SHALL retain their existing dispositions. Definitive stale rejection SHALL be recorded as a local stale/discard outcome with bounded cleanup, not success or infinite upload retry.
|
||||
|
||||
#### Scenario: Client restarts with a pending result
|
||||
- **WHEN** a client restarts before confirming server acceptance
|
||||
- **THEN** it recovers the pending result and retries/reconciles it before claiming replacement work for that occupied slot
|
||||
|
||||
#### Scenario: Scanner reports an existing deferred or terminal error
|
||||
- **WHEN** the existing scanner returns an error disposition
|
||||
- **THEN** the client/server handoff preserves that disposition and the server applies the existing queue policy rather than a transport-specific retry rule
|
||||
|
||||
### Requirement: Pre-bundle terminal reports are replay-safe and fenced
|
||||
Pre-bundle failure/release reports SHALL use the current issuance identity and existing outcome/accounting rules. Clients SHALL retain and retry the report identity until its authoritative outcome is acknowledged. Duplicate accepted reports SHALL return the original outcome without duplicate retry charges, counters, or quota/credit release; expired/superseded unaccepted reports SHALL NOT alter newer work.
|
||||
|
||||
#### Scenario: Failure or release reply is lost
|
||||
- **WHEN** the server commits a pre-bundle terminal disposition but its reply is lost and the client repeats the report
|
||||
- **THEN** the original disposition is acknowledged, remote quota/credits are reconciled once, and the client slot resolves once without an extra scanner retry charge
|
||||
|
||||
#### Scenario: Same worker reports failure for an old issuance
|
||||
- **WHEN** a worker reclaims a target under a new issuance and a delayed unaccepted terminal report for its expired issuance arrives
|
||||
- **THEN** the server rejects the stale report without changing the new issuance, its quota, or the target's current outcome
|
||||
|
||||
### Requirement: Worker observability derives from authoritative events
|
||||
The system SHALL expose safe per-worker unfinished/completed/failed/expired counts, last authenticated API contact, correlated target/source/reservation identities, issue/finish times, duration, and outcome/error category using existing logs and records. It SHALL distinguish result acceptance from later processing and MUST NOT infer online/offline status without a heartbeat.
|
||||
|
||||
#### Scenario: Duplicate or stale result arrives
|
||||
- **WHEN** a result is duplicated or rejected as stale
|
||||
- **THEN** a correlated safe event is recorded without inflating completed counts or exposing credentials/raw findings
|
||||
|
||||
#### Scenario: A valid worker is silent during a long scan
|
||||
- **WHEN** the worker makes no API request before its assignment deadline
|
||||
- **THEN** the admin view shows last contact and outstanding assignment state without declaring the worker dead solely from silence
|
||||
|
||||
### Requirement: Local validation is isolated and synthetic
|
||||
Implementation SHALL be developed and verified in the independent source-only workspace, with fresh test-owned storage and empty PostgreSQL initialized only for synthetic fixtures. Tests MUST NOT copy/restore/mount production data, reuse active runtime volumes/image tags, inherit production credentials/DSNs, or perform live discovery/provider verification. Windows/Linux real scanner parity and controlled-clock recovery tests SHALL precede release.
|
||||
|
||||
#### Scenario: Empty local end-to-end execution
|
||||
- **WHEN** a test runtime, Caddy, and clients are started for the worker scenario
|
||||
- **THEN** they use isolated project/image/volume/port/config identities, synthetic targets, mocked provider transports, and restricted external egress
|
||||
- **AND** existing production checkouts, containers, PGDATA, logs, and results are neither read as runtime inputs nor modified
|
||||
|
||||
#### Scenario: Unsafe inherited deployment defaults are detected
|
||||
- **WHEN** test setup would reuse the inherited production project, shared image tag, existing data volume, runtime bind mount, or live source configuration
|
||||
- **THEN** validation refuses to start runtime services until explicit isolation is established
|
||||
+75
@@ -0,0 +1,75 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Only authenticated edge routes are public
|
||||
The deployment SHALL expose HTTPS through Caddy only for `/api/v1/worker/*` and a configured random administrative prefix. The standalone dashboard, PostgreSQL, supervisor control, Caddy control API, and raw backend ports SHALL remain private. Unknown application paths SHALL return 404; an administrative URL MUST NOT replace authentication.
|
||||
|
||||
#### Scenario: Internet visitor requests backend or dashboard access
|
||||
- **WHEN** an unauthenticated visitor requests `/dashboard`, a normal `/admin` path, an unknown route, or a backend/control/database port
|
||||
- **THEN** no dashboard data or backend/control access is available
|
||||
|
||||
#### Scenario: Dashboard information appears inside admin
|
||||
- **WHEN** operational dashboard information is rendered in the administrative area
|
||||
- **THEN** every page, asset, and data request is protected by the same admin authentication boundary without publishing the standalone backend
|
||||
|
||||
### Requirement: Worker access has narrow device-scoped authority
|
||||
Every Worker API operation SHALL require a valid revocable opaque device token over certificate-validated HTTPS, scoped to its user/device assignments and quotas. The server SHALL store token hashes, not reusable plaintext tokens. Worker tokens MUST NOT authorize admin actions, another device's result mutation, database access, or generic commands.
|
||||
|
||||
#### Scenario: Invalid or revoked token submits work
|
||||
- **WHEN** an invalid/revoked device token claims a task or uploads a result
|
||||
- **THEN** the request is refused before task/result mutation and does not count toward the admin login jail
|
||||
|
||||
#### Scenario: Authorized device tries to finish somebody else's task
|
||||
- **WHEN** a device submits another device's reservation identity
|
||||
- **THEN** the request is refused without changing that reservation or disclosing task credentials
|
||||
|
||||
### Requirement: Admin routes require credentials and safe mutations
|
||||
The administrative prefix SHALL contain at least 128 bits of randomness and all administrative routes SHALL require login/password authentication using a supported password hash. Typed mutations SHALL enforce appropriate Origin/CSRF protection. The interface MUST NOT expose a shell or generic supervisor-command passthrough.
|
||||
|
||||
#### Scenario: Visitor knows the full administrative URL
|
||||
- **WHEN** a visitor accesses the correct administrative prefix without valid credentials
|
||||
- **THEN** no administrative page, data, asset, or mutation becomes available
|
||||
|
||||
#### Scenario: Cross-site request attempts a queue or token mutation
|
||||
- **WHEN** a browser submits an administrative mutation without valid same-origin/CSRF authorization
|
||||
- **THEN** the mutation is rejected even if browser-level password authentication is present
|
||||
|
||||
### Requirement: Two failed admin logins trigger an admin-only day ban
|
||||
Two actual invalid administrative credential submissions from one IP within ten minutes SHALL ban that IP from administrative routes for 24 hours. The ban SHALL persist across service restart, expire automatically, and support SSH/operator removal. Initial unauthenticated authentication challenges, Worker API failures, and unrelated requests MUST NOT count. Enforcement SHALL leave Worker API traffic from the same IP unaffected.
|
||||
|
||||
#### Scenario: Repeated invalid admin credentials
|
||||
- **WHEN** an IP submits a second invalid admin login within the ten-minute window
|
||||
- **THEN** subsequent admin requests from that IP are denied for 24 hours, including otherwise valid credentials, until expiry or operator unban
|
||||
|
||||
#### Scenario: Admin and worker share one NAT address
|
||||
- **WHEN** admin login failures trigger a ban for an IP also used by an authorized worker
|
||||
- **THEN** the worker can still claim/upload over HTTPS and administrative requests remain blocked
|
||||
|
||||
#### Scenario: Normal first visit receives an authentication challenge
|
||||
- **WHEN** a browser initially requests the protected area without sending credentials
|
||||
- **THEN** the normal challenge does not consume either of the two failed-login attempts
|
||||
|
||||
#### Scenario: Ban expires or operator recovers access
|
||||
- **WHEN** 24 hours elapse or an operator removes the ban through the documented SSH procedure
|
||||
- **THEN** credential-authenticated administrative access is restored without restarting or weakening Worker API authorization
|
||||
|
||||
### Requirement: IP identification and ban enforcement match the deployed edge
|
||||
The initial direct-DNS deployment SHALL derive client IP from the direct connection, ignoring untrusted forwarded headers. Any future proxy SHALL require an explicit trusted-proxy configuration and verification before enabling IP bans. The fail2ban action SHALL enforce the admin route boundary at Caddy rather than indiscriminately blocking shared port 443.
|
||||
|
||||
#### Scenario: Attacker supplies another address in a forwarded header
|
||||
- **WHEN** a direct client sends arbitrary `X-Forwarded-For` or equivalent headers with failed admin credentials
|
||||
- **THEN** another user's address is not selected for banning and the caller cannot evade its own ban
|
||||
|
||||
#### Scenario: Deployed ban is verified through Docker ingress
|
||||
- **WHEN** the admin jail applies its ban against the actual Caddy/Docker topology
|
||||
- **THEN** external administrative requests are denied and worker requests remain available, not merely a host firewall rule being present
|
||||
|
||||
### Requirement: Diagnostics and transport protect sensitive data
|
||||
All worker/admin transport SHALL use HTTPS without disabling certificate verification. Tokens, passwords, secret admin prefixes in routine access logs, task credentials, and raw findings SHALL be redacted or omitted from diagnostics. Admin responses SHALL disable caching and referrer disclosure and use a compatible restrictive CSP; production HTTPS SHALL use HSTS. Client disk/workspace encryption SHALL NOT be required.
|
||||
|
||||
#### Scenario: Auth or upload error is logged
|
||||
- **WHEN** authentication, payload validation, or upload handling fails
|
||||
- **THEN** logs contain only safe correlation/status/error metadata and enough redacted authentication outcome for the admin jail, not request credentials or raw result content
|
||||
|
||||
#### Scenario: Trusted client stores pending work locally
|
||||
- **WHEN** a worker downloads or persists a bundle before acknowledgement
|
||||
- **THEN** ordinary local storage is supported without an application encryption layer while HTTPS and post-acknowledgement cleanup remain enforced
|
||||
@@ -0,0 +1,49 @@
|
||||
## 1. Establish Safe Empty Test Isolation
|
||||
|
||||
- [x] 1.1 Replace unsafe inherited test launch defaults with explicit test-owned project/image/volume/port identities and refusal checks for production binds/volumes; do not run the copied default Compose or import overrides.
|
||||
- [x] 1.2 Reuse existing test harnesses to initialize empty PostgreSQL, scrub inherited secrets/DSNs/proxies before imports, disable live source/keycheck autostart, and provide synthetic targets/provider transports with external egress blocked.
|
||||
- [x] 1.3 Establish baseline synthetic scan/bundle/queue/keycheck fixtures and run the reviewed isolated unit selection before changing execution semantics; record relevant pre-existing failures separately.
|
||||
|
||||
## 2. Extract The Existing Scan Execution Boundary
|
||||
|
||||
- [x] 2.1 Trace `stage_claim`, `scan_target_result`, process authority, bundle candidate extraction, and Git/Docker exact-plan inputs; define the smallest DB-free job input using existing types/identities rather than a parallel task model.
|
||||
- [x] 2.2 Adapt one planned download/scan/bundle path to run on Windows/Linux without PostgreSQL or supervisor credentials while preserving `OwnedProcess`, native slot handling, source errors/dispositions, and structured Postman candidates.
|
||||
- [x] 2.3 Add centralized effective-config/plan delivery and protocol/scanner/detector-policy compatibility validation, supplying only task-needed credentials and forbidding arbitrary commands or client provider overrides.
|
||||
- [x] 2.4 Compare local and DB-free execution on existing source fixtures, including Git/Docker coverage, Postman evidence, and Xai/ZAI custom-detector direction regressions; preserve detailed keycheck exclusively on the server.
|
||||
|
||||
## 3. Extend Existing Admission And Recovery
|
||||
|
||||
- [x] 3.1 Add only necessary user/device token-hash/quota bindings and remote metadata on existing reservations, with revocation and no second queue or independent remote-job state machine.
|
||||
- [x] 3.2 Implement authenticated one-task claims through existing admission with atomic per-user caps across devices, existing capacity limits, stable request identity, ambiguous-claim reconciliation, and bounded empty/capacity polling.
|
||||
- [x] 3.3 Apply configurable server-clock fixed expiry (default 24 hours) to remote ownership and dependent target/plan/blob leases, separating it from local PID recovery and scanner timeouts without heartbeat or renewal endpoints.
|
||||
- [x] 3.4 Wire periodic expired-assignment recovery into runtime maintenance using existing refund/requeue accounting; permit same-worker reclaim, avoid new error/retry/blacklist policies, and release quota exactly once.
|
||||
- [x] 3.5 Test concurrent quota admission, lower-cap behavior, lost claim replies, restart recovery, fixed-clock expiry, dependent plan leases, and stale ownership fencing against isolated PostgreSQL.
|
||||
|
||||
## 4. Add Durable Canonical Bundle Transport
|
||||
|
||||
- [x] 4.1 Receive `.trb` binary streams into bounded server-owned partial files, enforcing existing capacity/size limits, upload timeouts, ownership, expiry, codec/path/identity validation, and content hash.
|
||||
- [x] 4.2 Reuse durable atomic publication and `mark_result_bundle_ready` before acknowledgement; reconcile upload/recovery races and reject stale/conflicting bodies without affecting current plan coverage or credits.
|
||||
- [x] 4.3 Preserve accepted identity/digest/receipt in existing records independently of spool cleanup and return the same receipt after ingestion, cleanup, restart, and expiry; reuse existing ingester/projector/candidate/keycheck behavior and exclude ready bundles from worker expiry.
|
||||
- [x] 4.4 Preserve current scan-failure dispositions and pre-bundle infrastructure release paths with issuance-fenced idempotent terminal-report replay; test lost/repeated/stale reports, single slot/quota/credit resolution, interrupted/invalid/conflicting uploads, and publication/ready/commit crashes.
|
||||
|
||||
## 5. Build The Minimal Client Loop
|
||||
|
||||
- [x] 5.1 Implement server-URL/token/N bootstrap and one claim per free slot, with node-local execution containment and no independent source/provider configuration, client keycheck, heartbeat, batching, or updater.
|
||||
- [x] 5.2 Persist assignment identity and pending bundles, recover them on restart, retry transport without scan retry charges, free work slots only after authoritative resolution, and handle definitive stale rejection with explicit bounded cleanup.
|
||||
- [x] 5.3 Package pinned scanner/config assets for a Windows portable client and Linux client/container; verify certificate-validating HTTPS and avoid secret/raw-finding logs without requiring local encryption.
|
||||
- [x] 5.4 Exercise real synthetic scans and complete bundle round trips on both Windows and Linux with N greater than one, restart during pending upload, server outage, and subsequent recovery; do not substitute mocks for cross-platform scanner verification.
|
||||
|
||||
## 6. Restrict Edge Access And Add Minimal Administration
|
||||
|
||||
- [x] 6.1 Wire API/admin into the single runtime and separate Caddy edge, publishing only authenticated worker routes and a random admin prefix; keep dashboard/backend/database/control ports private and remote admission disabled until configured.
|
||||
- [x] 6.2 Add device-scoped token authorization and hash-based admin password authentication with protected assets, typed CSRF/Origin-checked mutations, no-store/same-origin-referrer/CSP/production-HSTS headers, and redacted logs.
|
||||
- [x] 6.3 Add a host fail2ban admin jail for two actual bad logins in ten minutes and a 24-hour persisted ban; implement validated admin-only Caddy denylist updates, direct-IP handling, automatic expiry, and documented SSH unban without blocking worker traffic.
|
||||
- [x] 6.4 Build only necessary admin user/device-token/quota and existing queue controls plus read-only worker counts, durations, outcomes, and last-contact summaries from existing records; do not add online/offline guesses or a telemetry store.
|
||||
- [x] 6.5 Verify unknown/private routes, revoked/wrong-device tokens, cross-site mutations, absent-credentials challenges, spoofed forwarded headers, actual two-failure bans/restart/unban, and an authorized worker sharing the banned admin IP through the isolated edge.
|
||||
|
||||
## 7. Complete Regression Gates And Handoff
|
||||
|
||||
- [x] 7.1 Run the synthetic empty-database end-to-end flow through claim, real scan, complete bundle, ingestion, projection, and mocked detailed server keycheck; compare normalized output/dispositions with the existing local path.
|
||||
- [x] 7.2 Run combined disconnect/restart/expiry/reissue/upload races with controlled clocks; verify single authoritative acceptance, intact plan coverage, no credit leaks, and no duplicate worker statistics.
|
||||
- [x] 7.3 Verify test manifests, build context, logs, artifacts, and cleanup exclude production state/credentials and that active production containers/volumes were untouched; retain only test-owned failure evidence.
|
||||
- [x] 7.4 Document isolated run commands, client bootstrap, quota/deadline tuning, token revocation, admin unban, accepted-custody semantics, known 24-hour recovery trade-offs, and later rollout/drain rollback; validate OpenSpec and leave unrelated changes unarchived.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-06-02
|
||||
@@ -0,0 +1,80 @@
|
||||
## Context
|
||||
|
||||
The scanner is already organized around independent sources managed by `console_runner.py` and `supervisor.py`. Each source discovers targets, writes them to per-source queue files, scans targets through TruffleHog, records results in JSONL and SQLite, and can stop pagination when consecutive pages contain only known targets. Existing package sources already download and extract npm/PyPI artifacts before running TruffleHog filesystem scans.
|
||||
|
||||
Postman artifacts fit this model as filesystem scan targets, but they need separate discovery and enrichment. Public Postman web search is comparatively fragile, while GitHub code search exposes many real `*.postman_collection.json` and `*.postman_environment.json` files. npm and PyPI can also expose Postman artifacts during the package extraction windows that already exist.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Add a first-class `postman` source with standard queue, checked, supervisor, dashboard, and database behavior.
|
||||
- Discover Postman collection/environment JSON from GitHub code search using the existing GitHub auth pool.
|
||||
- Support initial backfill over up to the GitHub Search API result cap per query and ongoing tail scans over recently indexed pages.
|
||||
- Filter discovered GitHub code artifacts by last file commit age so stale files can be skipped during backfill.
|
||||
- Rotate across multiple GitHub tokens and pause only when all usable tokens are rate-limited.
|
||||
- Cache discovered Postman artifacts durably so package-derived artifacts survive temp directory cleanup.
|
||||
- Harvest Postman artifacts from npm and PyPI extraction flows without disrupting existing package scans.
|
||||
- Enrich findings with Postman-specific context derived from auth configuration, headers, query params, request bodies, environment variables, and endpoint hosts.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Scraping Postman's own web application in the first implementation.
|
||||
- Replacing TruffleHog detectors with custom regex-only detection.
|
||||
- Adding new keychecker services as part of this change.
|
||||
- Scanning private Postman workspaces through the Postman API.
|
||||
|
||||
## Decisions
|
||||
|
||||
1. Use GitHub code search as the primary discovery channel.
|
||||
|
||||
GitHub code search has authenticated API support, predictable pagination, and high-quality results for `filename:postman_collection.json <query>` and `filename:postman_environment.json <query>`. Public Postman web pages and `postman.com` links are lower-yield and more likely to change without notice. Direct Postman URL discovery can be added later as an additional provider without changing the scanner contract.
|
||||
|
||||
2. Treat Postman artifacts as durable filesystem scan targets.
|
||||
|
||||
The source will download or copy each discovered artifact into a runtime Postman cache and scan a temporary directory containing the cached JSON. This reuses the existing TruffleHog filesystem path and avoids keeping npm/PyPI extraction directories alive.
|
||||
|
||||
3. Identify GitHub-discovered targets by `repo:path:sha`.
|
||||
|
||||
The same file at the same SHA must not be rescanned, while a new SHA for the same path must be queued again. This matches the existing queue/checked model and makes `stop_on_seen_pages` useful for tail scans.
|
||||
|
||||
4. Identify package-harvested targets by content hash.
|
||||
|
||||
npm/PyPI packages can contain duplicate Postman artifacts across versions or package names. A SHA-256 content hash provides stable dedupe and allows different origins to point to the same cached artifact without rescanning identical content.
|
||||
|
||||
5. Filter GitHub code artifacts by path commit age after discovery.
|
||||
|
||||
The code search response does not include reliable file modification dates. The implementation will query the latest commit for each `repo:path` and skip artifacts older than `max_file_age_days`. This costs extra core API requests but keeps the code search query simple and reliable.
|
||||
|
||||
6. Use a source-local GitHub token pool for discovery.
|
||||
|
||||
Postman discovery may make many GitHub requests in one cycle. A per-request token pool can rotate across all configured GitHub auth entries, cool down only the token that failed, and sleep when no token remains available. This is more efficient than one token per source cycle.
|
||||
|
||||
7. Keep Postman enrichment separate from detection.
|
||||
|
||||
TruffleHog remains responsible for finding candidate secrets. Postman enrichment will add context and confidence by correlating findings with request auth, headers, variables, endpoints, and placeholder detection. This avoids increasing false positives from regex-only scans.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- GitHub code search is rate-limited to roughly 10 requests per minute per token -> throttle to a configurable safe RPM and rotate across the auth pool.
|
||||
- Commit-age filtering adds extra core API requests -> make `max_file_age_days` configurable and cache commit metadata per `repo:path:sha` within a cycle.
|
||||
- Broad queries such as `ai` can hit the 1000-result search cap and include noisy results -> use a Postman-specific query list and allow per-source query tuning.
|
||||
- Package harvesting adds small overhead during npm/PyPI scans -> limit file walking to reasonable extensions, max file size, and known Postman filename patterns.
|
||||
- Postman variables often contain placeholders rather than live secrets -> classify placeholders separately and keep TruffleHog verification/keycheckers as the authority for live/dead status.
|
||||
- All tokens may become unavailable -> sleep until the earliest known reset time, or a configured fallback such as 30 minutes when no reset is known.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add the `postman` source disabled by default in `config.yaml`.
|
||||
2. Add queue/dashboard/database support for `postman` without changing existing source behavior.
|
||||
3. Run a small verification cycle with `pages: 1`, `per_page: 10`, and `max_targets` set.
|
||||
4. Run the one-time backfill with `pages: 10`, `per_page: 100`, `stop_on_seen_pages: false`, and `max_file_age_days: 365`.
|
||||
5. Switch the source to tail mode with `pages: 1-3` and `stop_on_seen_pages: true`.
|
||||
6. Enable npm/PyPI harvesting after the base Postman source is verified.
|
||||
|
||||
Rollback is to disable `sources.postman.enabled`, leave its queues/cache intact, and continue running existing sources unchanged.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether the default backfill query list should include broad terms like `ai`, or keep only higher-intent terms such as `openai`, `anthropic`, `gemini`, `llm`, `rag`, and `agent`.
|
||||
- Whether package-harvested artifacts should always be enqueued, or only when they include auth/secret-related markers.
|
||||
@@ -0,0 +1,32 @@
|
||||
## Why
|
||||
|
||||
Public Postman collections and environments are a high-signal source for leaked API credentials because they often preserve request auth settings, headers, variables, and example payloads close to real API usage. The existing scanner already supports multi-source discovery, queues, TruffleHog filesystem scans, token rotation, and observability, so adding Postman can reuse the current architecture while expanding coverage beyond repositories, packages, containers, and HuggingFace Spaces.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a new `postman` source that discovers, queues, scans, and records Postman collection/environment artifacts.
|
||||
- Seed Postman targets from GitHub code search using public `*.postman_collection.json` and `*.postman_environment.json` files.
|
||||
- Support a one-time backfill mode that scans up to the GitHub Search API result limit per query while filtering out artifacts older than a configured age window.
|
||||
- Support a daily tail mode that fetches recently indexed pages and stops early when all targets on consecutive pages are already known.
|
||||
- Use the configured GitHub auth pool for Postman discovery, rotating across tokens and sleeping when all tokens are rate-limited.
|
||||
- Add durable Postman artifact caching so targets discovered from GitHub, npm, and PyPI can be scanned after temporary extraction directories are removed.
|
||||
- Harvest Postman artifacts from npm and PyPI packages during existing package extraction flows and enqueue them into the shared Postman queue.
|
||||
- Add Postman-aware result enrichment that classifies credentials using TruffleHog findings plus Postman auth/header/query/body/environment context.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `postman-source`: Discovery, queueing, scanning, caching, and enrichment for Postman collection and environment artifacts.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected scanner paths: `app/scanner.py`, `app/console_runner.py`, `app/scanner_db.py`, `app/dashboard.py`, and `app/config.yaml`.
|
||||
- Adds runtime files under `runtime/queues/` for `todo_postman.txt` and `checked_postman.txt`.
|
||||
- Adds durable artifact storage under a runtime Postman cache directory.
|
||||
- Uses existing GitHub auth pools from `secrets.yaml`; no new secret format is required for GitHub discovery.
|
||||
- Uses existing TruffleHog filesystem scanning and keychecker follow-up flows; no breaking changes to current sources are expected.
|
||||
@@ -0,0 +1,158 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Postman source registration
|
||||
The system SHALL provide a first-class `postman` source that can be configured, selected, supervised, queued, scanned, and displayed consistently with existing scanner sources.
|
||||
|
||||
#### Scenario: Configured Postman source is selectable
|
||||
- **WHEN** a config contains `sources.postman.enabled: true` and the runner is invoked with `--source postman`
|
||||
- **THEN** the runner SHALL execute only the Postman source cycle using Postman source settings
|
||||
|
||||
#### Scenario: Postman queue files are used
|
||||
- **WHEN** the Postman source prepares targets
|
||||
- **THEN** the system SHALL use `todo_postman.txt` and `checked_postman.txt` under the configured queue directory
|
||||
|
||||
### Requirement: GitHub code search discovery
|
||||
The Postman source SHALL discover public Postman artifacts from GitHub code search using configured queries and artifact kinds.
|
||||
|
||||
#### Scenario: Collection files are discovered
|
||||
- **WHEN** `search_kinds` includes `collection` and the query is `openai`
|
||||
- **THEN** discovery SHALL search GitHub code for `filename:postman_collection.json openai`
|
||||
|
||||
#### Scenario: Environment files are discovered
|
||||
- **WHEN** `search_kinds` includes `environment` and the query is `openai`
|
||||
- **THEN** discovery SHALL search GitHub code for `filename:postman_environment.json openai`
|
||||
|
||||
#### Scenario: Search pagination is bounded
|
||||
- **WHEN** `pages` is `10` and `per_page` is `100`
|
||||
- **THEN** discovery SHALL request no more than 1000 search results per query and kind
|
||||
|
||||
### Requirement: GitHub auth pool rotation
|
||||
The Postman GitHub discovery flow SHALL use all available tokens from the configured GitHub auth pool before sleeping for rate limits.
|
||||
|
||||
#### Scenario: Token rotates per GitHub request
|
||||
- **WHEN** multiple GitHub auth entries are available
|
||||
- **THEN** GitHub code search, commit lookup, and content download requests SHALL rotate across available tokens
|
||||
|
||||
#### Scenario: One token is rate-limited
|
||||
- **WHEN** a GitHub request returns a primary or secondary rate limit for the current token
|
||||
- **THEN** the system SHALL mark only that token unavailable until its reset time or configured cooldown and continue with another available token
|
||||
|
||||
#### Scenario: All tokens are unavailable
|
||||
- **WHEN** every configured GitHub token is rate-limited or temporarily unavailable
|
||||
- **THEN** the source SHALL sleep until the earliest known reset time, or for the configured fallback cooldown when no reset time is known
|
||||
|
||||
#### Scenario: Token is invalid
|
||||
- **WHEN** a GitHub request returns an authentication-invalid response for a token
|
||||
- **THEN** the system SHALL exclude that token from the current cycle and report the authentication failure without marking other tokens invalid
|
||||
|
||||
### Requirement: Backfill freshness filtering
|
||||
The Postman source SHALL support a backfill mode that can scan deep code search pages while skipping GitHub artifacts older than a configured file age.
|
||||
|
||||
#### Scenario: Recent file is queued
|
||||
- **WHEN** a GitHub code search result has a latest path commit within `max_file_age_days`
|
||||
- **THEN** the target SHALL be eligible for queueing
|
||||
|
||||
#### Scenario: Old file is skipped
|
||||
- **WHEN** a GitHub code search result has a latest path commit older than `max_file_age_days`
|
||||
- **THEN** the target SHALL not be queued and SHALL be counted as skipped by freshness filtering
|
||||
|
||||
#### Scenario: Freshness filtering is disabled
|
||||
- **WHEN** `max_file_age_days` is `0`
|
||||
- **THEN** the source SHALL not perform path commit age filtering
|
||||
|
||||
### Requirement: Tail mode early stop
|
||||
The Postman source SHALL support ongoing tail scans that stop pagination after consecutive known pages.
|
||||
|
||||
#### Scenario: Known page increments stop counter
|
||||
- **WHEN** `stop_on_seen_pages` is enabled and every normalized target on a fetched page already exists in `todo_postman.txt` or `checked_postman.txt`
|
||||
- **THEN** the source SHALL count that page as known
|
||||
|
||||
#### Scenario: Tail pagination stops
|
||||
- **WHEN** the known page count reaches `seen_page_threshold` after `min_pages_before_stop`
|
||||
- **THEN** the source SHALL stop fetching additional pages for that query and artifact kind
|
||||
|
||||
### Requirement: Postman target identity and deduplication
|
||||
The system SHALL normalize Postman targets so identical artifacts are not rescanned while changed artifacts are scanned again.
|
||||
|
||||
#### Scenario: GitHub target identity includes SHA
|
||||
- **WHEN** a Postman target is discovered from GitHub code search
|
||||
- **THEN** its normalized target SHALL include source, repository, path, and file SHA
|
||||
|
||||
#### Scenario: GitHub file changes
|
||||
- **WHEN** the same GitHub repository and path is discovered with a new SHA
|
||||
- **THEN** the system SHALL treat it as a new Postman target
|
||||
|
||||
#### Scenario: Package target identity uses content hash
|
||||
- **WHEN** a Postman artifact is harvested from npm or PyPI
|
||||
- **THEN** its normalized target SHALL include the artifact content SHA-256 hash
|
||||
|
||||
### Requirement: Durable Postman artifact cache
|
||||
The system SHALL store discovered Postman artifact content in durable runtime cache before scanning.
|
||||
|
||||
#### Scenario: GitHub content is cached
|
||||
- **WHEN** a GitHub code search target is queued for scanning
|
||||
- **THEN** the system SHALL download the artifact content and store it under the configured Postman cache directory
|
||||
|
||||
#### Scenario: Package content is cached before cleanup
|
||||
- **WHEN** npm or PyPI extraction finds a Postman artifact
|
||||
- **THEN** the system SHALL copy the artifact into the durable Postman cache before the extraction directory is removed
|
||||
|
||||
#### Scenario: Cache size is constrained
|
||||
- **WHEN** an artifact exceeds the configured maximum Postman artifact size
|
||||
- **THEN** the system SHALL skip the artifact and record a bounded error or skip reason
|
||||
|
||||
### Requirement: Postman artifact scanning
|
||||
The Postman source SHALL scan cached Postman collection and environment artifacts with TruffleHog filesystem scanning.
|
||||
|
||||
#### Scenario: Cached artifact is scanned
|
||||
- **WHEN** a Postman target points to a cached JSON artifact
|
||||
- **THEN** the scanner SHALL run TruffleHog against a temporary filesystem directory containing that artifact
|
||||
|
||||
#### Scenario: Findings are persisted
|
||||
- **WHEN** TruffleHog reports findings for a Postman target
|
||||
- **THEN** the system SHALL persist findings to existing JSONL outputs and scanner database tables with source `postman`
|
||||
|
||||
#### Scenario: Scan finishes
|
||||
- **WHEN** a Postman target scan completes with findings, errors, skipped status, or clean status
|
||||
- **THEN** the target SHALL be moved from `todo_postman.txt` to `checked_postman.txt`
|
||||
|
||||
### Requirement: npm and PyPI Postman harvesting
|
||||
The npm and PyPI source flows SHALL harvest Postman artifacts discovered during existing package extraction and enqueue them for the Postman source.
|
||||
|
||||
#### Scenario: npm package contains collection
|
||||
- **WHEN** an extracted npm package contains a file matching `*.postman_collection.json`
|
||||
- **THEN** the system SHALL cache the file and enqueue a Postman target with npm package origin metadata
|
||||
|
||||
#### Scenario: PyPI package contains environment
|
||||
- **WHEN** an extracted PyPI artifact contains a file matching `*.postman_environment.json`
|
||||
- **THEN** the system SHALL cache the file and enqueue a Postman target with PyPI package origin metadata
|
||||
|
||||
#### Scenario: Existing package scan continues
|
||||
- **WHEN** Postman harvesting fails for one package artifact
|
||||
- **THEN** the original npm or PyPI scan SHALL still complete and record the harvesting failure without failing unrelated package scanning
|
||||
|
||||
### Requirement: Postman-aware enrichment
|
||||
The system SHALL enrich Postman findings with contextual classification derived from Postman structure without replacing TruffleHog detection.
|
||||
|
||||
#### Scenario: Header credential is classified
|
||||
- **WHEN** a finding appears in a Postman request header such as `Authorization` or `x-api-key`
|
||||
- **THEN** enrichment SHALL record the context location and infer credential kind from header type, value shape, and endpoint host when possible
|
||||
|
||||
#### Scenario: Environment variable is classified
|
||||
- **WHEN** a finding appears in a Postman environment variable value
|
||||
- **THEN** enrichment SHALL record the variable name and classify provider or credential kind when supported by value shape or associated request endpoints
|
||||
|
||||
#### Scenario: Placeholder is detected
|
||||
- **WHEN** a Postman value is a placeholder such as `{{API_KEY}}`, `<api_key>`, `YOUR_API_KEY`, `example`, or `changeme`
|
||||
- **THEN** enrichment SHALL classify it as placeholder or low confidence rather than a live secret
|
||||
|
||||
### Requirement: Observability for Postman source
|
||||
The system SHALL expose Postman source activity through existing logs, queue counts, source cycle metrics, target scan records, findings, errors, and dashboard views.
|
||||
|
||||
#### Scenario: Source cycle is recorded
|
||||
- **WHEN** a Postman source cycle runs
|
||||
- **THEN** the scanner database SHALL record source cycle metrics including fetched, queued, scanned, found, error, skipped, and queue counts
|
||||
|
||||
#### Scenario: Dashboard shows Postman queues
|
||||
- **WHEN** Postman queue files exist
|
||||
- **THEN** the dashboard SHALL include Postman queue counts in the current queues view
|
||||
@@ -0,0 +1,69 @@
|
||||
## 1. Source Wiring
|
||||
|
||||
- [x] 1.1 Add `postman` to CLI `--source` and `--platform` choices and source-to-platform mapping.
|
||||
- [x] 1.2 Add `postman` queue file support through the existing `queue_files_for_args`, prepare, and mark-checked flow.
|
||||
- [x] 1.3 Add `postman` to supervisor source configuration and dashboard source lists.
|
||||
- [x] 1.4 Add disabled-by-default `sources.postman` configuration with GitHub auth pool, search kinds, backfill/tail controls, cache path, and rate-limit settings.
|
||||
|
||||
## 2. Target Model And Cache
|
||||
|
||||
- [x] 2.1 Define Postman target JSON formats for GitHub code search, npm package, PyPI package, local cache, and future URL targets.
|
||||
- [x] 2.2 Implement Postman target parsing and normalization in `console_runner.py` and `scanner_db.py`.
|
||||
- [x] 2.3 Implement durable Postman cache path resolution under the configured runtime directory.
|
||||
- [x] 2.4 Implement safe cache writes with SHA-256 content hashing, max artifact size checks, and origin metadata preservation.
|
||||
|
||||
## 3. GitHub Code Search Discovery
|
||||
|
||||
- [x] 3.1 Implement GitHub code search queries for `filename:postman_collection.json <query>` and `filename:postman_environment.json <query>` based on `search_kinds`.
|
||||
- [x] 3.2 Implement bounded pagination using configured `pages` and `per_page`, respecting the GitHub 1000-result search cap.
|
||||
- [x] 3.3 Convert GitHub code search items into Postman target JSON containing repository, path, SHA, kind, API URL, and HTML URL.
|
||||
- [x] 3.4 Implement latest path commit lookup for `max_file_age_days` filtering.
|
||||
- [x] 3.5 Integrate existing known-page early stop behavior for Postman tail scans.
|
||||
|
||||
## 4. GitHub Token Pool And Rate Limits
|
||||
|
||||
- [x] 4.1 Build a source-local GitHub token pool from configured `auth_pool` entries and fallback token settings.
|
||||
- [x] 4.2 Rotate tokens per GitHub code search, commit lookup, and content download request.
|
||||
- [x] 4.3 Mark only the failing token unavailable on primary rate limit, secondary rate limit, auth invalid, or auth forbidden responses.
|
||||
- [x] 4.4 Sleep until earliest known reset time, or configured fallback cooldown, when all GitHub tokens are unavailable.
|
||||
- [x] 4.5 Record token cooldown status in the source runtime state without exposing token values in logs or database snapshots.
|
||||
|
||||
## 5. Postman Artifact Scanning
|
||||
|
||||
- [x] 5.1 Implement Postman content download from GitHub Contents API and cache it before scanning.
|
||||
- [x] 5.2 Implement `scan_postman_target()` to stage cached JSON in a temporary directory and run TruffleHog filesystem scanning.
|
||||
- [x] 5.3 Add Postman branch to `scan_targets_batch()` and pass timeout, detectors, excluded detectors, and verification flags.
|
||||
- [x] 5.4 Preserve nearby file context and apply existing noisy finding filters to Postman scan results.
|
||||
- [x] 5.5 Ensure Postman findings, errors, skipped reasons, and clean scans are persisted through existing JSONL and scanner database writes.
|
||||
|
||||
## 6. npm And PyPI Harvesting
|
||||
|
||||
- [x] 6.1 Add a Postman artifact finder for extracted package directories that matches collection and environment filename patterns.
|
||||
- [x] 6.2 Cache npm package Postman artifacts before package temp directory cleanup and attach npm origin metadata.
|
||||
- [x] 6.3 Cache PyPI package Postman artifacts before package temp directory cleanup and attach PyPI origin metadata.
|
||||
- [x] 6.4 Enqueue harvested package artifacts into `todo_postman.txt` after package scan batches without failing the original package scan.
|
||||
- [x] 6.5 Deduplicate harvested package artifacts by content hash before enqueueing.
|
||||
|
||||
## 7. Postman-Aware Enrichment
|
||||
|
||||
- [x] 7.1 Parse Postman collection and environment JSON into request, auth, header, query, body, and variable context maps.
|
||||
- [x] 7.2 Correlate TruffleHog finding locations or nearby context with Postman context maps.
|
||||
- [x] 7.3 Classify credential kind and provider using DetectorName, value shape, auth/header type, variable name, and endpoint host.
|
||||
- [x] 7.4 Detect common placeholders and assign placeholder or low-confidence classification.
|
||||
- [x] 7.5 Persist enrichment fields using existing finding enrichment/database columns where possible.
|
||||
|
||||
## 8. Observability And Configuration
|
||||
|
||||
- [x] 8.1 Add Postman source cycle metrics, queue snapshots, target scan records, findings, and errors to existing database flows.
|
||||
- [x] 8.2 Add Postman queue counts to dashboard current queues and source health views.
|
||||
- [x] 8.3 Add redaction coverage for Postman/GitHub auth pool settings in config snapshots and logs.
|
||||
- [x] 8.4 Add backfill-friendly and tail-friendly config examples in `config.yaml` comments.
|
||||
|
||||
## 9. Verification
|
||||
|
||||
- [x] 9.1 Run a small Postman GitHub discovery cycle with `pages: 1`, `per_page: 10`, and `max_targets` set.
|
||||
- [x] 9.2 Re-run the same cycle and verify duplicate targets are skipped through `todo_postman.txt` and `checked_postman.txt`.
|
||||
- [x] 9.3 Verify all-token rate-limit fallback with a simulated or controlled token-unavailable state.
|
||||
- [x] 9.4 Verify npm and PyPI harvesting using a package fixture containing collection and environment JSON files.
|
||||
- [x] 9.5 Verify database and dashboard visibility for Postman source cycles, queues, target scans, findings, and errors.
|
||||
- [x] 9.6 Run `openspec status --change add-postman-source` and ensure all implementation tasks are complete before archive.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-19
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,155 @@
|
||||
## Context
|
||||
|
||||
The current runtime combines provider discovery, queue admission, PostgreSQL claiming, and local scanner execution in the same `console_runner.py` source cycle. Supervisor can pause a child in memory, but that state is lost on restart and does not fence concurrent Worker API claims. The remote protocol supports GitHub and GitLab exact-Git assignments, while DockerHub and HuggingFace exist only as local scan paths. The typed admin UI can mutate worker users, devices, and queue rows, but routine runtime control and file changes still require SSH.
|
||||
|
||||
The deployment is deliberately split across trust boundaries. Worker API/admin runs unprivileged in the read-only runtime container; either the standalone Truf edge or an existing root-owned host Caddy plus a loopback Truf edge owns public routing and injects a private edge marker; PostgreSQL is the durable queue authority; Supervisor exposes an authenticated loopback control protocol; systemd/Docker lifecycle control remains on the host. The exact ingress profile is root-installed policy, not runtime or request input. The design must retain those boundaries, keep uploads available during operational pauses, avoid a local server scanner, and never expose a shell or unrestricted host path.
|
||||
|
||||
The default distributed core profile changes to exactly `gitlab`, `dockerhub`, and `huggingface`. Existing protocol-1 Git assignments may still be in flight when the new server is deployed, so result compatibility and migration ordering matter.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Run provider discovery as server-side producer processes that only search, normalize, and enqueue.
|
||||
- Persist and atomically enforce independent discovery pause, dispatch pause, and drain controls.
|
||||
- Complete remote claim, scan, upload, ingestion, and projection for GitLab, DockerHub, and HuggingFace.
|
||||
- Prove the first new-source path with a bounded public DockerHub immutable-digest canary.
|
||||
- Provide typed, server-rendered admin operations for runtime state, controls, Supervisor, logs, configuration, plaintext secrets, managed files, and audit history.
|
||||
- Apply configuration and secrets through validated candidates, backups, coordinated restart, health checks, and automatic rollback.
|
||||
- Keep operator identity, operation status, and audit records durable across admin/runtime restarts.
|
||||
- Support exact standalone and shared-host ingress profiles while preserving route confinement, loopback-only shared ingress, and the closed host-agent request schema.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Local scanning or TruffleHog execution on the server.
|
||||
- Arbitrary shell commands, arbitrary Supervisor command strings, Docker socket access, or browsing host root.
|
||||
- PostgreSQL data-file access through the file page.
|
||||
- Private Docker registry credential delivery or Docker layer-plan transport in the first rollout.
|
||||
- Private HuggingFace credential delivery in the first canary; server-side discovery credentials remain supported.
|
||||
- A client-side single-page application or storage of configuration/secrets in browser persistence.
|
||||
- Replacing PostgreSQL queue, reservation, bundle, or projection authority.
|
||||
- Managing, restarting, or reconfiguring unrelated host Caddy sites, X-UI, or other shared-host services.
|
||||
- Runtime-selected topology, arbitrary ingress ports/upstreams, or wildcard/private-interface shared-edge binding.
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. Separate discovery into an explicit process role
|
||||
|
||||
Supervisor will launch a `discovery-producer` role for each enabled core source instead of launching the existing combined source cycle. The role is part of the authenticated runtime bootstrap identity and is visible in structured Supervisor state.
|
||||
|
||||
`console_runner.py` will expose a discovery-only cycle that performs provider requests, normalization, DockerHub retry/tag resolution where applicable, and idempotent queue admission. It will finish the source-cycle record with no scan requests and cannot call scan option preparation, scan-slot acquisition, target claiming, bundle staging, or scanner execution. Discovery runs according to its interval even when pending queue backlog exists.
|
||||
|
||||
This is preferred over a mutable `enqueue_only` flag on the existing combined cycle because a distinct bootstrap role and call graph make accidental local scanner entry testable and fail-closed. The old combined path may remain for non-server/local workflows, but the server profile will not launch it.
|
||||
|
||||
### 2. Make PostgreSQL the durable control authority
|
||||
|
||||
An additive singleton control row will hold:
|
||||
|
||||
- monotonically increasing `revision`;
|
||||
- explicit `discovery_paused` and `dispatch_paused` flags;
|
||||
- `drain_state` (`normal`, `draining`, or `drained`);
|
||||
- actor, operation ID, and update timestamps.
|
||||
|
||||
Mutations use compare-and-swap on the expected revision and append an audit event in the same transaction. Explicit pause flags remain independent; entering drain overlays both effective gates, and cancelling drain does not clear pauses that the operator set explicitly.
|
||||
|
||||
Discovery producers check the effective discovery gate before provider I/O and enforce it again in the same transaction as discovery queue admission. Source retry and Docker tag-resolution claims are also gated. Generic enqueue functions used by ingestion and maintenance are not globally disabled.
|
||||
|
||||
`reserve_and_claim_target()` enforces the effective dispatch gate inside its existing claim transaction. This closes the race between an API pre-check and target reservation. Authentication, status, terminal reports, assignment expiry, result upload, receipt replay, `mark_result_bundle_ready()`, and ingestion remain available while dispatch is paused or draining.
|
||||
|
||||
Drain is complete when there are no live remote assignments and no accepted result bundles that have not reached database commit. Pending target backlog, projection work, and keycheck work do not prevent `drained`; those workers can safely resume after restart. A reconciler advances `draining` to `drained` from database state.
|
||||
|
||||
This is preferred over stopping Worker API or Supervisor-only pause state because in-flight workers must retain their upload path and all enforcement must survive process restart.
|
||||
|
||||
### 3. Generalize assignments through source adapters and protocol 2
|
||||
|
||||
The Git-specific assignment builder will become an adapter registry. Each adapter declares the canonical queue source, worker scan platform, planning kind, package capability, execution-snapshot validator, and claim/reconciliation behavior:
|
||||
|
||||
| Queue source | Worker platform | Planning kind | Initial execution |
|
||||
| --- | --- | --- | --- |
|
||||
| `gitlab` | `gitlab` | `exact_git_v1` | Existing exact commit/snapshot path |
|
||||
| `dockerhub` | `docker` | `docker_direct_v1` | Public immutable digest reference |
|
||||
| `huggingface` | `huggingface` | `huggingface_space_v1` | Public Space identifier |
|
||||
|
||||
The assignment API paths remain stable, but worker protocol/package compatibility advances to version 2 and manifests advertise explicit source/planning capabilities. Protocol-1 packages receive no new claims after cutover. Status, terminal report, upload, receipt replay, and immutable snapshot reconciliation remain available for already-issued protocol-1 assignments through their fixed expiry.
|
||||
|
||||
Source selection will try other eligible configured sources when one queue has no claimable target instead of permanently choosing one source by request-ID modulo. Every successful claim still binds one fenced reservation, fixed expiry, device identity, immutable execution snapshot, and result-bundle identity.
|
||||
|
||||
The first DockerHub canary uses an image resolved to an immutable digest and direct worker execution. Digest resolution is assignment planning, not proof that the worker can access the registry. The worker is the final access check and reports an inaccessible image using the bounded provider-failure result contract. The assignment does not serialize `DockerRegistryAuth`, process-local monotonic deadlines, server blob leases, or a Docker layer plan. The first HuggingFace canary follows the same worker-authoritative access model. Discovery drops Spaces that its existing provider response explicitly marks private, protected, gated, or disabled; the worker reports an inaccessible repository as non-retryable. Existing GitLab credential behavior remains, but provider discovery credentials are not assignment fields.
|
||||
|
||||
Source adapters SHALL remain minimal. The server validates canonical target and assignment shape, performs only planning needed to identify the target, and leaves real provider access to the worker. A new per-target server access probe, durable public-access proof, proof freshness schema, broad child-environment credential scrubbing, credential sandbox, or post-hoc redaction pipeline is not implied by the credential non-transfer rule. Any such mechanism requires separate operator approval and an explicit OpenSpec requirement and task before implementation. Existing defensive code is not precedent for adding the same machinery to another source.
|
||||
|
||||
### 4. Keep the admin interface typed and server-rendered
|
||||
|
||||
`admin_api.py` will add explicit GET and POST routes for overview, search, dispatch/workers, Supervisor, logs, config, secrets, files, audit, and operation status. Forms retain exact field sets, bounded URL-encoded bodies, exact HTTPS Origin checks, CSRF, escaped output, CSP/HSTS/no-store headers, and POST/redirect/GET behavior. Unknown methods, route shapes, action names, source IDs, and file-root IDs fail closed.
|
||||
|
||||
Caddy will strip any inbound operator header and inject the authenticated Basic-auth username alongside the existing trusted edge marker. The backend accepts the actor only with that marker and records it in control/audit rows. Plaintext secrets are rendered only in the dedicated no-store page; no JavaScript, local storage, or audit payload receives their values.
|
||||
|
||||
The root-owned deployment profile is exactly `standalone-edge-v1` or `shared-host-edge-v1`. Standalone remains the default and owns host port 443. In shared-host mode, the existing host Caddy remains the sole owner of ports 80/443 and imports a fixed route-only snippet for only Worker API and the exact random admin prefix. It strips private/transit headers, injects an independent ingress marker, and proxies to the Truf edge at fixed loopback `127.0.0.1:18766`. The Truf edge rejects a missing marker before trusting the forwarded client address, binds only loopback, and retains Basic authentication, operator attribution, private backend marker, denylist, redacted logging, and security headers. It has no catch-all route for unrelated host applications.
|
||||
|
||||
Admin will call new exact Supervisor actions for structured snapshot, one managed-source lifecycle action, and bounded log tail. Web input will never be forwarded to Supervisor's generic command parser. Long-running apply/restart operations return an operation ID and status page because the process serving the POST may be restarted.
|
||||
|
||||
### 5. Share one strict configuration/secrets validator
|
||||
|
||||
A side-effect-free validator will be used by preview, runtime startup, and the host operations agent. It will enforce bounded UTF-8 YAML, duplicate-key rejection, mapping roots, strict scalar types and bounds, known keys, the exact core profile, credential-pool entry schemas and unique names, reference integrity, package capabilities, and managed deployment paths. Validation errors identify fields but never echo secret values.
|
||||
|
||||
Edits are candidate revisions, not direct active-file writes. Preview shows a structural/text diff with secret values redacted in audit and operation records. Candidate save and apply use expected SHA-256 hashes as compare-and-swap guards against stale forms or concurrent SSH changes.
|
||||
|
||||
Configuration and secrets remain separate logical resources and are not exposed through the generic file browser.
|
||||
|
||||
### 6. Use a narrow host operations agent for privileged lifecycle work
|
||||
|
||||
A root-owned systemd socket/service will accept local requests from the runtime UID over a Unix socket. Its request schema contains only an operation UUID, one action enum (`apply-config`, `apply-secrets`, `apply-both`, or `restart`), and expected active/candidate hashes. It accepts no command, service name, path, environment, Compose argument, or shell text.
|
||||
|
||||
Candidates live under a fixed host-managed bind directory shared read-only/read-write as required; active config/secrets will migrate from Docker-volume-only storage to fixed host-managed files before the agent is enabled. Immutable worker package manifests use a separate root-owned `/etc/truf/worker-packages` authority mapped read-only at `/data/worker-packages`; they never share the runtime-writable active-document trust root. Runtime-generated initialization state, lock, and PostgreSQL password remain in the private `/data` volume rather than the read-only active-document bind. The agent reads one root-owned exact profile and uses its fixed Compose files, network, ports, volume, capabilities, and Caddyfile. It never accepts that profile through its six-field request. The agent uses a singleton lock, revalidates candidate bytes, verifies all hashes, takes byte-identical backups, stops the fixed Truf runtime and edge, atomically replaces fixed files, recreates and attests that exact profile, and waits for Supervisor ACTIVE, PostgreSQL READY, Worker API, ingester, projector, and edge health. It never performs lifecycle actions on host Caddy, X-UI, or unrelated services. Failure restores backups and verifies the previous runtime. If both forward start and rollback fail, it enters a failed hold without deleting evidence or retrying indefinitely.
|
||||
|
||||
The admin request and operation row are committed before the agent begins. The agent writes a bounded result envelope that the runtime reconciles into PostgreSQL after restart. The Docker socket is never mounted into the runtime container.
|
||||
|
||||
### 7. Restrict managed files by logical root and descriptor-safe traversal
|
||||
|
||||
The file page exposes configured logical roots for logs, backups, and selected result/export directories. The client submits a root ID plus canonical relative path, never an absolute root. Config, secrets, PostgreSQL storage, application code, sockets, host-agent metadata, and raw result bundles are excluded.
|
||||
|
||||
Paths reject empty/absolute/drive-qualified/backslash/NUL/dot components and enforce byte, depth, listing, and file-size bounds. Linux traversal retains a root directory descriptor and uses component-wise `openat`/`dir_fd` operations with `O_NOFOLLOW`. Only single-link regular files are readable or replaceable; symlinks, hardlinks, reparse points, devices, FIFOs, and sockets are rejected. Writes use an exclusive same-directory temporary file, fsync, atomic replacement, directory fsync, and final owner/type/mode verification.
|
||||
|
||||
Typed operations are limited to list, view/download, create/replace, and delete within roots that explicitly allow each action. Every mutation records actor, logical root/path, before/after hashes, byte counts, operation result, and timestamp, never file content.
|
||||
|
||||
### 8. Persist operations and append-only audit records
|
||||
|
||||
Additive PostgreSQL tables will store operation lifecycle and audit events. Operation rows contain typed action/target, requested/started/completed timestamps, safe status/category/detail, expected and resulting revisions/hashes, and host-agent reconciliation state. Audit events are append-only and include actor, operation ID, action, logical target, before/after identities, result, and a previous-event/hash-chain identity.
|
||||
|
||||
Control mutation and its audit event commit atomically. File/config operation requests are audited when accepted and again when completed. Secret values, authorization headers, device tokens, provider tokens, CSRF values, and uploaded file bytes are forbidden from both schemas and logs.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [A discovery code path accidentally reaches local scanning] -> Use a separate bootstrap role and call graph, omit TruffleHog from the server image/profile, and test that scan/claim functions are never invoked.
|
||||
- [Pause races enqueue or claim] -> Enforce gates transactionally at queue admission and reservation, not only in UI or process state.
|
||||
- [A runtime restart interrupts uploads] -> Drain to database-committed bundles before planned apply; preserve upload/status routes during pause; rely on fixed-expiry replay for network failures.
|
||||
- [Protocol-2 rollout strands old work] -> Stop protocol-1 issuance first, retain its reconciliation/upload readers until no unresolved assignments remain, and only then remove compatibility in a later change.
|
||||
- [Direct Docker scanning is less efficient than layer reuse] -> Accept the bandwidth cost for the first bounded canary; add layer-plan transport only after the simpler authority path is proven.
|
||||
- [A source-specific access check grows into preventive server or worker security infrastructure] -> Keep provider access worker-authoritative and require separate operator approval plus an explicit requirement/task before adding probes, durable proofs, broad environment scrubbing, sandboxes, or redaction pipelines.
|
||||
- [Plaintext secret editing exposes values to an operator browser] -> Require the existing protected admin boundary, no-store responses, no client persistence/scripts, bounded rendering, and value-free audit/log records.
|
||||
- [The host agent becomes a root command proxy] -> Use a closed action enum and fixed paths/units, peer-credential checks, hash CAS, no shell, and adversarial request-schema tests.
|
||||
- [Filesystem containment has TOCTOU or link attacks] -> Use retained directory descriptors and no-follow operations for every component; reject multi-link and non-regular files.
|
||||
- [Rollback binary cannot read an additive schema] -> Keep migrations additive, preserve old markers, avoid incompatible constraint rewrites, and test old-image rollback before production cutover.
|
||||
- [Shared-host ingress exposes a private listener] -> Require host networking only in the exact shared profile, no Docker-published ports, explicit loopback bind, an independent ingress marker, and metadata attestation before lifecycle work.
|
||||
- [A Truf route captures or disrupts another host application] -> Install only a fixed route-only host-Caddy snippet with no listener, catch-all, global policy, or unrelated lifecycle authority.
|
||||
- [Shared profile drift changes the trust boundary] -> Read one stable root-owned mode-0444 profile, reject unknown values and metadata, and attest exact Compose labels, mounts, network, ports, capabilities, and Caddyfile.
|
||||
- [Server-rendered pages are less dynamic] -> Prefer explicit refresh/status pages over JavaScript to preserve the current CSP and reduce secret-retention surface.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add and test the control/audit/operation schema, discovery role, source adapters, protocol-2 package support, typed admin routes, validator, file service, and host agent while all new controls remain disabled.
|
||||
2. Build new server and Windows/Linux worker artifacts and verify code-authority/package manifests.
|
||||
3. Set dispatch paused, stop discovery, and allow or expire all protocol-1 assignments while continuing to accept their uploads.
|
||||
4. Stop runtime through the authenticated deployment path, create database and file backups, and run the additive migration under existing offline migration guards.
|
||||
5. Migrate active config/secrets to the fixed host-managed bind directory; install the exact root-owned ingress profile and socket-activated agent; start the new runtime and Truf edge with discovery and dispatch paused. In shared-host mode, install and validate the fixed route-only snippet in the existing host Caddy without granting the agent authority over that service.
|
||||
6. Verify health, admin actor attribution, control CAS/audit, Supervisor typed actions, managed-file containment, and rollback using a non-secret candidate.
|
||||
7. Publish protocol-2 worker packages. Enable a low-cap DockerHub public immutable-digest canary and verify search, enqueue, claim, scan, upload, ingestion, projection, expiry/replay, and drain.
|
||||
8. Enable GitLab and then public HuggingFace after the canary gates pass. Switch the default core profile exactly once and keep GitHub disabled.
|
||||
|
||||
Rollback restores byte-identical config/secrets and the previous runtime/edge images while retaining additive database tables and audit evidence. Shared-host rollback does not modify or restart host Caddy or unrelated services. Rollback must not begin while protocol-2 DockerHub/HuggingFace assignments or pre-commit bundles are unresolved. If rollback health also fails, the agent leaves the Truf deployment stopped/held with backups intact for SSH recovery.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- The retention duration and byte caps for browsable logs/backups/results need deployment defaults, but remain configurable within validator bounds.
|
||||
- Private DockerHub and HuggingFace worker credential delivery is deferred. It must not be designed or implemented without separate operator approval and a dedicated change defining only the agreed delivery and failure semantics.
|
||||
- Multi-operator authorization roles are deferred. This change records the Caddy Basic-auth username as actor but grants the existing admin policy uniformly.
|
||||
@@ -0,0 +1,32 @@
|
||||
## Why
|
||||
|
||||
The remote deployment can accept GitHub and GitLab worker assignments, but it cannot continuously discover targets without also entering the local scan path, and routine operation still requires SSH and direct file edits. The server needs a web-operated control plane that keeps discovery, dispatch, remote scanning, configuration, and runtime supervision separate and makes the intended distributed pipeline usable end to end.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add discovery-only server producers for the core source set `gitlab`, `dockerhub`, and `huggingface`; they may search and enqueue targets but must never claim or scan them locally.
|
||||
- Add persistent controls for pausing discovery, pausing new assignment dispatch, and draining the system while continuing to accept uploads for existing assignments.
|
||||
- Extend remote assignments, worker packages, scan execution, and result acceptance to support full GitLab, DockerHub, and HuggingFace claim-to-ingestion cycles, with DockerHub as the first deployed end-to-end canary.
|
||||
- Add authenticated admin pages for overview/search controls, workers and dispatch, Supervisor status/commands/logs, runtime configuration, plaintext `secrets.yaml`, managed files, and an operation audit trail.
|
||||
- Apply configuration and secret changes as validated, backed-up operations with coordinated runtime restart, health verification, and automatic rollback rather than in-place live mutation.
|
||||
- Support two exact root-selected production ingress profiles: the standalone Truf edge and a shared-host edge behind an existing root-owned host Caddy, without adding topology, port, path, service, or command fields to the host-agent request.
|
||||
- Expose only explicitly managed project directories through the file page; host root, PostgreSQL data, Docker control sockets, and arbitrary shell execution remain outside the web interface.
|
||||
- **BREAKING** Replace GitHub in the default core source set with HuggingFace; the new default core set is exactly GitLab, DockerHub, and HuggingFace.
|
||||
- **BREAKING** Advance worker compatibility so packages that support only the current GitHub/GitLab assignment contract are not eligible for the new core-source profile and must be rebuilt.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `distributed-core-source-processing`: Discovery-only production, persistent discovery/dispatch/drain controls, and remote-only processing for the configured core sources.
|
||||
- `multisource-worker-assignments`: Compatible worker packaging and fenced assignment/result lifecycles for GitLab, DockerHub, and HuggingFace.
|
||||
- `web-operations-console`: Authenticated web views and mutations for runtime overview, search, dispatch, workers, Supervisor, logs, and audited operations.
|
||||
- `managed-runtime-editing`: Validated editing and coordinated application of runtime configuration, plaintext secrets, and allowlisted project files with backup and rollback.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
The change affects source-cycle separation in `app/console_runner.py` and `app/supervisor.py`; queue and control state in `app/scanner_db.py`; worker assignment, package, client, scan execution, and API modules; the typed admin API and its HTML/CSS; standalone and shared-host Caddy/Compose profiles; profile-specific host lifecycle, installer, and denylist validation; runtime configuration and worker package manifests; and focused unit, integration, browser, and deployment tests. Existing queue and result authority remains PostgreSQL-backed, existing uploads remain accepted during drain, and no host filesystem or generic shell API is introduced.
|
||||
+102
@@ -0,0 +1,102 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Exact distributed core source set
|
||||
The server core profile SHALL contain exactly `gitlab`, `dockerhub`, and `huggingface`, and SHALL NOT start GitHub or any other discovery source as part of that profile.
|
||||
|
||||
#### Scenario: Core profile starts
|
||||
- **WHEN** Supervisor starts the distributed core profile
|
||||
- **THEN** it starts one discovery producer for GitLab, DockerHub, and HuggingFace and no GitHub producer
|
||||
|
||||
### Requirement: Discovery-only producer isolation
|
||||
Each core discovery producer SHALL search its provider, normalize targets, and admit them to PostgreSQL without claiming targets, acquiring scan slots, invoking a scanner, or staging result bundles locally.
|
||||
|
||||
#### Scenario: GitLab discovery finds targets
|
||||
- **WHEN** the GitLab producer completes a provider search
|
||||
- **THEN** it enqueues the normalized targets and records zero local scan requests
|
||||
|
||||
#### Scenario: DockerHub discovery resolves targets
|
||||
- **WHEN** the DockerHub producer processes pages, retries, tags, or digests
|
||||
- **THEN** it may persist discovery progress and immutable targets but never invokes TruffleHog or reserves those targets locally
|
||||
|
||||
#### Scenario: HuggingFace discovery finds Spaces
|
||||
- **WHEN** the HuggingFace producer returns Space identifiers
|
||||
- **THEN** it drops records explicitly marked private, protected, gated, or disabled and enqueues the remaining identifiers without entering a local scan path or making a second per-Space verification request
|
||||
|
||||
#### Scenario: Pending backlog exists
|
||||
- **WHEN** a scheduled discovery interval arrives while pending targets already exist
|
||||
- **THEN** the producer still performs the configured discovery cycle unless discovery is paused
|
||||
|
||||
### Requirement: Durable operations control state
|
||||
The system SHALL persist discovery pause, dispatch pause, drain state, revision, actor, operation identity, and update timestamps in PostgreSQL so that control state survives process and host restarts.
|
||||
|
||||
#### Scenario: Runtime restarts while paused
|
||||
- **WHEN** discovery and dispatch are paused and the runtime restarts
|
||||
- **THEN** both effective gates remain paused after startup
|
||||
|
||||
#### Scenario: Stale control form is submitted
|
||||
- **WHEN** a mutation supplies a revision older than the current control revision
|
||||
- **THEN** the system rejects it without changing control state or writing a success audit event
|
||||
|
||||
#### Scenario: Explicit pause coexists with drain
|
||||
- **WHEN** an operator explicitly pauses discovery, enters drain, and later cancels drain
|
||||
- **THEN** the explicit discovery pause remains set
|
||||
|
||||
### Requirement: Transactional discovery admission gate
|
||||
The system SHALL enforce the effective discovery gate in the same transaction that admits discovered targets or claims discovery-specific retry work.
|
||||
|
||||
#### Scenario: Pause races target admission
|
||||
- **WHEN** discovery pause commits before a producer admission transaction commits
|
||||
- **THEN** no newly discovered target is admitted by that transaction
|
||||
|
||||
#### Scenario: Discovery is paused before provider request
|
||||
- **WHEN** a producer begins a cycle while discovery is effectively paused
|
||||
- **THEN** it performs no provider request and records a paused cycle outcome
|
||||
|
||||
#### Scenario: Upload-derived work arrives during pause
|
||||
- **WHEN** result ingestion creates projection or keycheck work while discovery is paused
|
||||
- **THEN** that work remains admissible because it is not provider discovery
|
||||
|
||||
### Requirement: Transactional dispatch gate
|
||||
The system SHALL enforce the effective dispatch gate within the reservation transaction so that no new remote assignment can be issued after dispatch pause or drain commits.
|
||||
|
||||
#### Scenario: Dispatch pause races claim
|
||||
- **WHEN** dispatch pause commits before a worker claim transaction commits
|
||||
- **THEN** the claim returns a paused or no-work response and creates no reservation
|
||||
|
||||
#### Scenario: Existing worker uploads while paused
|
||||
- **WHEN** dispatch is paused and a worker with an existing assignment reports status or uploads its result
|
||||
- **THEN** the server accepts the valid request under the existing assignment authority
|
||||
|
||||
#### Scenario: Assignment expires while paused
|
||||
- **WHEN** an existing assignment expires during dispatch pause
|
||||
- **THEN** the reaper processes it normally without issuing replacement work
|
||||
|
||||
### Requirement: Drain lifecycle
|
||||
Entering drain SHALL effectively pause discovery and dispatch while preserving status, terminal report, upload, receipt replay, ingestion, projection, and maintenance paths needed to finish accepted work.
|
||||
|
||||
#### Scenario: Drain begins with active assignments
|
||||
- **WHEN** drain is requested while remote assignments are active
|
||||
- **THEN** the state becomes `draining`, no new targets or assignments are admitted, and existing workers retain their result path
|
||||
|
||||
#### Scenario: Drain reaches completion
|
||||
- **WHEN** no live remote assignments remain and every accepted result bundle has reached database commit
|
||||
- **THEN** the reconciler advances the state to `drained`
|
||||
|
||||
#### Scenario: Pending targets remain
|
||||
- **WHEN** pending queue targets remain but all issued assignments and pre-commit bundles are resolved
|
||||
- **THEN** drain may still become `drained`
|
||||
|
||||
#### Scenario: Projection work remains
|
||||
- **WHEN** projection or keycheck work remains after its result bundle is database-committed
|
||||
- **THEN** that work does not prevent the control state from becoming `drained`
|
||||
|
||||
### Requirement: Discovery process observability
|
||||
Supervisor and the operations console SHALL expose each discovery producer's source, role, lifecycle state, last cycle result, last successful discovery time, next scheduled run, and bounded safe error category.
|
||||
|
||||
#### Scenario: Provider rejects credentials
|
||||
- **WHEN** a discovery producer receives a provider authorization error
|
||||
- **THEN** operations state reports the source and safe authorization category without exposing the credential or provider response body containing secrets
|
||||
|
||||
#### Scenario: Producer is stopped
|
||||
- **WHEN** an operator stops a managed producer through a typed Supervisor action
|
||||
- **THEN** structured state identifies it as stopped without changing the persisted discovery pause flag
|
||||
+182
@@ -0,0 +1,182 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Shared strict document validation
|
||||
Configuration and secrets preview, startup, and privileged apply SHALL use the same side-effect-free validator for bounded UTF-8 YAML, duplicate keys, mapping roots, strict scalar types and bounds, known keys, exact core profile, auth-pool schemas, unique entry names, reference integrity, package capabilities, and deployment paths.
|
||||
|
||||
#### Scenario: Valid configuration is previewed
|
||||
- **WHEN** an operator submits a candidate satisfying the complete schema
|
||||
- **THEN** preview returns normalized validation success and a bounded diff without activating the candidate
|
||||
|
||||
#### Scenario: Duplicate or unknown key is submitted
|
||||
- **WHEN** a candidate contains a duplicate mapping key or unsupported field
|
||||
- **THEN** validation fails before staging or restart and identifies the field without echoing secret values
|
||||
|
||||
#### Scenario: Secret reference is invalid
|
||||
- **WHEN** configuration selects an auth-pool entry that does not exist in the candidate secrets
|
||||
- **THEN** combined validation fails and neither document becomes active
|
||||
|
||||
### Requirement: Separate immutable package authority
|
||||
Worker package manifests SHALL resolve only beneath a separate root-owned, non-writable package authority, while runtime-generated initialization state, locks, and PostgreSQL credentials SHALL remain in the private writable data volume rather than the active config/secrets bind.
|
||||
|
||||
#### Scenario: Package manifest uses the active-document directory
|
||||
- **WHEN** configuration references a package manifest beneath the config/secrets authority or outside the fixed package root
|
||||
- **THEN** validation fails before startup or privileged apply
|
||||
|
||||
#### Scenario: Runtime initializes with active documents mounted read-only
|
||||
- **WHEN** the runtime initializes or restarts with the active config/secrets directory mounted read-only
|
||||
- **THEN** its initialization marker, singleton lock, and generated PostgreSQL password remain writable only through fixed private data paths
|
||||
|
||||
### Requirement: Candidate revisions and compare-and-swap
|
||||
The system SHALL stage validated candidate revisions separately from active files and SHALL require expected active and candidate SHA-256 identities when saving or applying them.
|
||||
|
||||
#### Scenario: Candidate is saved
|
||||
- **WHEN** an operator saves valid bytes against the current active hash
|
||||
- **THEN** the candidate is durably staged with a new hash and the active file is unchanged
|
||||
|
||||
#### Scenario: Active file changed through SSH
|
||||
- **WHEN** the active hash differs from the expected hash submitted by a stale page
|
||||
- **THEN** save or apply fails without replacing either active file
|
||||
|
||||
#### Scenario: Candidate changed concurrently
|
||||
- **WHEN** the supplied candidate hash is no longer current
|
||||
- **THEN** apply is rejected before host lifecycle changes begin
|
||||
|
||||
### Requirement: Plaintext secrets editing without persistence leakage
|
||||
The protected secrets page SHALL allow authorized operators to view and edit the complete plaintext YAML document while responses remain no-store and secret values remain absent from browser persistence, application logs, diffs outside that page, operation status, and audit records.
|
||||
|
||||
#### Scenario: Operator opens secrets page
|
||||
- **WHEN** an authenticated operator requests the dedicated secrets editor
|
||||
- **THEN** the current document is rendered in a server-side form over the protected no-store response
|
||||
|
||||
#### Scenario: Secrets candidate is validated
|
||||
- **WHEN** an operator previews or saves changed secrets
|
||||
- **THEN** validation results name safe field paths and hashes but do not repeat credential values
|
||||
|
||||
#### Scenario: Operator leaves the page
|
||||
- **WHEN** the browser navigates to another admin page
|
||||
- **THEN** the application has written no secret value to local storage, session storage, service workers, or client-side application state
|
||||
|
||||
### Requirement: Closed host-agent protocol
|
||||
The privileged host operations agent SHALL accept only a canonical operation UUID, one fixed action enum, expected active hashes, and expected candidate hashes from the authorized runtime peer over a local Unix socket. Deployment profile is root-installed host policy and SHALL NOT be added to that request.
|
||||
|
||||
#### Scenario: Valid apply request arrives
|
||||
- **WHEN** the authorized runtime UID submits an exact valid request
|
||||
- **THEN** the agent verifies the persisted operation and hashes before acquiring the singleton apply lock
|
||||
|
||||
#### Scenario: Request includes path or command data
|
||||
- **WHEN** a request includes a service name, path, shell text, environment, Docker argument, or unknown field
|
||||
- **THEN** the agent rejects it before any privileged action
|
||||
|
||||
#### Scenario: Request attempts topology selection
|
||||
- **WHEN** a request includes a profile, ingress port, upstream, Caddy path, unit, service, command, environment, Compose file, or Compose argument
|
||||
- **THEN** the agent rejects it before reading candidate documents or stopping the deployment
|
||||
|
||||
#### Scenario: Unauthorized local peer connects
|
||||
- **WHEN** a process with an unapproved peer identity uses the socket
|
||||
- **THEN** the agent rejects the request regardless of its JSON body
|
||||
|
||||
### Requirement: Coordinated apply with health verification
|
||||
For config, secrets, or combined apply, the host agent SHALL revalidate fixed candidate files, verify compare-and-swap hashes, create byte-identical backups, stop the fixed Truf deployment, atomically replace active files, recreate the exact root-selected runtime and edge profile, attest its fixed topology, and wait for required health before reporting success. In shared-host mode it SHALL NOT stop, restart, reload, reconfigure, or remove host Caddy, X-UI, or another unrelated service.
|
||||
|
||||
#### Scenario: Combined apply succeeds
|
||||
- **WHEN** both candidates validate and the restarted deployment reaches Supervisor ACTIVE, PostgreSQL READY, Worker API, ingester, and projector health
|
||||
- **THEN** the operation completes successfully with resulting hashes and retained rollback evidence
|
||||
|
||||
#### Scenario: Validation changes between preview and apply
|
||||
- **WHEN** host-side revalidation or hash verification differs from the accepted candidate operation
|
||||
- **THEN** the agent aborts before stopping the healthy runtime
|
||||
|
||||
#### Scenario: Apply is requested while another is active
|
||||
- **WHEN** the singleton operation lock is held
|
||||
- **THEN** the second request is rejected or remains queued without overlapping lifecycle mutations
|
||||
|
||||
#### Scenario: Shared-host profile is healthy
|
||||
- **WHEN** shared-host apply recreates a host-network runtime with no published ports and an edge bound only to fixed loopback, and all runtime and edge health checks pass
|
||||
- **THEN** the operation succeeds without a host-Caddy or X-UI lifecycle action
|
||||
|
||||
#### Scenario: Shared-host topology drifts
|
||||
- **WHEN** runtime publishes a port, edge binds a non-loopback address, profile metadata changes, or Compose labels, mounts, network, capabilities, or Caddyfile differ from the fixed profile
|
||||
- **THEN** the agent rejects the deployment before candidate replacement or reports failed health without claiming success
|
||||
|
||||
#### Scenario: Shared-host rollback succeeds
|
||||
- **WHEN** candidate health fails and the previous Truf runtime and edge are restored
|
||||
- **THEN** rollback completes exactly once without changing host Caddy, X-UI, or another unrelated service
|
||||
|
||||
### Requirement: Automatic rollback and failed hold
|
||||
If the new deployment fails its bounded health check, the agent SHALL restore byte-identical backups and verify the previous deployment; if rollback also fails, it SHALL stop retrying and retain a failed-hold state and all evidence for SSH recovery.
|
||||
|
||||
#### Scenario: New configuration fails startup
|
||||
- **WHEN** the restarted runtime cannot reach required health within the deadline
|
||||
- **THEN** the agent restores the prior active files and restarts the previous deployment
|
||||
|
||||
#### Scenario: Rollback succeeds
|
||||
- **WHEN** the restored deployment reaches required health
|
||||
- **THEN** the operation records rolled-back status and safe failure category without claiming apply success
|
||||
|
||||
#### Scenario: Rollback fails
|
||||
- **WHEN** neither the candidate nor restored deployment becomes healthy
|
||||
- **THEN** the agent enters failed hold, performs no replacement loop or forced authority release, and preserves backups and diagnostics
|
||||
|
||||
### Requirement: Logical managed roots
|
||||
The generic file page SHALL address only configured logical roots with explicit read, create/replace, and delete permissions, and SHALL never accept an absolute root from a client.
|
||||
|
||||
#### Scenario: Operator lists a managed runtime root
|
||||
- **WHEN** a valid logical root ID for logs, keycheck projections, or result projections and a canonical relative directory are requested
|
||||
- **THEN** the service returns a bounded read-only listing of permitted regular files and directories under that root
|
||||
|
||||
#### Scenario: Operator downloads rotated result projections
|
||||
- **WHEN** the operator requests a permitted active or rotated result projection within its configured file-size bound
|
||||
- **THEN** the service returns that regular single-link file without granting mutation access or exposing the backing runtime path
|
||||
|
||||
#### Scenario: Result projection names are allowlisted
|
||||
- **WHEN** the result-projection root is listed or read
|
||||
- **THEN** only active `scan_results.jsonl` and `found_secrets.jsonl` files and their exact six-digit generation names are visible, while locks, databases, ledgers, scan errors, temporary/quarantine directories, malformed generations, and recovery artifacts remain unavailable
|
||||
|
||||
#### Scenario: Large result projection is downloaded
|
||||
- **WHEN** an allowed result projection is within the larger result-root byte bound
|
||||
- **THEN** the service copies and hashes an unchanged source revision into an anonymous same-volume snapshot using bounded chunks, permits only one such snapshot at a time, streams the snapshot with bounded memory, and releases the snapshot and concurrency slot after response completion or failure
|
||||
|
||||
#### Scenario: Operator requests excluded storage
|
||||
- **WHEN** a request targets application code, config/secrets through the generic page, PostgreSQL storage, sockets, host-agent metadata, raw result bundles, or an unknown root
|
||||
- **THEN** the service rejects it without revealing host paths or existence details
|
||||
|
||||
### Requirement: Descriptor-safe path containment
|
||||
Managed-file traversal and mutation SHALL use a retained root directory descriptor, component-wise no-follow operations, canonical relative components, and regular single-link file checks.
|
||||
|
||||
#### Scenario: Relative traversal is attempted
|
||||
- **WHEN** a path contains an empty, dot, dot-dot, absolute, drive-qualified, backslash, NUL, over-depth, or over-length component
|
||||
- **THEN** the request is rejected before filesystem access outside the retained root
|
||||
|
||||
#### Scenario: Symlink is swapped during access
|
||||
- **WHEN** a path component becomes a symlink between validation and open
|
||||
- **THEN** no-follow descriptor traversal fails without accessing the link target
|
||||
|
||||
#### Scenario: Hardlink or special file is targeted
|
||||
- **WHEN** the final object is multi-linked or is not a regular file
|
||||
- **THEN** view, download, replace, and delete are rejected
|
||||
|
||||
### Requirement: Durable bounded file mutation
|
||||
An allowed managed-file create or replace SHALL use an exclusive same-directory temporary regular file, bounded bytes, fsync, atomic descriptor-relative replacement, directory fsync, and final ownership/type/mode verification.
|
||||
|
||||
#### Scenario: File replacement succeeds
|
||||
- **WHEN** an authorized bounded replacement is submitted against the current file hash
|
||||
- **THEN** readers observe either the complete old file or complete new file and audit records the safe before/after hashes
|
||||
|
||||
#### Scenario: File is concurrently changed
|
||||
- **WHEN** the current file hash differs from the expected hash
|
||||
- **THEN** replacement fails without overwriting the concurrent change
|
||||
|
||||
#### Scenario: Upload exceeds the root limit
|
||||
- **WHEN** submitted bytes exceed the configured bounded file size
|
||||
- **THEN** the service rejects and removes temporary data without changing the target
|
||||
|
||||
### Requirement: Content-free operational audit
|
||||
Configuration, secrets, and managed-file operations SHALL record actor, typed action, logical target, timestamps, result, safe category, byte counts where applicable, and before/after hashes, but SHALL NOT record file contents or credentials.
|
||||
|
||||
#### Scenario: Managed file is deleted
|
||||
- **WHEN** an allowed delete succeeds against the expected hash
|
||||
- **THEN** audit records the logical root/path and previous hash without retaining deleted content
|
||||
|
||||
#### Scenario: Secrets apply fails
|
||||
- **WHEN** a secrets operation fails validation, startup, or rollback
|
||||
- **THEN** status and audit expose only the bounded failure category and document hashes
|
||||
+120
@@ -0,0 +1,120 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Protocol 2 source capabilities
|
||||
Protocol-2 worker packages SHALL advertise explicit source, worker-platform, and planning-kind capabilities for GitLab, DockerHub, and HuggingFace, and the server SHALL issue work only when the selected package supports the complete assignment capability.
|
||||
|
||||
#### Scenario: Compatible package requests work
|
||||
- **WHEN** a protocol-2 package advertising the required capability requests a supported target
|
||||
- **THEN** the server may create an assignment using that capability
|
||||
|
||||
#### Scenario: Package lacks planning capability
|
||||
- **WHEN** a package advertises the source but not the required planning kind
|
||||
- **THEN** the server rejects the claim without reserving a target
|
||||
|
||||
#### Scenario: Package advertises unknown capability
|
||||
- **WHEN** a package manifest contains an unknown source, platform, or planning kind
|
||||
- **THEN** package validation fails closed
|
||||
|
||||
### Requirement: Canonical multisource execution plans
|
||||
The server SHALL create and validate immutable execution snapshots using `exact_git_v1` for GitLab, `docker_direct_v1` for DockerHub, and `huggingface_space_v1` for HuggingFace.
|
||||
|
||||
#### Scenario: GitLab assignment is issued
|
||||
- **WHEN** a GitLab target is claimed
|
||||
- **THEN** the assignment binds the existing exact commit and Git scan plan under `exact_git_v1`
|
||||
|
||||
#### Scenario: DockerHub assignment is issued
|
||||
- **WHEN** a public DockerHub target is claimed
|
||||
- **THEN** the assignment binds an immutable digest reference under `docker_direct_v1` and does not depend on a mutable tag
|
||||
|
||||
#### Scenario: HuggingFace assignment is issued
|
||||
- **WHEN** a public HuggingFace Space is claimed
|
||||
- **THEN** the assignment binds its canonical Space identifier under `huggingface_space_v1`
|
||||
|
||||
#### Scenario: Snapshot shape does not match source
|
||||
- **WHEN** an execution snapshot's source, worker platform, or planning kind combination is invalid
|
||||
- **THEN** the server and worker reject it before scanner execution
|
||||
|
||||
### Requirement: Fenced assignment authority for every source
|
||||
Every supported source assignment SHALL bind one user, device, target, fixed expiry, immutable execution snapshot, result reservation, and result bundle identity using the existing PostgreSQL authority model.
|
||||
|
||||
#### Scenario: Lost claim response is retried
|
||||
- **WHEN** the server committed an assignment but the worker did not receive the response
|
||||
- **THEN** retrying the same admission request returns the same assignment and immutable execution snapshot
|
||||
|
||||
#### Scenario: Stale worker uploads
|
||||
- **WHEN** a worker uploads with an expired, replaced, or mismatched reservation token
|
||||
- **THEN** the server rejects the upload without changing queue or bundle authority
|
||||
|
||||
#### Scenario: Valid result commits
|
||||
- **WHEN** a valid assigned worker uploads and finalizes its bundle
|
||||
- **THEN** ingestion commits the queue result and downstream projection work exactly once
|
||||
|
||||
### Requirement: Eligible source fallback
|
||||
The assignment service SHALL try other eligible configured sources when one supported source has no claimable target, while still issuing at most one assignment for an admission request.
|
||||
|
||||
#### Scenario: Initially selected source is empty
|
||||
- **WHEN** the first eligible source has no claimable target and another eligible source does
|
||||
- **THEN** the same claim request may receive one assignment from the other source
|
||||
|
||||
#### Scenario: All eligible sources are empty
|
||||
- **WHEN** no compatible source has a claimable target
|
||||
- **THEN** the claim returns no work and creates no reservation
|
||||
|
||||
### Requirement: Legacy protocol-1 completion compatibility
|
||||
After protocol-2 cutover, the server SHALL stop issuing new claims to protocol-1 packages but SHALL continue status, terminal report, upload, receipt replay, and immutable snapshot reconciliation for already-issued protocol-1 assignments until they resolve or expire.
|
||||
|
||||
#### Scenario: Protocol-1 package requests a new claim
|
||||
- **WHEN** a legacy GitHub/GitLab-only package requests new work after cutover
|
||||
- **THEN** the server returns an incompatibility response and creates no assignment
|
||||
|
||||
#### Scenario: Existing protocol-1 assignment uploads
|
||||
- **WHEN** a legacy worker uploads a valid result for an assignment issued before cutover
|
||||
- **THEN** the server accepts and ingests it under its original immutable authority
|
||||
|
||||
#### Scenario: Legacy snapshot is reconciled
|
||||
- **WHEN** the server reconstructs a lost response for an existing protocol-1 assignment
|
||||
- **THEN** it reads the original snapshot without rewriting it into protocol 2
|
||||
|
||||
### Requirement: DockerHub end-to-end canary
|
||||
The rollout SHALL prove a bounded DockerHub `search -> enqueue -> claim -> scan -> upload -> ingestion -> projection` cycle using a public image resolved to an immutable digest before broader new-source enablement.
|
||||
|
||||
#### Scenario: DockerHub canary succeeds
|
||||
- **WHEN** a canary producer discovers the configured public image and a compatible worker processes it
|
||||
- **THEN** the target reaches database-committed ingestion and projection under one fenced assignment
|
||||
|
||||
#### Scenario: Mutable tag changes during canary
|
||||
- **WHEN** the discovered tag changes after queue admission
|
||||
- **THEN** the worker still scans the immutable digest bound in its assignment
|
||||
|
||||
#### Scenario: Registry credentials would be required
|
||||
- **WHEN** the DockerHub canary target cannot be scanned without private registry credentials
|
||||
- **THEN** the worker returns a bounded inaccessible-provider result, the server applies its declared retryability, and no discovery or registry credential is transferred in the assignment
|
||||
|
||||
### Requirement: HuggingFace remote processing
|
||||
The system SHALL support the same fenced claim-to-ingestion lifecycle for public HuggingFace Spaces without invoking the scanner on the server.
|
||||
|
||||
#### Scenario: Public Space completes
|
||||
- **WHEN** a compatible worker claims and scans a public HuggingFace Space
|
||||
- **THEN** its result is uploaded, ingested, and projected under the bound assignment
|
||||
|
||||
#### Scenario: Space is inaccessible without worker credentials
|
||||
- **WHEN** a tokenless worker cannot read a claimed HuggingFace Space because its repository is private, protected, removed, or otherwise unavailable
|
||||
- **THEN** it returns a non-retryable inaccessible result, the server does not retry that target, and the server discovery token is never exposed
|
||||
|
||||
### Requirement: Worker-authoritative provider access
|
||||
The server SHALL validate canonical target and assignment authority but SHALL treat worker execution as the final provider-access check. A source SHALL NOT require a per-target server access probe, durable public-access proof, proof-freshness state, broad child-environment credential scrubbing, credential sandbox, or post-hoc redaction pipeline unless the operator separately approves an explicit OpenSpec requirement and implementation task.
|
||||
|
||||
#### Scenario: Provider accessibility changes after discovery
|
||||
- **WHEN** a canonical target becomes inaccessible before worker execution
|
||||
- **THEN** the worker returns the source's bounded permanent or retryable provider-failure result and the server settles or retries it according to that result
|
||||
|
||||
#### Scenario: Another source adapter is proposed
|
||||
- **WHEN** implementation would add preventive access proof or source-specific security infrastructure beyond the assignment's declared fields
|
||||
- **THEN** implementation pauses until the operator approves a dedicated requirement and task
|
||||
|
||||
### Requirement: Credential and result secrecy
|
||||
Provider credentials, worker device tokens, authorization headers, and result contents SHALL NOT appear in operation status, audit records, Supervisor snapshots, or routine assignment logs.
|
||||
|
||||
#### Scenario: Assignment logging occurs
|
||||
- **WHEN** any supported source assignment is created, retried, rejected, or completed
|
||||
- **THEN** logs identify bounded source and authority metadata without credential values or result payload bytes
|
||||
+142
@@ -0,0 +1,142 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Protected typed admin routes
|
||||
The operations console SHALL expose explicit server-rendered routes and exact mutation forms behind the existing random admin path, Caddy Basic authentication, trusted edge marker, exact same-origin check, CSRF validation, no-store responses, and restrictive security headers.
|
||||
|
||||
#### Scenario: Authorized operator opens a page
|
||||
- **WHEN** Caddy authenticates the request and injects the trusted marker and operator identity
|
||||
- **THEN** the requested operations page renders escaped server-side HTML with no client-side secret persistence
|
||||
|
||||
#### Scenario: Direct backend request lacks marker
|
||||
- **WHEN** a request reaches an admin route without the trusted edge marker
|
||||
- **THEN** the backend rejects it regardless of supplied operator headers
|
||||
|
||||
#### Scenario: Mutation has stale or invalid CSRF
|
||||
- **WHEN** a POST has a missing, duplicate, or invalid CSRF value or wrong Origin
|
||||
- **THEN** the backend rejects the mutation without side effects
|
||||
|
||||
#### Scenario: Unknown route or form action is submitted
|
||||
- **WHEN** a request contains an unsupported method, route shape, action, field, or duplicate field
|
||||
- **THEN** it fails closed without invoking Supervisor, database mutations, or host operations
|
||||
|
||||
### Requirement: Trusted operator attribution
|
||||
Caddy SHALL strip any inbound operator identity header and inject the authenticated Basic-auth username, and the backend SHALL trust that identity only with the private edge marker.
|
||||
|
||||
#### Scenario: Client spoofs operator header
|
||||
- **WHEN** a public request supplies its own operator identity header
|
||||
- **THEN** Caddy removes it and the audit actor is the authenticated Basic-auth user
|
||||
|
||||
#### Scenario: Mutation is accepted
|
||||
- **WHEN** an authenticated operator performs a valid mutation
|
||||
- **THEN** the control or operation record and its audit event identify that operator
|
||||
|
||||
### Requirement: Exact production ingress profiles
|
||||
The production deployment SHALL use exactly one root-installed profile: `standalone-edge-v1` or `shared-host-edge-v1`. The selected profile SHALL NOT be supplied by an admin request, host-agent request, runtime document, or other unprivileged input.
|
||||
|
||||
#### Scenario: Standalone edge is selected
|
||||
- **WHEN** `standalone-edge-v1` is installed
|
||||
- **THEN** the managed Truf edge remains the sole Truf listener on host port 443 and retains the exact runtime-network-namespace contract
|
||||
|
||||
#### Scenario: Shared-host edge is selected
|
||||
- **WHEN** `shared-host-edge-v1` is installed
|
||||
- **THEN** the existing root-owned host Caddy remains the sole owner of ports 80/443 and proxies only fixed Truf routes to a managed edge bound at `127.0.0.1:18766`
|
||||
|
||||
#### Scenario: A request attempts to select topology
|
||||
- **WHEN** a request supplies a deployment mode, upstream, port, Caddy path, unit, service, command, or Compose argument
|
||||
- **THEN** it is rejected before lifecycle work
|
||||
|
||||
### Requirement: Shared-host route confinement
|
||||
The shared-host profile SHALL install a fixed root-owned route-only host-Caddy snippet. It SHALL claim only `/api/v1/worker/*`, the exact random admin-prefix root, and that prefix's subtree. It SHALL strip inbound private and transit headers, inject an independent ingress marker and canonical client address, and preserve the managed edge's authentication, operator attribution, denylist, redacted logging, and security-header behavior without adding a listener, TLS policy, global error handler, trusted-proxy policy, catch-all, or unrelated route.
|
||||
|
||||
#### Scenario: An unrelated host route is requested
|
||||
- **WHEN** a request does not match a Truf worker or admin path
|
||||
- **THEN** the Truf snippet does not handle or alter the request
|
||||
|
||||
#### Scenario: The loopback Truf edge is unavailable
|
||||
- **WHEN** a matching route cannot reach `127.0.0.1:18766`
|
||||
- **THEN** host Caddy fails that Truf request without forwarding it to X-UI or another fallback upstream
|
||||
|
||||
#### Scenario: A client supplies transit headers
|
||||
- **WHEN** a public request supplies an ingress marker, forwarded address, private edge marker, or operator identity
|
||||
- **THEN** host Caddy strips those values and injects only its reviewed ingress marker and observed client address
|
||||
|
||||
### Requirement: Runtime overview
|
||||
The overview page SHALL report bounded structured health for Supervisor, PostgreSQL, required pipeline workers, discovery producers, queue status, active remote assignments, result bundles, operation controls, and recent operation outcomes.
|
||||
|
||||
#### Scenario: Runtime is healthy
|
||||
- **WHEN** all required components hold valid authority and health
|
||||
- **THEN** the overview reports the runtime active and identifies each required component without exposing secrets
|
||||
|
||||
#### Scenario: Component is unavailable
|
||||
- **WHEN** a health source times out or returns malformed state
|
||||
- **THEN** the overview reports that component unavailable without blocking the rest of the page
|
||||
|
||||
### Requirement: Search controls
|
||||
The search page SHALL expose each core producer's structured state and typed start, stop, restart, pause, resume, and interval controls while clearly separating process lifecycle from the persistent discovery gate.
|
||||
|
||||
#### Scenario: Operator pauses search
|
||||
- **WHEN** an operator submits pause with the current control revision
|
||||
- **THEN** the persistent discovery gate changes atomically and every producer stops admitting new discovered targets
|
||||
|
||||
#### Scenario: Operator restarts one producer
|
||||
- **WHEN** an operator selects restart for an allowed producer ID
|
||||
- **THEN** only that managed discovery process restarts and the persistent pause state is unchanged
|
||||
|
||||
### Requirement: Dispatch and drain controls
|
||||
The workers/dispatch page SHALL expose persistent dispatch pause/resume, drain start/cancel, drain progress, compatible package state, worker users/devices, assignment counts, and upload availability.
|
||||
|
||||
#### Scenario: Operator pauses dispatch
|
||||
- **WHEN** the current revision is submitted to the pause action
|
||||
- **THEN** no new worker assignment can commit while valid existing uploads remain accepted
|
||||
|
||||
#### Scenario: Operator starts drain
|
||||
- **WHEN** drain is started
|
||||
- **THEN** the page reports draining progress from authoritative assignment and bundle counts until the state becomes drained
|
||||
|
||||
#### Scenario: Stale page attempts resume
|
||||
- **WHEN** another operator has changed the control revision before resume is submitted
|
||||
- **THEN** the console reports a revision conflict and does not overwrite the newer state
|
||||
|
||||
### Requirement: Typed Supervisor operations
|
||||
The console SHALL use a closed Supervisor protocol for structured snapshot, allowlisted managed-source lifecycle actions, and bounded log tail, and SHALL NOT forward generic command strings.
|
||||
|
||||
#### Scenario: Operator requests source status
|
||||
- **WHEN** the Supervisor page loads
|
||||
- **THEN** it displays structured source IDs, roles, phases, process state, restart state, and safe errors without parsing a text dashboard
|
||||
|
||||
#### Scenario: Operator tails logs
|
||||
- **WHEN** an allowed managed source and bounded line count are submitted
|
||||
- **THEN** Supervisor returns only that source's bounded log tail
|
||||
|
||||
#### Scenario: Input resembles a shell command
|
||||
- **WHEN** an operator submits command text, a path, or an unrecognized source ID
|
||||
- **THEN** the request is rejected and no generic Supervisor command or operating-system shell is called
|
||||
|
||||
### Requirement: Durable asynchronous operation status
|
||||
Long-running restart and apply actions SHALL create a PostgreSQL operation record before execution and SHALL remain queryable by operation ID across runtime/admin restarts.
|
||||
|
||||
#### Scenario: Apply restarts the admin process
|
||||
- **WHEN** the process that accepted an apply request terminates during the coordinated restart
|
||||
- **THEN** the operator can reopen the operation URL and observe reconciled success, rollback, or failure state
|
||||
|
||||
#### Scenario: Unknown operation is requested
|
||||
- **WHEN** an operator requests an operation ID that does not exist or is not canonical
|
||||
- **THEN** the console returns not found without searching filesystem paths or host-agent state by user input
|
||||
|
||||
### Requirement: Append-only audit view
|
||||
The audit page SHALL show bounded append-only events for accepted and completed controls, Supervisor actions, configuration/secrets operations, and managed-file mutations, including actor, action, logical target, time, result, and safe before/after identity.
|
||||
|
||||
#### Scenario: Secret apply is audited
|
||||
- **WHEN** a secrets candidate is accepted and later applied or rolled back
|
||||
- **THEN** audit events record hashes and outcomes but no secret value, candidate bytes, authorization data, or CSRF value
|
||||
|
||||
#### Scenario: Audit pagination is requested
|
||||
- **WHEN** an operator navigates audit history
|
||||
- **THEN** the backend returns a bounded deterministic page without unbounded database or browser output
|
||||
|
||||
### Requirement: Existing worker API availability
|
||||
Adding the operations console SHALL NOT weaken or couple public worker endpoints to admin page availability.
|
||||
|
||||
#### Scenario: Admin feature is disabled or unhealthy
|
||||
- **WHEN** the admin console is disabled or a Supervisor/host-agent status dependency is unavailable
|
||||
- **THEN** authenticated worker status, upload, terminal report, and receipt paths continue under their existing authority
|
||||
@@ -0,0 +1,92 @@
|
||||
## 1. PostgreSQL Operations Authority
|
||||
|
||||
- [x] 1.1 Add an additive runtime-safety migration for the singleton operations control row, durable operation records, and append-only audit events
|
||||
- [x] 1.2 Add schema invariants, final-cutover checks, import/export handling, and migration-count fixtures for the new tables
|
||||
- [x] 1.3 Implement ScannerDB control-state reads and revision-checked discovery, dispatch, and drain mutations with atomic audit insertion
|
||||
- [x] 1.4 Enforce the discovery gate transactionally in provider admission, discovery retry, and Docker tag-resolution paths without blocking ingestion-derived work
|
||||
- [x] 1.5 Enforce the dispatch gate transactionally in remote reservation admission while preserving status, terminal report, upload, replay, expiry, and ingestion
|
||||
- [x] 1.6 Implement drain progress queries and reconciliation from live assignments and pre-commit result bundles
|
||||
- [x] 1.7 Add concurrent PostgreSQL tests for pause/admission races, stale revisions, restart persistence, drain completion, and uninterrupted uploads
|
||||
|
||||
## 2. Discovery-Only Server Producers
|
||||
|
||||
- [x] 2.1 Extract a discovery-only cycle for GitLab, DockerHub, and HuggingFace that performs provider work and enqueueing without scan preparation or claiming
|
||||
- [x] 2.2 Preserve DockerHub page cursors, retry-lane processing, tag resolution, and immutable target admission in the discovery role
|
||||
- [x] 2.3 Add an authenticated `discovery-producer` runtime bootstrap role and Supervisor managed-process type with structured state
|
||||
- [x] 2.4 Change the default distributed core profile to exactly GitLab, DockerHub, and HuggingFace and remove GitHub from that profile
|
||||
- [x] 2.5 Add tests proving each producer runs with backlog, respects persistent pause, reports safe state, and never enters local scanner/claim/bundle code
|
||||
|
||||
## 3. Protocol-2 Multisource Assignments
|
||||
|
||||
- [x] 3.1 Replace the Git-only assignment switch with source adapters that declare queue source, worker platform, planning kind, snapshot validation, and package capability
|
||||
- [x] 3.2 Implement GitLab `exact_git_v1`, DockerHub `docker_direct_v1`, and HuggingFace `huggingface_space_v1` execution-snapshot models and canonical validation
|
||||
- [x] 3.3 Generalize ScannerDB claim recovery and result-ready validation for the three planning kinds while retaining fixed reservation and device fences
|
||||
- [x] 3.4 Update source selection to try other compatible eligible queues while issuing at most one assignment per admission request
|
||||
- [x] 3.5 Advance worker/package manifests to protocol 2 with explicit source/platform/planning capabilities and fail-closed manifest validation
|
||||
- [x] 3.6 Update Windows and Linux worker clients to dispatch the bound Docker and HuggingFace scan platforms and validate protocol-2 snapshots before execution
|
||||
- [x] 3.7 Retain protocol-1 status, upload, terminal-report, receipt, and immutable reconciliation for existing assignments while refusing new protocol-1 claims
|
||||
- [x] 3.8 Add unit and PostgreSQL integration tests for capability matching, fallback, replay, expiry, stale upload rejection, source aliases, and legacy completion
|
||||
|
||||
## 4. New-Source End-to-End Canaries
|
||||
|
||||
Implementation guardrail: provider accessibility is finalized by the worker. New per-target server preflight/proof state, broad worker-environment credential scrubbing, or other source-specific defensive infrastructure requires separate operator approval and an explicit OpenSpec requirement/task before implementation.
|
||||
|
||||
- [x] 4.1 Implement public DockerHub immutable-digest assignment execution without server registry-credential or layer-plan transport
|
||||
- [x] 4.2 Implement tokenless HuggingFace Space assignment execution with explicit discovery visibility filtering and non-retryable inaccessible results, without leaking server discovery credentials
|
||||
- [x] 4.3 Extend packaged-worker verification for Windows and Linux with synthetic DockerHub and HuggingFace claim-to-ingestion flows
|
||||
- [x] 4.4 Add bounded canary configuration and assertions for search, enqueue, claim, worker-classified provider failures, upload, ingestion, projection, replay, expiry, and drain without adding per-target server access proofs
|
||||
|
||||
## 5. Shared Runtime Document Validation
|
||||
|
||||
- [x] 5.1 Add a side-effect-free bounded YAML loader with duplicate-key rejection and secret-safe errors
|
||||
- [x] 5.2 Define strict configuration, core-profile, auth-pool, reference-integrity, package-capability, and deployment-path validation
|
||||
- [x] 5.3 Use the shared validator in preview, runtime startup, and secrets import without weakening existing runtime security checks
|
||||
- [x] 5.4 Implement fixed config/secrets candidate storage with private durable writes, SHA-256 compare-and-swap, and bounded redacted diffs
|
||||
- [x] 5.5 Add validation and concurrency tests for unknown keys, duplicate keys, invalid references, stale active hashes, stale candidates, and error redaction
|
||||
|
||||
## 6. Typed Supervisor and Operations Services
|
||||
|
||||
- [x] 6.1 Extend Supervisor control protocol with structured runtime/source snapshots and exact managed-source lifecycle actions
|
||||
- [x] 6.2 Add bounded log-tail actions keyed only by allowlisted managed source IDs and reject generic web command forwarding
|
||||
- [x] 6.3 Implement operation creation, lifecycle transition, bounded result reconciliation, and append-only audit service methods
|
||||
- [x] 6.4 Add Supervisor and operation-service tests for malformed actions, unknown sources, log bounds, restart races, and content-free audit records
|
||||
|
||||
## 7. Web Operations Console
|
||||
|
||||
- [x] 7.1 Add trusted Caddy operator-header stripping/injection and backend actor validation tied to the private edge marker
|
||||
- [x] 7.2 Add shared admin navigation and overview page with bounded runtime, queue, assignment, bundle, control, and operation health
|
||||
- [x] 7.3 Add Search pages and exact forms for persistent discovery controls plus typed producer lifecycle and interval actions
|
||||
- [x] 7.4 Add Workers/Dispatch pages and exact forms for pause/resume, drain start/cancel/progress, package compatibility, users, devices, and assignments
|
||||
- [x] 7.5 Add Supervisor status/action and bounded log pages without shell, path, or generic command inputs
|
||||
- [x] 7.6 Add config and plaintext secrets preview/save/apply pages with no-store rendering, revision/hash conflicts, and no value leakage outside the editor
|
||||
- [x] 7.7 Add durable operation-status and bounded paginated audit pages that survive runtime restart
|
||||
- [x] 7.8 Add admin API and browser tests for routes, methods, exact form shapes, actor spoofing, Origin/CSRF, stale forms, navigation, CSP, and secret non-retention
|
||||
|
||||
## 8. Managed File Service
|
||||
|
||||
- [x] 8.1 Define logical managed roots and per-root list/read/create-replace/delete permissions with bounded path, listing, and byte limits
|
||||
- [x] 8.2 Implement descriptor-relative Linux traversal with no-follow component opens and rejection of absolute, dot, drive, backslash, symlink, hardlink, and special-file targets
|
||||
- [x] 8.3 Implement bounded download, durable compare-and-swap create/replace, and expected-hash delete with private temporary files and directory fsync
|
||||
- [x] 8.4 Add typed Files pages/forms and content-free mutation audit events while keeping config, secrets, database, sockets, agent metadata, and raw bundles excluded
|
||||
- [x] 8.5 Add adversarial traversal, encoded traversal, symlink-swap, hardlink, special-file, limit, concurrent-replacement, and forbidden-root tests
|
||||
|
||||
## 9. Privileged Host Operations Agent
|
||||
|
||||
- [x] 9.1 Implement a root-owned Unix-socket agent with peer-credential checks and a closed request schema containing only operation ID, action enum, and expected hashes
|
||||
- [x] 9.2 Implement singleton locking, persisted-operation verification, host-side revalidation, fixed-path backups, and atomic config/secrets replacement
|
||||
- [x] 9.3 Implement fixed runtime/edge stop and recreation plus bounded health verification for Supervisor, PostgreSQL, Worker API, ingester, and projector
|
||||
- [x] 9.4 Implement byte-identical automatic rollback and failed-hold behavior without arbitrary services, paths, commands, or retry loops
|
||||
- [x] 9.5 Add systemd socket/service units, fixed host-managed config/candidate directories, permissions, and deployment installer validation
|
||||
- [x] 9.6 Add crash-boundary and hostile-request tests for validation, backup, replacement, restart, health failure, rollback success, rollback failure, and request-field injection
|
||||
|
||||
## 10. Deployment Migration and Verification
|
||||
|
||||
- [x] 10.1 Update container, code-authority, package, Caddy, systemd, and deployment fixtures for all new modules, profiles, routes, headers, sockets, and fixed managed paths
|
||||
- [x] 10.2 Add an offline migration/rollback test proving the previous image can coexist with additive tables after protocol-2 work is drained
|
||||
- [x] 10.3 Run focused unit and PostgreSQL integration suites, packaged worker verification, edge E2E, browser coverage, strict OpenSpec validation, and diff checks
|
||||
- [x] 10.4 On the approved production host, deploy one exact supported ingress profile with discovery and dispatch paused, verify protocol-2 packages plus runtime/admin/edge/agent health, complete one reconciled lifecycle operation and one successful automatic rollback drill, and only then issue or resume new work
|
||||
- [x] 10.5 Run the bounded DockerHub end-to-end canary and retain evidence for queue authority, ingestion, projection, expiry/replay, and drain
|
||||
- [x] 10.6 Enable GitLab and public HuggingFace only after canary gates pass, confirm GitHub remains excluded, and document rollback evidence
|
||||
- [x] 10.7 Add exact root-owned `standalone-edge-v1` and `shared-host-edge-v1` deployment profiles without changing the host-agent request schema
|
||||
- [x] 10.8 Add the shared-host Compose/Caddy route confinement and profile-specific lifecycle, installer, and denylist validation while preserving standalone behavior
|
||||
- [x] 10.9 Add dual-profile tests, deployment documentation, strict validation, and real hardened Caddy/Compose checks
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-23
|
||||
@@ -0,0 +1,206 @@
|
||||
## Context
|
||||
|
||||
The protocol-2 remote worker has proven its authoritative data path under real production load, including native Windows concurrency 3 and simultaneous Windows/WSL execution. Its operator experience has not reached the same level: the package exposes only `--server`, `--token`, and `--parallelism`; successful work is mostly silent; scanner output is captured out of view; slot state remains `assigned` during synchronous execution; and the server sees no phase progress between claim and result.
|
||||
|
||||
Two production DockerHub assignments remained unresolved until the fixed 7,200-second assignment deadline even though the configured Docker target timeout was 600 seconds. The current watchdog terminates the owned TruffleHog process tree, but permit acquisition and surrounding Python work such as cleanup, filtering, result serialization, bundle staging, and handoff are not one hard-preemptible unit. Current evidence cannot identify which phase stalled.
|
||||
|
||||
The administration page compounds the problem by labeling `result_reservations.last_error_code` as `Error category`. That field describes assignment transport failures and expiry, while accepted scan errors live in `target_scans`, `errors`, and result metadata and are not shown as assignment diagnostics.
|
||||
|
||||
Legacy `supervisor.py --attach` demonstrates a useful interaction model: a verified background instance, an authenticated local control channel, a live status table, bounded logs, and detach without shutdown. The remote-worker package does not contain that supervisor and needs a smaller worker-specific implementation.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Deliver one coherent operator product rather than temporary UI over incomplete fields.
|
||||
- Give owners a cross-platform lifecycle CLI with detached operation, attach, status, logs, history, diagnostics, and doctor commands.
|
||||
- Define one phase/event model used locally, over the Worker API, in PostgreSQL, and in the admin UI.
|
||||
- Apply a hard local scan-stage deadline to the complete assignment execution unit, not only its scanner child process.
|
||||
- Preserve exact diagnostic material when available and explicitly describe size truncation or other transformations.
|
||||
- Separate assignment transport outcome, scan outcome, and diagnostics in storage and UI.
|
||||
- Make global and per-source assignment deadline policy visible and measurable with phase duration percentiles.
|
||||
- Ship a from-zero operator guide and fault-injection coverage as part of the same change.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Replacing PostgreSQL assignment authority, immutable assignment expiry, durable receipts, or bundle ingestion.
|
||||
- Renewing assignment ownership from progress events.
|
||||
- Reporting fabricated percentage completion when the scanner has no reliable denominator.
|
||||
- Rebuilding the full server runtime supervisor inside the worker package.
|
||||
- Adding a temporary compatibility-only error page that will be removed after diagnostics land.
|
||||
- Adding a new diagnostic masking or redaction subsystem.
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. One worker supervisor owns lifecycle and authoritative local projections
|
||||
|
||||
The package will expose a `truf-worker` command with `install`, `run`, `start`, `stop`, `status`, `attach`, `logs`, `history`, and `doctor` subcommands. `run` executes the supervisor in the foreground; `start` launches the same supervisor detached and waits for a startup handshake. Docker continues to run the supervisor in the foreground under Tini while Docker supplies detachment.
|
||||
|
||||
The supervisor owns the singleton lock, slot controllers, local control endpoint, rotating human log, event JSONL, current status snapshot, and terminal local history. A verified instance record binds PID, process creation identity, executable, package identity, local control endpoint, and lifecycle state. `attach` reads status/events through the local control channel; it does not attach to arbitrary stdout and detaching never stops the worker.
|
||||
|
||||
The current three-flag invocation remains a foreground-run migration alias for already shipped package launch definitions, but all new documentation and generated launchers use explicit subcommands.
|
||||
|
||||
Alternative considered: add more prints to `remote_worker_client.py`. Rejected because prints cannot provide detached lifecycle control, reliable concurrent-slot rendering, machine output, history, or process identity.
|
||||
|
||||
### 2. A single append-only event model drives every projection
|
||||
|
||||
Each state transition emits a versioned event with a monotonic local sequence:
|
||||
|
||||
```json
|
||||
{
|
||||
"schema": 1,
|
||||
"sequence": 42,
|
||||
"timestamp": "2026-09-23T12:00:00Z",
|
||||
"instance_id": "...",
|
||||
"slot_id": 0,
|
||||
"reservation_id": 123,
|
||||
"source": "dockerhub",
|
||||
"type": "slot.phase",
|
||||
"phase": "scanning",
|
||||
"phase_started_at": "...",
|
||||
"scan_deadline_at": "...",
|
||||
"assignment_deadline_at": "...",
|
||||
"progress": {}
|
||||
}
|
||||
```
|
||||
|
||||
Canonical phases are `idle`, `claiming`, `assigned`, `waiting_permit`, `preparing`, `resolving`, `downloading`, `cloning`, `scanning`, `filtering`, `cleaning`, `bundling`, `uploading`, `awaiting_receipt`, `backoff`, `draining`, and `stopped`.
|
||||
|
||||
The local status JSON is a rebuildable projection of the event stream, not a second independently written state machine. Server progress uses the same event schema with a server-assigned receive timestamp. Terminal history records authoritative receipt/prebundle/stale outcomes plus durations.
|
||||
|
||||
Alternative considered: define separate local, API, and admin status shapes. Rejected because they would drift and recreate the current ambiguity.
|
||||
|
||||
### 3. Execute each scan stage in a supervised assignment runner process
|
||||
|
||||
Threads in the controller continue to own claim/recovery/upload state, but scanner execution moves into a package-local assignment runner subprocess. The controller writes one validated input record and starts the runner in a contained process tree with a dedicated work/output directory.
|
||||
|
||||
The hard scan-stage deadline starts before permit acquisition and covers:
|
||||
|
||||
- permit wait;
|
||||
- assignment preparation and source resolution performed locally;
|
||||
- downloads/clones;
|
||||
- scanner subprocess execution;
|
||||
- filtering and result conversion;
|
||||
- cleanup;
|
||||
- result and bundle staging.
|
||||
|
||||
The runner emits phase events over a local pipe/file protocol. At the deadline the controller terminates the runner process tree, atomically detaches its work directory for later janitor handling, and creates a normal timeout result bundle with the phase and diagnostic envelope. The controller then uploads that result while the immutable assignment deadline still permits it.
|
||||
|
||||
Normal cleanup remains cooperative. Cleanup that exceeds the deadline cannot keep the slot occupied; abandoned work is moved into a janitor-owned tree and reported in status.
|
||||
|
||||
Alternative considered: add deadline checks around existing Python calls. Rejected because blocking Python/filesystem/library calls cannot be hard-preempted reliably in the controller process.
|
||||
|
||||
### 4. Progress events are informational and never renew authority
|
||||
|
||||
The Worker API gains an authenticated progress endpoint accepting ordered phase events for the worker's current reservation. It validates reservation/device ownership and monotonically advances the latest accepted event sequence. Duplicate events are idempotent.
|
||||
|
||||
Progress updates do not change `remote_expires_at`, queue leases, bundle capacity, or receipt authority. Failure to send progress does not abort a healthy local scan; events remain in the local journal and retry with bounded backoff. The server can therefore display `last phase` and `last progress` without turning progress into a lease heartbeat.
|
||||
|
||||
Alternative considered: renewable leases. Rejected because a stuck client could retain work indefinitely and because the fixed-expiry fencing model is already proven.
|
||||
|
||||
### 5. Server-owned deadlines support global fallback and per-source policy
|
||||
|
||||
`supervisor.worker_api.assignment_ttl_seconds` remains the global fallback. A managed per-source override map adds GitLab, DockerHub, and HuggingFace assignment TTL values. The server selects and commits the immutable deadline when issuing an assignment and includes the effective scan and assignment deadlines in the response.
|
||||
|
||||
The editor labels these separately as `Target scan timeout`, `Assignment deadline (end-to-end)`, and `Result upload body deadline`. Validation retains absolute bounds and checks that each effective assignment deadline covers its source scan timeout, upload deadline, and handoff margin.
|
||||
|
||||
The server aggregates phase and end-to-end durations by source and outcome as p50, p95, and p99. Configuration remains explicit; metrics inform changes but do not silently rewrite policy.
|
||||
|
||||
Alternative considered: let each worker select or renew its TTL. Rejected because workload policy belongs to the assigning server and must be consistent for queue fencing.
|
||||
|
||||
### 6. One diagnostic envelope spans scan and prebundle failures
|
||||
|
||||
Diagnostics use one versioned envelope with indexed dimensions and exact optional payloads:
|
||||
|
||||
- diagnostic ID and schema;
|
||||
- reservation, scan event, slot, attempt, and source;
|
||||
- phase, kind, category, stable code, summary, and retryable disposition;
|
||||
- provider operation and HTTP status/content type/request ID;
|
||||
- process name, exit code, signal, and timeout state;
|
||||
- exception type/message/fingerprint;
|
||||
- raw body material;
|
||||
- log stream head/tail material;
|
||||
- occurred/captured/received timestamps;
|
||||
- original byte count, stored byte count, content hash, and truncation state.
|
||||
|
||||
Categories are broad query dimensions such as `authorization`, `rate_limit`, `not_found`, `network`, `timeout`, `scanner`, `storage`, `protocol`, and `assignment_expired`. Stable codes express the concrete cause, such as `docker.manifest_http_403` or `trufflehog.exit_nonzero`. Phase is independent of category.
|
||||
|
||||
Text/bytes are preserved as captured. Non-text bodies use an explicit encoding field. Storage bounds are deterministic: body excerpt 16 KiB, combined log head/tail 32 KiB, one transmitted envelope 64 KiB, at most 32 diagnostics and 256 KiB per assignment. Metadata records every truncation; no truncation is presented as a complete body.
|
||||
|
||||
Accepted result bundles gain a diagnostic frame. Prebundle reports carry the same envelope under the existing JSON body limit. The old E-frame error string remains ingestible during rollout but is projected into the new model exactly once.
|
||||
|
||||
Alternative considered: keep scanner error strings, reservation error codes, and source metadata as separate taxonomies. Rejected because operators cannot correlate or filter them consistently.
|
||||
|
||||
### 7. Local diagnostics retain complete operator evidence when available
|
||||
|
||||
The supervisor writes:
|
||||
|
||||
```text
|
||||
control/worker.instance.json
|
||||
control/worker.status.json
|
||||
events/worker-events.jsonl
|
||||
history/worker-history.jsonl
|
||||
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.json
|
||||
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.body
|
||||
diagnostics/YYYY-MM-DD/<reservation>/<diagnostic>.log
|
||||
logs/worker.log
|
||||
```
|
||||
|
||||
The JSON envelope points to optional body/log files and records their hashes and sizes. Local retention is configurable by age and total bytes and is reported by `status` and `doctor`. Rotation never mutates terminal history entries; it changes attached-artifact availability explicitly.
|
||||
|
||||
Alternative considered: store every raw artifact directly in one JSONL. Rejected because large multiline/process output makes append recovery and bounded tailing expensive.
|
||||
|
||||
### 8. PostgreSQL stores diagnostics as first-class records
|
||||
|
||||
Add `worker_diagnostics` with an idempotent diagnostic UID, reservation FK, optional target-scan FK, indexed phase/category/code/kind/retryable columns, summary, canonical envelope JSON, received timestamp, and optional bounded body/log payload columns. Bundle ingestion writes diagnostics in the same transaction as the target scan and error projection. Prebundle diagnostics attach to the reservation before a target scan exists.
|
||||
|
||||
Existing `errors` rows remain the compatibility scan-error projection. Existing `last_error_code` is retained as assignment failure code but is no longer labeled as the complete error category.
|
||||
|
||||
Alternative considered: put all envelopes only into `metadata_json`. Rejected because filtering, detail lookup, idempotency, and prebundle diagnostics require first-class rows.
|
||||
|
||||
### 9. Admin views assignment outcome, scan outcome, progress, and diagnostics separately
|
||||
|
||||
The worker assignment table presents:
|
||||
|
||||
- assignment outcome: unfinished, accepted, prebundle failed, expired;
|
||||
- scan outcome: clean, found, degraded, error, skipped, or not available;
|
||||
- diagnostic count and highest-priority category/code;
|
||||
- current/latest phase, phase age, assignment deadline, and last progress age;
|
||||
- worker/device/package identity and slot where available.
|
||||
|
||||
Each row links to a detail page with an event timeline, duration breakdown, transport/receipt data, scan summary, diagnostics, raw body/log tabs, canonical JSON copy/download, and explicit truncation metadata. Filters operate independently on assignment outcome, scan outcome, source, phase, category, code, retryability, and time.
|
||||
|
||||
Alternative considered: make the existing `Error category` cell open a modal. Rejected because the list model itself conflates transport and scan semantics.
|
||||
|
||||
### 10. Human output and machine output are equal product contracts
|
||||
|
||||
Human status/attach uses a stable table and event stream. It displays elapsed time, configured scan deadline, assignment time remaining, last progress age, child state, and trustworthy counters. It never fabricates completion percentages.
|
||||
|
||||
`--json` commands emit one versioned JSON document. Follow modes emit NDJSON with one event per line and no decorative output. Tests treat both output forms as contracts.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Runner process split touches scanner lifecycle and recovery paths] -> Introduce it behind the same `execute_protocol2_remote_claim` contract, preserve deterministic bundle validation, and fault-test every boundary before replacing in-process execution.
|
||||
- [Progress traffic increases database writes] -> Persist only monotonic phase transitions and coarse progress changes, deduplicate by reservation/sequence, and keep high-frequency local samples local.
|
||||
- [Diagnostic payloads increase bundle and database volume] -> Enforce deterministic per-item/per-assignment byte and count limits, expose truncation metadata, and track storage usage.
|
||||
- [One broad change can take too long] -> Implement as large vertical chunks that each finish a final architecture slice; do not ship throwaway status/error models.
|
||||
- [Per-source policy adds configuration complexity] -> Keep one global fallback, explicit source overrides, editor-derived effective values, and validation based on existing scan/upload settings.
|
||||
- [Local full-stage termination can leave work trees] -> Atomically detach them to janitor ownership and surface retained bytes/counts in status and doctor.
|
||||
- [Old packages do not emit progress/diagnostics] -> Admin renders legacy records from existing fields and marks phase/diagnostic availability explicitly until packages are upgraded.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add schema/event/diagnostic libraries, PostgreSQL tables, and read paths without changing current assignment execution.
|
||||
2. Add the worker supervisor CLI and local projections while the current foreground invocation remains a migration alias.
|
||||
3. Add assignment runner execution, full-stage watchdog, local phase events, and fault-injection tests.
|
||||
4. Add Worker API progress and diagnostic transport, then enable server persistence and detail queries.
|
||||
5. Replace the worker admin list/detail presentation and add percentile/deadline editor views.
|
||||
6. Rebuild Windows/Linux packages, run protocol compatibility tests, then run bounded Windows and WSL production validation.
|
||||
7. Update generated launchers and the from-zero operator guide; migrate the production worker launch definition to explicit `run`.
|
||||
8. Remove the migration alias only in a separately declared breaking change after all known deployments use subcommands.
|
||||
|
||||
Rollback disables progress ingestion and runner selection while retaining additive event/diagnostic tables. Existing immutable assignment, bundle, receipt, and legacy E-frame paths remain authoritative throughout rollout.
|
||||
|
||||
## Open Questions
|
||||
|
||||
No blocking product questions remain. Exact local retention defaults and percentile windows can be selected from implementation benchmarks without changing the external contracts above.
|
||||
@@ -0,0 +1,38 @@
|
||||
## Why
|
||||
|
||||
Remote workers execute and settle production work correctly, but they behave as opaque background processes: an owner cannot tell whether a slot is downloading, scanning, cleaning up, bundling, uploading, stalled, or close to its deadline, while the admin console conflates assignment failures with scan errors. The next change should turn the proven worker engine into an understandable operator-facing product without building temporary UI around the current incomplete status and error fields.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a worker-specific supervisor and CLI for install, foreground run, detached start, graceful stop, status, attach, live logs, local history, and diagnostics.
|
||||
- Add one versioned assignment phase/event model shared by local status, attach, server progress, history, diagnostics, and admin rendering.
|
||||
- Apply a hard local scan-stage deadline across permit acquisition, scanner execution, cleanup, filtering, and result staging instead of bounding only the scanner subprocess path.
|
||||
- Make assignment deadlines observable and configurable as server-owned global and per-source policy, with explicit scan, upload, and assignment deadline semantics.
|
||||
- Add a versioned diagnostic envelope for provider responses, process failures, exceptions, timeout state, structured categories/codes, and bounded raw body/log material.
|
||||
- Persist local worker events and diagnostics as JSON/JSONL and transmit the same diagnostic model through terminal reports and accepted result bundles.
|
||||
- Replace the ambiguous assignment-table `Error category` presentation with distinct assignment outcome, scan outcome, diagnostic count, and a clickable detail/timeline view.
|
||||
- Add machine-readable `--json`/NDJSON output and human-readable status/attach views without inventing progress percentages.
|
||||
- Add a from-zero operator guide covering installation, first start, attach/status, interpreting phases and errors, graceful drain/stop, recovery, update, and troubleshooting.
|
||||
- Do not add a separate masking/redaction feature or silently generalize captured diagnostics. Storage bounds and any unavoidable transformation must be explicit in diagnostic metadata.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `worker-operator-supervisor`: Lifecycle CLI, detached supervision, attach/status/log/history/doctor commands, local projections, and operator documentation.
|
||||
- `worker-progress-deadlines`: Canonical assignment phases, local and server progress events, full scan-stage watchdog behavior, deadline policy, and duration percentile observability.
|
||||
- `worker-diagnostics`: Unified diagnostic envelope, local diagnostic archive, terminal/bundle transport, taxonomy, retention bounds, and exact transformation metadata.
|
||||
- `worker-admin-experience`: Assignment and scan outcome separation, diagnostic persistence/querying, clickable timeline/detail UI, and human/machine-readable diagnostic views.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None. There is no synchronized main-spec directory in this workspace; the new capabilities define the operator-facing contract over the existing distributed worker implementation.
|
||||
|
||||
## Impact
|
||||
|
||||
- Worker client/bootstrap/package entrypoints and generated Windows/Linux launchers.
|
||||
- Scanner subprocess lifecycle, scan-slot ownership, cleanup, bundle staging, and local runtime state.
|
||||
- Worker API progress, prebundle report, result-bundle, and assignment-status contracts.
|
||||
- PostgreSQL reservation, scan, error, diagnostic, and observability queries/schema.
|
||||
- Server-rendered worker administration pages and runtime deadline editor fields.
|
||||
- Windows process supervision and Linux/WSL/Docker foreground/detached operation.
|
||||
- Worker package tests, protocol tests, fault-injection tests, admin UI tests, and operator documentation.
|
||||
@@ -0,0 +1,207 @@
|
||||
# Remote Worker Operator Experience Findings
|
||||
|
||||
Recorded: 2026-09-23.
|
||||
|
||||
This file preserves the investigation that led to `add-worker-operator-experience` so implementation does not have to rediscover current behavior and terminology.
|
||||
|
||||
## Production Validation Facts
|
||||
|
||||
- Native Windows protocol-2 worker reached exact concurrency 3 and never exceeded it.
|
||||
- Its bounded production cohort accepted 180 assignments: DockerHub 63, GitLab 65, HuggingFace 52.
|
||||
- All 180 accepted results ingested, queue-settled, projected, and reconciled without global lineage violations or quarantine.
|
||||
- A later dual-worker run proved Windows max 1, WSL max 1, combined max 2.
|
||||
- Two WSL DockerHub assignments in separate runs stayed unresolved until fixed two-hour lease expiry and produced no accepted bundle.
|
||||
- The server expired/refunded those reservations correctly; the missing information is where the worker spent the time before expiry.
|
||||
|
||||
Durable validation report:
|
||||
|
||||
`docs/worker-parallelism-validation-2026-09-23.md`
|
||||
|
||||
## Timeout and Lease Map
|
||||
|
||||
### Assignment deadline
|
||||
|
||||
Production explicitly used:
|
||||
|
||||
```yaml
|
||||
supervisor:
|
||||
worker_api:
|
||||
assignment_ttl_seconds: 7200
|
||||
```
|
||||
|
||||
- Production value: 7,200 seconds.
|
||||
- Code/template fallback: 86,400 seconds.
|
||||
- Managed validation bounds: 60 through 604,800 seconds.
|
||||
- Validation requires the effective TTL to cover the maximum configured source scan timeout, bundle upload deadline, and 60 seconds of handoff margin.
|
||||
- The server commits one immutable expiry at assignment issuance.
|
||||
- Ordinary worker contacts and progress do not renew it.
|
||||
- Relevant code: `app/runtime_document.py`, `app/worker_api.py`, `app/worker_assignment.py`, `app/scanner_db.py`.
|
||||
|
||||
### Docker target timeout
|
||||
|
||||
Canonical managed full-runtime and production value:
|
||||
|
||||
```yaml
|
||||
sources:
|
||||
dockerhub:
|
||||
timeout: 600
|
||||
```
|
||||
|
||||
- Managed legacy Windows/full runtime: 600 seconds.
|
||||
- Current production profile: 600 seconds.
|
||||
- Direct CLI/function fallback without managed config: 1,800 seconds.
|
||||
- Legacy Windows Job containment terminates the owned TruffleHog process tree at the subprocess deadline.
|
||||
- Historical evidence: `tests/test_scanner_queue_high_fixes.py`, `tests/test_validated_high_scanner_fixes.py`, and `openspec/changes/stabilize-dockerhub-trufflehog-lifecycle/`.
|
||||
|
||||
### Other relevant timers
|
||||
|
||||
- Worker API client socket timeout: 120 seconds.
|
||||
- Result upload absolute body deadline: 1,800 seconds.
|
||||
- Result upload idle timeout: 30 seconds.
|
||||
- Remote assignment expiry reaper interval: 60 seconds.
|
||||
- Empty-claim `Retry-After`: normally 5 seconds.
|
||||
- Result-ingester/projector leases: 300 seconds after upload, unrelated to pre-upload scanning.
|
||||
|
||||
### Why ten minutes became two hours
|
||||
|
||||
The 600-second budget hard-preempts the TruffleHog subprocess path, but it is not one hard preemptive boundary around every surrounding Python operation. Permit acquisition, process startup/termination recovery, cleanup, filtering, serialization, bundle staging/fsync, and handoff can outlive the subprocess deadline. The remote client checks assignment expiry before synchronous execution and then has no phase heartbeat or cancellation loop.
|
||||
|
||||
The evidence does not prove which phase stalled. Increasing the assignment TTL would only make the unknown stall retain ownership longer. The implementation needs phase instrumentation and a complete supervised execution-unit deadline.
|
||||
|
||||
## Current Worker Experience
|
||||
|
||||
Current supported CLI flags in `app/remote_worker_client.py`:
|
||||
|
||||
```text
|
||||
--server
|
||||
--token
|
||||
--parallelism
|
||||
```
|
||||
|
||||
There are no worker `start`, `stop`, `status`, `attach`, `logs`, `history`, `doctor`, or JSON-output commands.
|
||||
|
||||
Current behavior:
|
||||
|
||||
- one daemon thread per slot;
|
||||
- slot recovery state in `slot-N.json`;
|
||||
- ready bundles retained until an authoritative receipt;
|
||||
- scanner call is synchronous from the slot controller;
|
||||
- state remains broadly `assigned` until bundle readiness;
|
||||
- client loop failures print generic lines without a structured timeline;
|
||||
- successful claim/scan/upload/receipt is mostly silent;
|
||||
- scanner stdout/stderr is captured in temporary files and returned only after process completion;
|
||||
- exact Git execution can emit no useful start line;
|
||||
- concurrent messages can interleave;
|
||||
- server `last_contact_at` does not prove or disprove active scanner progress.
|
||||
|
||||
Default paths:
|
||||
|
||||
- Windows: `%LOCALAPPDATA%/TRUF/RemoteWorker`.
|
||||
- Linux: `$XDG_STATE_HOME/truf/remote-worker` plus `$XDG_DATA_HOME/truf/remote-worker`.
|
||||
- Docker production convention: persistent `/data` volume with separate state/data roots.
|
||||
|
||||
## Legacy Attach Pattern
|
||||
|
||||
The old full-runtime supervisor has an `--attach` implementation in `app/supervisor.py` and a wrapper `attach_runtime.ps1`. The wrapper is currently disabled and the supervisor is not in the remote-worker package.
|
||||
|
||||
Useful semantics to reuse:
|
||||
|
||||
- exact background-instance identity;
|
||||
- startup and loopback control handshake;
|
||||
- initial status table;
|
||||
- interactive `attach>` prompt;
|
||||
- alternate-screen `watch` table;
|
||||
- bounded log tail;
|
||||
- `q`/EOF/Ctrl-C detach without worker shutdown;
|
||||
- explicit coordinated shutdown command.
|
||||
|
||||
The worker needs a smaller implementation over its own event/status model, not a copy of the complete server supervisor.
|
||||
|
||||
## Current Error Model
|
||||
|
||||
The admin `Error category` column is rendered in `app/admin_api.py` from:
|
||||
|
||||
```sql
|
||||
result_reservations.last_error_code AS error_category
|
||||
```
|
||||
|
||||
It therefore describes assignment/transport errors such as remote prebundle failure or assignment expiry. It is not `errors.category`, scanner `error_class`, or `source_failure_category`.
|
||||
|
||||
Consequences:
|
||||
|
||||
- an accepted bundle is shown as assignment `completed` even when scan status is `error`;
|
||||
- accepted scan errors usually leave assignment `Error category` blank;
|
||||
- process versus storage prebundle failure survives in resolution JSON but is not shown;
|
||||
- permanent provider skips can exist in metadata/warnings without an `errors` row;
|
||||
- provider response bodies are inconsistently reduced or discarded.
|
||||
|
||||
Useful data already persisted but not presented together:
|
||||
|
||||
- `result_reservations`: resolution kind/JSON, receipt, issue/expiry/resolve timestamps, last error code/detail;
|
||||
- `target_scans`: status, error count, skipped reason, first error summary, start/end/duration;
|
||||
- `errors`: category, summary, raw selected error line;
|
||||
- `scan_result_compat.metadata_json`: error class, retryability, source failure category, warnings, degraded/skipped flags, process return/timeout/output metadata;
|
||||
- `keycheck_results`: provider status group/message/metadata.
|
||||
|
||||
## Diagnostic Direction
|
||||
|
||||
Use separate dimensions rather than one overloaded category:
|
||||
|
||||
```text
|
||||
assignment outcome: accepted | prebundle_failed | expired | unfinished
|
||||
scan outcome: clean | found | degraded | error | skipped | unavailable
|
||||
phase: scanning | cleaning | bundling | uploading | ...
|
||||
kind: provider_http | scanner_process | exception | storage | protocol | ...
|
||||
category: authorization | rate_limit | timeout | network | scanner | ...
|
||||
code: stable concrete identifier
|
||||
retryable: true | false
|
||||
```
|
||||
|
||||
Candidate transmitted limits from the investigation:
|
||||
|
||||
- provider body material: 16 KiB;
|
||||
- process log head/tail: 32 KiB combined;
|
||||
- one diagnostic envelope: 64 KiB;
|
||||
- at most 32 diagnostics and 256 KiB total per assignment;
|
||||
- prebundle envelope profile sized to fit the existing Worker API JSON limit.
|
||||
|
||||
The local worker archive can retain larger/full artifacts under configurable age and byte rotation. Every transport transformation must state original size, stored size, hash, encoding, and truncation state.
|
||||
|
||||
## Progress and Duration Direction
|
||||
|
||||
Minimum phase transitions needed to explain the two-hour event:
|
||||
|
||||
```text
|
||||
assignment_received
|
||||
scan_permit_acquired
|
||||
runner_started
|
||||
source_prepare_started/completed
|
||||
scanner_started/exited
|
||||
filtering_started/completed
|
||||
cleanup_started/completed
|
||||
bundle_started/ready
|
||||
upload_started/acknowledged
|
||||
```
|
||||
|
||||
Progress is evidence only and does not renew the immutable lease.
|
||||
|
||||
Duration percentiles:
|
||||
|
||||
- p50: median duration;
|
||||
- p95: 95 percent of observations finish at or below this duration;
|
||||
- p99: 99 percent finish at or below it.
|
||||
|
||||
Compute them separately by source, phase, outcome, and time window, with sample counts. They describe observed behavior and inform policy; they do not silently set policy.
|
||||
|
||||
## Product Decisions
|
||||
|
||||
- Build one final operator architecture rather than a temporary admin patch.
|
||||
- Implement it in large vertical chunks that each remain part of the final system.
|
||||
- Use a worker-specific supervisor and local event stream.
|
||||
- Split scanner execution into a supervised per-assignment runner process so the complete scan stage can be hard-preempted.
|
||||
- Keep immutable server assignment ownership and make progress non-renewing.
|
||||
- Add global fallback plus per-source assignment TTL policy.
|
||||
- Use one diagnostic envelope across local files, terminal reports, bundles, PostgreSQL, API, and UI.
|
||||
- Preserve diagnostic fidelity and expose all explicit transformations.
|
||||
- Separate assignment outcome, scan outcome, and diagnostics in the admin UI.
|
||||
- Include from-zero operator documentation and fault injection in the same change.
|
||||
@@ -0,0 +1,71 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Separate assignment and scan outcomes
|
||||
The worker administration list SHALL display assignment transport outcome, scan outcome, and diagnostic count as separate fields and SHALL not label `last_error_code` as the complete scan error category.
|
||||
|
||||
#### Scenario: Accepted scan has an error outcome
|
||||
- **WHEN** a reservation has a durable accepted bundle whose target scan status is `error`
|
||||
- **THEN** the list SHALL show assignment `accepted`, scan `error`, and the diagnostic count/categories
|
||||
|
||||
#### Scenario: Assignment expires before scan ingestion
|
||||
- **WHEN** a reservation expires without an accepted bundle
|
||||
- **THEN** the list SHALL show assignment `expired`, scan outcome unavailable, and the expiry diagnostic
|
||||
|
||||
### Requirement: Assignment detail timeline
|
||||
Each worker assignment row SHALL link to a detail page that reconstructs issued, phase-progress, bundle/terminal report, receipt, ingestion, queue settlement, and projection timestamps that exist for that assignment.
|
||||
|
||||
#### Scenario: Administrator opens an active assignment
|
||||
- **WHEN** progress events exist for an unresolved assignment
|
||||
- **THEN** the page SHALL show current phase, phase age, last progress age, scan deadline, assignment deadline, and ordered prior phases
|
||||
|
||||
#### Scenario: Administrator opens a settled assignment
|
||||
- **WHEN** the assignment has been accepted and projected
|
||||
- **THEN** the timeline SHALL distinguish acceptance, ingestion, queue settlement, and projection completion rather than collapsing them into one completion time
|
||||
|
||||
### Requirement: Clickable diagnostic detail
|
||||
The assignment detail page SHALL list diagnostics and SHALL provide human summary, canonical envelope JSON, raw body view, process log view, transformation metadata, and copy/download actions for each diagnostic.
|
||||
|
||||
#### Scenario: Diagnostic body is complete
|
||||
- **WHEN** an HTTP diagnostic contains an untruncated body
|
||||
- **THEN** the raw-body view SHALL identify it as complete and display the captured content and metadata
|
||||
|
||||
#### Scenario: Diagnostic material is truncated
|
||||
- **WHEN** body or log material was size-truncated
|
||||
- **THEN** the view SHALL prominently display original/stored sizes, hash, and truncation state
|
||||
|
||||
#### Scenario: Legacy scan error has no diagnostic envelope
|
||||
- **WHEN** an older scan has only existing `errors.raw_error` or result metadata
|
||||
- **THEN** the detail page SHALL display those fields as legacy evidence and SHALL not invent a new envelope
|
||||
|
||||
### Requirement: Diagnostic filtering and grouping
|
||||
The admin UI SHALL filter independently by source, time, worker/device, assignment outcome, scan outcome, phase, category, stable code, and retryability and SHALL group repeated diagnostic fingerprints without hiding individual occurrences.
|
||||
|
||||
#### Scenario: Administrator filters rate-limit errors
|
||||
- **WHEN** category `rate_limit` and a time window are selected
|
||||
- **THEN** results SHALL include matching diagnostics regardless of whether their assignments were accepted or prebundle-failed
|
||||
|
||||
#### Scenario: Repeated diagnostics are grouped
|
||||
- **WHEN** multiple diagnostics share a fingerprint
|
||||
- **THEN** the UI SHALL show aggregate count and affected assignments while retaining links to each occurrence
|
||||
|
||||
### Requirement: Worker fleet status
|
||||
The admin UI SHALL display each worker's configured cap, active slots, current phases, package identity, latest contact/progress ages, pending local-recovery indication when reported, and known idle/backoff reason.
|
||||
|
||||
#### Scenario: Worker is scanning without recent API contact
|
||||
- **WHEN** a worker has an active assignment and recent progress events but its authentication contact timestamp is old
|
||||
- **THEN** fleet status SHALL show active progress rather than classifying the worker as idle solely from contact age
|
||||
|
||||
#### Scenario: Worker cannot claim due to capacity
|
||||
- **WHEN** the server rejects claims because a pipeline capacity axis is closed
|
||||
- **THEN** fleet status SHALL show the capacity reason instead of a generic offline/idle state
|
||||
|
||||
### Requirement: Deadline and duration administration
|
||||
The runtime editor and worker observability pages SHALL explain effective scan, upload, and assignment deadlines and SHALL show p50/p95/p99 duration metrics with sample counts by source and phase.
|
||||
|
||||
#### Scenario: Administrator edits a source assignment deadline
|
||||
- **WHEN** a per-source TTL candidate is previewed
|
||||
- **THEN** the editor SHALL show the effective policy, validation relationship to scan/upload bounds, and that only future assignments are affected
|
||||
|
||||
#### Scenario: Administrator compares policy to observations
|
||||
- **WHEN** sufficient phase-duration samples exist
|
||||
- **THEN** the page SHALL show the configured deadline alongside source-specific percentile values without automatically changing configuration
|
||||
@@ -0,0 +1,86 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Unified versioned diagnostic envelope
|
||||
Worker scan errors, provider failures, process failures, local exceptions, timeouts, prebundle failures, and assignment expiry context SHALL use one versioned diagnostic envelope with independent phase, kind, category, stable code, summary, retryability, attempt, timestamps, and optional HTTP/process/exception material.
|
||||
|
||||
#### Scenario: Provider returns an HTTP error
|
||||
- **WHEN** a provider operation receives an unsuccessful HTTP response
|
||||
- **THEN** the diagnostic SHALL identify the phase, provider operation, HTTP status/content type, stable category/code, retryability, and captured response material
|
||||
|
||||
#### Scenario: Scanner process fails
|
||||
- **WHEN** a scanner process exits unsuccessfully or is terminated at its deadline
|
||||
- **THEN** the diagnostic SHALL identify its process result, timeout/signal state, phase, stable category/code, and captured log material
|
||||
|
||||
#### Scenario: Assignment expires without a worker result
|
||||
- **WHEN** the server expires an unresolved assignment
|
||||
- **THEN** it SHALL create or expose a diagnostic describing assignment expiry and the last accepted progress phase without claiming a scanner error occurred
|
||||
|
||||
### Requirement: Diagnostic fidelity and explicit transformation
|
||||
Captured body and log bytes SHALL be preserved without silent semantic rewriting. Every size limit, encoding conversion, or truncation SHALL record original bytes, stored bytes, content hash, encoding, and truncation state.
|
||||
|
||||
#### Scenario: Text body fits the bound
|
||||
- **WHEN** a captured provider response body fits the configured diagnostic body bound
|
||||
- **THEN** the transmitted diagnostic SHALL contain the complete captured text and SHALL mark it untruncated
|
||||
|
||||
#### Scenario: Body exceeds the bound
|
||||
- **WHEN** captured body bytes exceed the transmitted bound
|
||||
- **THEN** the diagnostic SHALL carry the bounded material plus original/stored sizes, full captured-content hash when available, and `truncated=true`
|
||||
|
||||
#### Scenario: Body is not text
|
||||
- **WHEN** captured diagnostic body bytes are not valid text in the declared encoding
|
||||
- **THEN** the envelope SHALL use an explicit binary encoding representation and SHALL preserve the same transformation metadata
|
||||
|
||||
### Requirement: Bounded diagnostic transport
|
||||
The protocol SHALL enforce deterministic per-body, per-log, per-envelope, diagnostic-count, and aggregate diagnostic bounds while rejecting envelopes whose declared and actual sizes disagree.
|
||||
|
||||
#### Scenario: Accepted bundle contains diagnostics
|
||||
- **WHEN** a worker uploads a result bundle with diagnostic frames within all bounds
|
||||
- **THEN** bundle acceptance and ingestion SHALL validate and persist each diagnostic idempotently with the scan
|
||||
|
||||
#### Scenario: Diagnostic aggregate exceeds its limit
|
||||
- **WHEN** a bundle or terminal report exceeds a diagnostic count or byte limit
|
||||
- **THEN** the API SHALL reject it with a stable protocol error and SHALL NOT partially persist diagnostics
|
||||
|
||||
### Requirement: Prebundle and accepted-result parity
|
||||
The same diagnostic envelope SHALL be usable in prebundle terminal reports and accepted scan-result bundles, with only transport-size profiles differing.
|
||||
|
||||
#### Scenario: Worker storage fails before bundle creation
|
||||
- **WHEN** the worker cannot create a result bundle
|
||||
- **THEN** its terminal report SHALL include a diagnostic envelope rather than replacing the exception with one generic fixed detail string
|
||||
|
||||
#### Scenario: Scan returns errors in a valid bundle
|
||||
- **WHEN** scanning completes with structured errors and a valid bundle
|
||||
- **THEN** those errors SHALL be represented as diagnostics attached to the ingested target scan and SHALL remain distinct from assignment transport outcome
|
||||
|
||||
### Requirement: Deterministic diagnostic identity
|
||||
Each diagnostic SHALL have a deterministic UID derived from its canonical identity and content so retries and replay cannot create duplicates.
|
||||
|
||||
#### Scenario: Accepted upload is replayed
|
||||
- **WHEN** an identical accepted result bundle is uploaded again
|
||||
- **THEN** the server SHALL return the durable receipt and SHALL NOT insert duplicate diagnostic rows
|
||||
|
||||
#### Scenario: Same code occurs twice in one assignment
|
||||
- **WHEN** two distinct occurrences share category and code but differ in occurrence identity or content
|
||||
- **THEN** both SHALL be retained as distinct diagnostics with stable UIDs
|
||||
|
||||
### Requirement: Local diagnostic archive
|
||||
The worker SHALL retain a queryable local JSON diagnostic envelope and optional body/log artifacts per assignment, with configurable age/byte rotation and explicit artifact-availability state in history.
|
||||
|
||||
#### Scenario: Operator opens a local diagnostic
|
||||
- **WHEN** `truf-worker history` or `logs` selects a retained diagnostic
|
||||
- **THEN** the worker SHALL present the canonical envelope and exact paths/availability of its body and log artifacts
|
||||
|
||||
#### Scenario: Artifact rotates out
|
||||
- **WHEN** a body or log artifact is removed by configured local rotation
|
||||
- **THEN** terminal history SHALL remain and SHALL state that the artifact is no longer locally retained
|
||||
|
||||
### Requirement: Orthogonal error taxonomy
|
||||
The diagnostic model SHALL keep phase, kind, broad category, stable code, retryability, assignment outcome, and scan outcome as separate dimensions.
|
||||
|
||||
#### Scenario: Accepted scan has provider errors
|
||||
- **WHEN** a result bundle is durably accepted but the scan outcome is `error`
|
||||
- **THEN** the assignment outcome SHALL remain `accepted`, scan outcome SHALL be `error`, and provider diagnostics SHALL retain their own categories/codes
|
||||
|
||||
#### Scenario: Assignment expires
|
||||
- **WHEN** an assignment expires before bundle acceptance
|
||||
- **THEN** assignment outcome SHALL be `expired`, scan outcome SHALL be unavailable, and the expiry diagnostic SHALL not be categorized as a provider scan failure
|
||||
+74
@@ -0,0 +1,74 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Unified worker lifecycle CLI
|
||||
The worker package SHALL provide `install`, `run`, `start`, `stop`, `status`, `attach`, `logs`, `history`, and `doctor` commands with equivalent lifecycle semantics on supported Windows and Linux/WSL platforms.
|
||||
|
||||
#### Scenario: Operator starts a detached worker
|
||||
- **WHEN** an installed operator invokes `truf-worker start`
|
||||
- **THEN** the command SHALL launch the worker supervisor, wait for its startup handshake, and return the verified instance identity and current state
|
||||
|
||||
#### Scenario: Operator runs in the foreground
|
||||
- **WHEN** an operator invokes `truf-worker run`
|
||||
- **THEN** the same supervisor implementation SHALL run in the foreground and SHALL begin graceful drain on the first interrupt
|
||||
|
||||
### Requirement: Verified detached supervisor lifecycle
|
||||
The supervisor SHALL publish a versioned instance record, status projection, local control endpoint, startup result, and shutdown receipt tied to the exact running process identity.
|
||||
|
||||
#### Scenario: Status finds a stale instance record
|
||||
- **WHEN** the recorded process no longer matches the recorded executable, creation identity, or live control handshake
|
||||
- **THEN** `status` SHALL report the instance as stale and SHALL NOT represent it as a running worker
|
||||
|
||||
#### Scenario: Graceful stop has active slots
|
||||
- **WHEN** `stop` is requested while one or more slots own assignments
|
||||
- **THEN** the supervisor SHALL stop new claims, display the draining slots, and wait for terminal local reconciliation up to the requested stop deadline
|
||||
|
||||
### Requirement: Attachable live operator view
|
||||
The supervisor SHALL provide an `attach` session that renders current worker and per-slot state and follows new events without making attachment own the worker lifetime.
|
||||
|
||||
#### Scenario: Operator detaches
|
||||
- **WHEN** the operator presses `q`, sends EOF, or interrupts the attach client
|
||||
- **THEN** only the attach session SHALL end and the worker supervisor SHALL continue running
|
||||
|
||||
#### Scenario: Concurrent slots update
|
||||
- **WHEN** multiple slots emit interleaved phase events
|
||||
- **THEN** attach SHALL render one coherent row per slot and SHALL preserve event ordering by local sequence
|
||||
|
||||
### Requirement: Honest per-slot status
|
||||
Status and attach SHALL display source, phase, phase elapsed time, scan deadline, assignment time remaining, last progress age, and only counters measured by the execution path. They SHALL NOT synthesize percentage completion without a reliable denominator.
|
||||
|
||||
#### Scenario: Long scanner execution
|
||||
- **WHEN** a slot remains in `scanning` with a live runner process
|
||||
- **THEN** status SHALL continue updating elapsed time, deadline remaining, and last-progress age rather than appearing frozen
|
||||
|
||||
#### Scenario: Worker has no assignment
|
||||
- **WHEN** a slot is idle because of server backoff, cap, paused dispatch, capacity, or an empty queue
|
||||
- **THEN** status SHALL report the known idle/backoff reason and next claim time when supplied by the server
|
||||
|
||||
### Requirement: Human and machine output contracts
|
||||
Every non-interactive inspection command SHALL support versioned JSON output, and every follow command SHALL support versioned NDJSON output containing no human decoration.
|
||||
|
||||
#### Scenario: Automation requests status
|
||||
- **WHEN** `truf-worker status --json` is invoked
|
||||
- **THEN** stdout SHALL contain exactly one parseable versioned status object representing the same state as the human view
|
||||
|
||||
#### Scenario: Automation follows events
|
||||
- **WHEN** `truf-worker logs --follow --json` is invoked
|
||||
- **THEN** stdout SHALL contain one complete versioned event object per line in sequence order
|
||||
|
||||
### Requirement: Local history and operational diagnosis
|
||||
The supervisor SHALL retain terminal assignment history, rotating worker logs, event history, diagnostic references, and local storage usage, and SHALL expose them through `history`, `logs`, and `doctor`.
|
||||
|
||||
#### Scenario: Operator investigates a completed assignment
|
||||
- **WHEN** the operator requests history for a terminal reservation
|
||||
- **THEN** the worker SHALL show its terminal local/receipt outcome, durations, phase timeline, and available diagnostic artifact references
|
||||
|
||||
#### Scenario: Operator runs doctor
|
||||
- **WHEN** `truf-worker doctor` is invoked
|
||||
- **THEN** it SHALL inspect package identity, singleton/process state, local state readability, disk usage, server reachability, and retained work without claiming an assignment
|
||||
|
||||
### Requirement: From-zero operator documentation
|
||||
The release SHALL include one canonical guide from package acquisition through installation, first start, attach/status interpretation, graceful stop, recovery, update, diagnostics, and removal.
|
||||
|
||||
#### Scenario: New operator follows the guide
|
||||
- **WHEN** an operator starts with a supported worker package and issued server enrollment data
|
||||
- **THEN** the documented commands SHALL lead to a running verified worker and explain every state visible before the first assignment
|
||||
+79
@@ -0,0 +1,79 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Canonical assignment phase model
|
||||
The worker SHALL represent assignment execution with one versioned phase/event model shared by local status, local history, server progress, and administrative views.
|
||||
|
||||
#### Scenario: Assignment completes normally
|
||||
- **WHEN** a slot claims, executes, stages, uploads, and receives acceptance for an assignment
|
||||
- **THEN** it SHALL emit monotonic phase events sufficient to reconstruct the time spent from `assigned` through `awaiting_receipt`
|
||||
|
||||
#### Scenario: Process restarts during an assignment
|
||||
- **WHEN** a worker restarts with persisted slot state
|
||||
- **THEN** recovered events SHALL continue from the persisted sequence and SHALL record recovery without rewriting the prior timeline
|
||||
|
||||
### Requirement: Complete scan-stage watchdog
|
||||
The worker SHALL enforce one hard scan-stage deadline across permit acquisition, local preparation/resolution, data acquisition, scanner execution, filtering, cleanup, and result staging by supervising the complete execution unit outside the controller process.
|
||||
|
||||
#### Scenario: Scanner child exceeds the deadline
|
||||
- **WHEN** the assignment runner remains active at the scan-stage deadline
|
||||
- **THEN** the controller SHALL terminate its complete process tree, release/detach local resources, and produce a timeout result identifying the final phase
|
||||
|
||||
#### Scenario: Cleanup blocks after scanner exit
|
||||
- **WHEN** scanner execution has ended but cleanup or staging remains blocked at the deadline
|
||||
- **THEN** the same hard deadline SHALL terminate the runner and SHALL prevent the slot from remaining occupied until assignment expiry
|
||||
|
||||
#### Scenario: Permit acquisition consumes the budget
|
||||
- **WHEN** no scan permit is acquired before the scan-stage deadline
|
||||
- **THEN** the worker SHALL produce a phase-specific timeout result without starting the scanner
|
||||
|
||||
### Requirement: Non-renewing server progress
|
||||
The Worker API SHALL accept idempotent monotonic progress events for the current reservation while preserving the original immutable assignment and queue deadlines.
|
||||
|
||||
#### Scenario: Progress is accepted
|
||||
- **WHEN** the assigned device submits the next valid event sequence for its unresolved reservation
|
||||
- **THEN** the server SHALL persist the event/latest phase and SHALL NOT alter assignment expiry or ownership
|
||||
|
||||
#### Scenario: Duplicate progress is retried
|
||||
- **WHEN** an already accepted event sequence is submitted again
|
||||
- **THEN** the server SHALL return the prior acceptance without creating a duplicate timeline event
|
||||
|
||||
#### Scenario: Progress cannot reach the server
|
||||
- **WHEN** local phase transitions occur during a temporary connection failure
|
||||
- **THEN** execution SHALL continue under the fixed deadline and events SHALL remain available locally for ordered retry
|
||||
|
||||
### Requirement: Observable deadline semantics
|
||||
Assignments SHALL carry distinct effective target-scan, result-upload, and end-to-end assignment deadlines, and every operator/admin view SHALL label them by those meanings.
|
||||
|
||||
#### Scenario: Operator inspects active work
|
||||
- **WHEN** status or admin renders an active reservation
|
||||
- **THEN** it SHALL show the effective scan deadline, assignment deadline, time remaining, and current phase without conflating them
|
||||
|
||||
#### Scenario: Assignment expires
|
||||
- **WHEN** the immutable assignment deadline passes without an accepted terminal result
|
||||
- **THEN** expiry evidence SHALL include the last accepted phase and last-progress timestamp when available
|
||||
|
||||
### Requirement: Global and per-source assignment policy
|
||||
The managed runtime configuration SHALL provide a global assignment TTL fallback and optional explicit overrides for GitLab, DockerHub, and HuggingFace, selected by the server at issuance.
|
||||
|
||||
#### Scenario: Source override exists
|
||||
- **WHEN** a DockerHub assignment is issued and a DockerHub assignment TTL override is configured
|
||||
- **THEN** its immutable expiry SHALL use the override and the assignment SHALL report that effective policy
|
||||
|
||||
#### Scenario: Source override is absent
|
||||
- **WHEN** an assignment is issued for a source without an override
|
||||
- **THEN** the global assignment TTL SHALL be used
|
||||
|
||||
#### Scenario: Invalid deadline policy is previewed
|
||||
- **WHEN** an effective assignment deadline cannot cover its source scan timeout, upload deadline, and required handoff margin
|
||||
- **THEN** managed configuration preview SHALL reject the candidate with a field-specific explanation
|
||||
|
||||
### Requirement: Phase duration percentiles
|
||||
The server SHALL expose p50, p95, and p99 durations by source, phase, outcome, and selected time window, based only on completed observations appropriate to that metric.
|
||||
|
||||
#### Scenario: Administrator reviews DockerHub latency
|
||||
- **WHEN** duration metrics are requested for DockerHub
|
||||
- **THEN** the result SHALL separate end-to-end, scanning, cleanup, bundling, and upload percentiles and SHALL report sample counts
|
||||
|
||||
#### Scenario: Insufficient samples exist
|
||||
- **WHEN** a percentile does not have the configured minimum sample count
|
||||
- **THEN** the UI/API SHALL label it insufficient rather than presenting it as a stable policy recommendation
|
||||
@@ -0,0 +1,44 @@
|
||||
## 1. Canonical Contracts and Persistence
|
||||
|
||||
- [x] 1.1 Implement the versioned worker phase/event model, canonical phase transitions, monotonic sequence validation, JSON/NDJSON serialization, and contract tests shared by worker, API, and admin code.
|
||||
- [x] 1.2 Implement the unified diagnostic envelope, orthogonal taxonomy, deterministic diagnostic UID, exact body/log representation, explicit truncation metadata, aggregate limits, and serialization/validation tests.
|
||||
- [x] 1.3 Add PostgreSQL progress-event and worker-diagnostic persistence, indexes, idempotent writes, reservation/scan joins, migration coverage, and authoritative query methods.
|
||||
- [x] 1.4 Add managed global/per-source assignment deadline policy, effective-value validation against scan/upload bounds, assignment serialization, and runtime-document/editor tests.
|
||||
|
||||
## 2. Complete Worker Supervisor
|
||||
|
||||
- [x] 2.1 Build the `truf-worker` command surface (`install`, `run`, `start`, `stop`, `status`, `attach`, `logs`, `history`, `doctor`) over one supervisor implementation, including the existing foreground invocation migration alias.
|
||||
- [x] 2.2 Implement verified Windows and Linux/WSL supervisor instance lifecycle, startup handshake, local control endpoint, graceful drain/stop, shutdown receipt, stale-instance handling, and lifecycle tests.
|
||||
- [x] 2.3 Implement append-only local events, rebuildable status projection, terminal history, rotating logs, per-assignment diagnostic/body/log files, retention accounting, and crash/restart recovery tests.
|
||||
- [x] 2.4 Implement human status/attach views and versioned JSON/NDJSON modes with coherent concurrent-slot rendering, honest phase/deadline/progress fields, bounded follow/tail behavior, and command-level tests.
|
||||
- [x] 2.5 Update Windows portable and Linux/Docker package entrypoints, manifests, launchers, and package self-tests so the supervisor is the supported runtime on every platform.
|
||||
|
||||
## 3. Assignment Runner and Full-Stage Watchdog
|
||||
|
||||
- [x] 3.1 Introduce the contained per-assignment runner process and controller protocol while preserving existing claim state, deterministic bundle authority, source capabilities, and receipt/recovery behavior.
|
||||
- [x] 3.2 Instrument permit wait, preparation, source resolution, download/clone, scanner execution, filtering, cleanup, bundle staging, upload, and receipt transitions with the canonical phase events and measured durations.
|
||||
- [x] 3.3 Enforce one hard scan-stage deadline across the runner process tree, produce a normal phase-specific timeout result, detach abandoned work to janitor ownership, and return the slot without waiting for assignment expiry.
|
||||
- [x] 3.4 Implement controller and runner crash recovery for persisted assignments, incomplete runner outputs, ready bundles, stale results, lowered parallelism, and supervisor restart.
|
||||
- [x] 3.5 Add deterministic fault-injection tests for blocking/failure in every phase, including permit starvation, child non-exit, cleanup stall, staging/fsync failure, upload retry, deadline crossing, and process restart.
|
||||
|
||||
## 4. Worker API and Result Pipeline
|
||||
|
||||
- [x] 4.1 Add the authenticated non-renewing progress endpoint with ownership checks, monotonic/idempotent sequencing, latest-phase projection, bounded retry behavior, and API/database tests.
|
||||
- [x] 4.2 Add diagnostic frames to protocol-2 bundles and the same diagnostic envelope to prebundle terminal reports, including exact size accounting, deterministic replay, and protocol compatibility tests.
|
||||
- [x] 4.3 Ingest diagnostics transactionally with target scans/errors and attach prebundle diagnostics to reservations, while preserving durable receipt, queue settlement, projection, and replay invariants.
|
||||
- [x] 4.4 Extend assignment/status responses with effective deadlines, latest phase/progress, known idle/backoff reason, and diagnostic availability, and cover old-package records explicitly in compatibility tests.
|
||||
|
||||
## 5. Final Administration Experience
|
||||
|
||||
- [x] 5.1 Replace the worker-list query/view model with separate assignment outcome, scan outcome, diagnostic summary, active phase, phase/progress age, effective deadlines, slot/cap, and package fields.
|
||||
- [x] 5.2 Build the assignment detail page with ordered phase/receipt/ingestion/settlement/projection timeline, duration breakdown, scan summary, diagnostic list, exact body/log views, canonical JSON copy/download, and explicit transformation metadata.
|
||||
- [x] 5.3 Add independent filters and repeated-diagnostic grouping for source, worker/device, assignment outcome, scan outcome, phase, category, stable code, retryability, and time window without hiding individual occurrences.
|
||||
- [x] 5.4 Add source/phase/outcome p50, p95, and p99 duration queries and admin views with sample counts, and present them beside effective scan/upload/assignment deadline policy in the runtime editor.
|
||||
- [x] 5.5 Add end-to-end admin tests for active progress, accepted scan errors, prebundle failures, assignment expiry, legacy records, complete/truncated bodies, repeated fingerprints, and machine-readable detail output.
|
||||
|
||||
## 6. Operator Release and Production Proof
|
||||
|
||||
- [x] 6.1 Write and validate the canonical from-zero operator guide covering package acquisition, install, first run, start/status/attach/logs/history, phase/deadline interpretation, diagnostics, graceful stop/drain, recovery, update, and removal.
|
||||
- [x] 6.2 Run the complete unit/integration/protocol/package test matrix and build reproducible Windows and Linux worker artifacts with registered manifests and documented identities.
|
||||
- [x] 6.3 Perform bounded production validation on native Windows and WSL/Docker covering multi-slot progress, attach while active, a forced full-stage timeout, diagnostic body/log inspection, restart recovery, accepted/ingested/projected reconciliation, and final production restoration.
|
||||
- [x] 6.4 Record final duration percentiles, watchdog evidence, diagnostic/admin screenshots or snapshots, operator command transcript, known limits, and rollout/rollback results in a durable dated report.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-06
|
||||
@@ -0,0 +1,80 @@
|
||||
## Context
|
||||
|
||||
The completed keyword-pruning change removed 38 globally zero-yield terms from discovery configuration, but `target_queue` admission still considers previously queued `pending` and due `deferred` rows without consulting current query policy. Since deployment, Docker work attributed to retired terms consumed about 20.6 scanner-hours and produced no strict-usable credential. GitHub Actions independently consumed one core worker while both fresh and retained work produced no strict-usable credential and its active backlog grew.
|
||||
|
||||
Queue rows are durable authority referenced by scan history, result reservations, Git and Docker coverage, deduplication, and retry state. Physical deletion or overloading quarantine would destroy or misrepresent that authority. Runtime configuration and PostgreSQL schema are lifecycle-protected and require coordinated offline changes.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Make policy-retired pending and deferred work explicitly unclaimable while preserving every durable record and field needed for audit or reversal.
|
||||
- Apply and reverse policy holds only through exact, reviewed, idempotent manifests under stopped-source lifecycle authority.
|
||||
- Prevent stale Docker resolver anchors from generating new digest children after the initial cold transition.
|
||||
- Pause GitHub Actions without deleting or rewriting its backlog.
|
||||
- Expose cold rows separately from active backlog in operational counts.
|
||||
|
||||
**Non-Goals:**
|
||||
- Deleting target queue rows or changing scan, finding, credential, result, deduplication, or coverage history.
|
||||
- Age-based retirement of mutable GitLab or HuggingFace targets.
|
||||
- Automatically colding every disabled source or all dormant historical backlogs in the first deployment.
|
||||
- Changing scanner concurrency, Docker scan policy, discovery keywords, or keycheck behavior.
|
||||
- Treating query retirement as evidence that a target can never become useful.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Add a first-class `cold` queue status
|
||||
|
||||
`target_queue.status='cold'` is the policy hold. Existing scan, legacy, and Docker resolver claim paths explicitly allow only `pending` and due `deferred`, so cold rows remain fail-closed even if a caller does not understand policy metadata. Cold is not a worker result disposition and ordinary scans cannot produce it.
|
||||
|
||||
Alternative: add a nullable metadata flag. Rejected because every current and older admission path would need to remember an additional predicate, making accidental claims likely. `deferred` is also unsuitable because it is a timed retry, and `quarantined` is reserved for pipeline-integrity failures with capacity accounting and review semantics.
|
||||
|
||||
### Use exact reviewed manifests and append-only audit events
|
||||
|
||||
Add an append-only `target_queue_policy_events` table recording each cold/reactivate transition, prior and next status, exact source/platform/query attribution, configuration and policy hashes, manifest hash, prior update fence, reason code, and reversal linkage. A private dry-run manifest lists only queue IDs and non-sensitive policy attribution; it never contains target values.
|
||||
|
||||
Apply locks rows in deterministic ID order and atomically validates the manifest fence, inserts one audit event, and changes only `status` and `updated_at`. Eligible cold transitions are limited to unfenced `pending` or `deferred` rows with no active queue/result/resolver lease or reservation. Reactivation restores the exact audited prior status and is explicit rather than an enqueue side effect.
|
||||
|
||||
Alternative: issue a one-off SQL update and infer reversal from `available_after`. Rejected because it is not reviewable, cannot prove the selected cohort, and loses the exact prior lifecycle state.
|
||||
|
||||
### Derive stale policy from canonical exact queries
|
||||
|
||||
Policy uses case-sensitive exact `(source, platform, query)` triples from canonical source configuration. Null/blank queries, unknown source/platform pairs, non-rotation provenance, and sole operational sentinel queries are not inferred as stale. Disabled sources retain their configured query policy; source disablement is not itself a retirement action.
|
||||
|
||||
The initial reviewed transition is scoped to DockerHub platform `docker`, covering every eligible pending/deferred resolver anchor and digest row whose own exact query is no longer configured. Dormant source families remain preserved and unclaimable by virtue of having no worker; they can be reviewed separately later rather than mutating a six-figure backlog in this deployment.
|
||||
|
||||
### Filter periodic Docker resolver claims by current query policy
|
||||
|
||||
Docker digest children inherit the resolver anchor query, and completed anchors can be periodically reclaimed. The runtime therefore passes the exact configured DockerHub query allowlist to resolver admission and excludes non-null anchors whose query is not allowed. This prevents completed stale anchors, which are outside the pending/deferred cold migration, from recreating policy-retired work.
|
||||
|
||||
The filter is exact and does not rewrite query attribution. Null legacy provenance remains excluded from automatic policy retirement and requires separate review.
|
||||
|
||||
### Preserve cold state until explicit reactivation
|
||||
|
||||
Ordinary queue synchronization, enqueue/upsert, rediscovery, retry, and completion logic must not turn a cold row back into pending/deferred. If a target becomes relevant through a retained query, an operator can review its audit lineage and explicitly reactivate it; implicit reactivation would make the policy hold non-deterministic and unaudited.
|
||||
|
||||
### Pause GitHub Actions at all configuration layers
|
||||
|
||||
Remove `github_actions` from `supervisor.enabled_sources` and set both supervisor-source and source-level enablement false. Keep its query and all queue/history rows unchanged. No cold transition is required for GHA because no worker remains able to claim its source/platform rows.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Earliest query attribution can cold a target rediscovered by a retained term] -> Preserve full audit lineage and require explicit reactivation; do not delete the row or overwrite its original query.
|
||||
- [A live or fenced row could be transitioned] -> Require coordinated source shutdown, exact row-update fences, reservation/lease checks, deterministic locks, and atomic all-or-nothing apply.
|
||||
- [Completed Docker anchors could bypass the migration] -> Filter retry and periodic resolver admission by the current exact query allowlist.
|
||||
- [Cold rows could inflate active backlog displays] -> Count `cold` separately and exclude it from pending/deferred operational backlog totals.
|
||||
- [A policy/config change between review and apply could invalidate the cohort] -> Fence the manifest with canonical configuration, query-policy, selection, and manifest hashes.
|
||||
- [Pausing GHA may miss future useful credentials] -> Preserve its entire queue and configuration for a coordinated future re-enable after explicit review.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add schema support, policy transition APIs, exact manifest planning/apply commands, Docker resolver query filtering, cold-preserving upsert behavior, and focused tests.
|
||||
2. Stop the authenticated runtime coordinately and verify sources are stopped and no active target/result/resolver fences block the selected Docker cohort.
|
||||
3. Start maintenance PostgreSQL, apply the additive schema migration, and generate a bounded private DockerHub stale-query cold manifest.
|
||||
4. Review aggregate counts and hashes, then apply the exact manifest atomically. Retain only the append-only database audit; remove the temporary private manifest after verification.
|
||||
5. Deploy the GHA-disabled configuration and restart through the canonical runtime lifecycle.
|
||||
6. Verify no cold Docker row is claimable, no stale completed anchor is resolver-claimable, cold and active counts reconcile, GHA has no child process, and the remaining pipeline is healthy.
|
||||
7. Roll back by stopping sources, generating an exact reactivation manifest from unreversed cold events, applying it atomically, restoring GHA configuration if desired, and restarting canonically.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None.
|
||||
@@ -0,0 +1,23 @@
|
||||
## Why
|
||||
|
||||
Removing zero-yield discovery terms stopped new discovery but left their pending and deferred targets claimable, so retired policy continued consuming scanner time. GitHub Actions also produced no strict-usable credential from either fresh work or its large retained backlog while that backlog kept growing, so it should no longer occupy a core worker.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add an explicit, auditable, reversible `cold` lifecycle state for policy-retired target queue rows without deleting queue history, scans, deduplication, reservations, or coverage records.
|
||||
- Cold only unfenced `pending` and `deferred` rows whose exact source query is absent from the canonical configured query policy, using a reviewed offline manifest and coordinated lifecycle authority.
|
||||
- Prevent completed Docker resolver anchors attributed to retired queries from being periodically reclaimed and creating new stale digest children.
|
||||
- Preserve cold rows across ordinary enqueue, rediscovery, retry, and completion paths; require an explicit audited action to reactivate them.
|
||||
- Pause GitHub Actions by removing it from the active core and disabling both supervisor and source configuration, while preserving its complete queue and history.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `target-queue-policy-holds`: Auditable cold and reactivation transitions for policy-stale target queue work, including claim exclusion and Docker resolver filtering.
|
||||
|
||||
### Modified Capabilities
|
||||
- `discovery-keyword-pruning`: Retired queries stop both future discovery and claimable pending/deferred work while preserving every historical authority record.
|
||||
|
||||
## Impact
|
||||
|
||||
The change affects PostgreSQL target queue schema and migration authority in `app/scanner_db.py` and `app/migrate_runtime_safety.py`, Docker resolver admission and query plumbing in `app/console_runner.py`, core source selection in `app/config.yaml`, focused lifecycle/configuration tests, and dashboard/status aggregation where queue states are enumerated. Deployment requires a coordinated runtime stop, schema migration, reviewed cold manifest application, and canonical restart. No target, scan, finding, credential, result, deduplication, or coverage row is deleted.
|
||||
@@ -0,0 +1,20 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Operational and historical authority is preserved
|
||||
Keyword retirement SHALL stop future discovery and SHALL permit existing unfenced pending/deferred targets attributed to retired exact queries to enter an audited, reversible cold state without deleting or rewriting historical authority.
|
||||
|
||||
#### Scenario: Dedicated source sentinels remain
|
||||
- **WHEN** archive and gist source rotations are loaded
|
||||
- **THEN** `gharchive`, `gharchive-files`, and `gists` SHALL remain as their sole configured query tokens
|
||||
|
||||
#### Scenario: Persisted rotation index remains valid
|
||||
- **WHEN** an existing query index exceeds a shortened query list
|
||||
- **THEN** normal modulo-based rotation SHALL select a valid configured query without a state-file edit
|
||||
|
||||
#### Scenario: Existing backlog is preserved but held
|
||||
- **WHEN** a previously admitted unfenced target is attributed to a retired exact query and selected by reviewed policy
|
||||
- **THEN** its queue row SHALL remain present with all attribution, retry, deduplication, scan, reservation, and coverage history preserved while its status becomes unclaimable `cold`
|
||||
|
||||
#### Scenario: Historical records remain unchanged
|
||||
- **WHEN** a stale-query cold transition is applied
|
||||
- **THEN** existing target scans, findings, credentials, keycheck results, completed queue rows, and coverage records SHALL NOT be deleted or rewritten
|
||||
@@ -0,0 +1,78 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Cold queue rows are not claimable
|
||||
The system SHALL represent policy-held target work with `target_queue.status='cold'`, and no scanner, legacy queue consumer, or Docker resolver SHALL claim a cold row.
|
||||
|
||||
#### Scenario: Scanner checks active backlog
|
||||
- **WHEN** a queue row has status `cold`
|
||||
- **THEN** the row SHALL be excluded from pending, due-deferred, retry, and periodic resolver admission
|
||||
|
||||
#### Scenario: Worker completes ordinary work
|
||||
- **WHEN** a scan result is ingested
|
||||
- **THEN** its queue disposition SHALL remain limited to ordinary lifecycle outcomes and SHALL NOT create a cold status
|
||||
|
||||
### Requirement: Policy holds are exact and audited
|
||||
The system SHALL transition queue rows into or out of cold state only through a reviewed, hash-fenced, append-only policy event under stopped-source lifecycle authority.
|
||||
|
||||
#### Scenario: Eligible stale row is held
|
||||
- **WHEN** an exact manifest entry still matches an unfenced `pending` or `deferred` row whose source query is absent from canonical policy
|
||||
- **THEN** the system SHALL atomically record the audit event and change only the row status and update timestamp to `cold`
|
||||
|
||||
#### Scenario: Selected row has an active fence
|
||||
- **WHEN** a selected row has a queue lease, resolver lease, current claim, active result reservation, or submitted content lease
|
||||
- **THEN** the policy apply SHALL fail closed without partially applying the manifest
|
||||
|
||||
#### Scenario: Exact apply is repeated
|
||||
- **WHEN** the same reviewed manifest is applied again after a successful transition
|
||||
- **THEN** the system SHALL report an idempotent duplicate without creating a second state transition
|
||||
|
||||
### Requirement: Cold state is explicitly reversible
|
||||
The system SHALL retain the exact prior queue status in policy audit history and SHALL require a reviewed reactivation manifest to restore a cold row.
|
||||
|
||||
#### Scenario: Cold row is rediscovered normally
|
||||
- **WHEN** enqueue, synchronization, rediscovery, retry, or completion logic encounters an existing cold row
|
||||
- **THEN** the row SHALL remain cold and its historical attribution SHALL remain unchanged
|
||||
|
||||
#### Scenario: Reviewed cold event is reactivated
|
||||
- **WHEN** an exact reactivation manifest references an unreversed cold event and all row fences still match
|
||||
- **THEN** the system SHALL restore the audited prior `pending` or `deferred` status and append a linked reactivation event
|
||||
|
||||
### Requirement: Stale-query selection follows canonical policy
|
||||
The system SHALL evaluate query staleness using case-sensitive exact source, platform, and query policy derived from canonical configuration.
|
||||
|
||||
#### Scenario: Configured query remains active
|
||||
- **WHEN** a queue row's exact source/platform/query triple remains configured
|
||||
- **THEN** automatic stale-policy planning SHALL NOT select the row
|
||||
|
||||
#### Scenario: Attribution cannot be classified safely
|
||||
- **WHEN** query attribution is null, blank, operational, non-rotation, or belongs to an unknown source/platform pair
|
||||
- **THEN** automatic planning SHALL skip and report the row rather than inferring retirement
|
||||
|
||||
#### Scenario: Initial Docker stale cohort is planned
|
||||
- **WHEN** DockerHub policy planning is scoped to platform `docker`
|
||||
- **THEN** it SHALL include eligible pending/deferred resolver anchors and digest rows attributed to removed exact queries and SHALL expose only aggregate counts plus non-sensitive manifest fields
|
||||
|
||||
### Requirement: Docker resolvers honor current query policy
|
||||
The system SHALL prevent completed or retryable Docker resolver anchors attributed to retired non-null queries from generating new digest queue rows.
|
||||
|
||||
#### Scenario: Periodic stale anchor becomes due
|
||||
- **WHEN** a completed Docker resolver anchor is due but its exact query is not in the configured DockerHub allowlist
|
||||
- **THEN** resolver admission SHALL leave the anchor unclaimed
|
||||
|
||||
#### Scenario: Configured anchor becomes due
|
||||
- **WHEN** a Docker resolver anchor's exact query remains configured and all ordinary claim fences pass
|
||||
- **THEN** resolver admission SHALL preserve the existing retry and periodic behavior
|
||||
|
||||
### Requirement: Source pause preserves backlog authority
|
||||
Pausing a source SHALL remove its worker from the active core without deleting or rewriting that source's queue or historical records.
|
||||
|
||||
#### Scenario: GitHub Actions is paused
|
||||
- **WHEN** canonical runtime configuration is loaded after this change
|
||||
- **THEN** GitHub Actions SHALL be absent from the supervisor core and disabled at both supervisor-source and source configuration layers while its configured query and persisted backlog remain intact
|
||||
|
||||
### Requirement: Cold work is separately observable
|
||||
Operational queue summaries SHALL report cold rows separately and SHALL NOT include them in active pending or deferred backlog counts.
|
||||
|
||||
#### Scenario: Queue state is summarized
|
||||
- **WHEN** an operator inspects canonical target queue status
|
||||
- **THEN** the summary SHALL expose a distinct cold count without representing those rows as retryable or claimable work
|
||||
@@ -0,0 +1,21 @@
|
||||
## 1. Durable Policy State
|
||||
|
||||
- [x] 1.1 Add the `cold` target queue lifecycle state, append-only policy-event schema, indexes, additive migration marker, and schema validation coverage.
|
||||
- [x] 1.2 Implement atomic, fenced, idempotent cold and reactivation database transitions that preserve all non-lifecycle queue authority.
|
||||
|
||||
## 2. Reviewed Policy Operations
|
||||
|
||||
- [x] 2.1 Implement exact canonical query-policy derivation and privacy-safe stale-row/reversal manifest planning.
|
||||
- [x] 2.2 Add stopped-source migration CLI dry-run/apply paths with manifest/config/policy/selection hash validation and bounded deterministic scope.
|
||||
|
||||
## 3. Runtime Enforcement
|
||||
|
||||
- [x] 3.1 Preserve cold rows across ordinary queue enqueue/synchronization and expose cold separately in canonical queue summaries.
|
||||
- [x] 3.2 Pass configured DockerHub query policy into retry/periodic resolver admission and exclude retired-query anchors.
|
||||
- [x] 3.3 Remove GitHub Actions from the active core and disable both supervisor and source layers without altering its query or persisted backlog.
|
||||
|
||||
## 4. Verification And Deployment
|
||||
|
||||
- [x] 4.1 Add focused unit and PostgreSQL integration tests for claim exclusion, exact selection, fenced transitions, idempotency, reversal, Docker resolver filtering, cold preservation, observability, and GHA pause.
|
||||
- [x] 4.2 Run targeted test suites and strict OpenSpec validation with no forbidden application bytecode artifacts.
|
||||
- [x] 4.3 Coordinately stop runtime, apply the additive schema migration and reviewed Docker stale-query cold manifest, remove temporary artifacts, restart canonically, and verify active backlog and pipeline health.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-08-26
|
||||
@@ -0,0 +1,78 @@
|
||||
## Context
|
||||
|
||||
The exact `openai` rollout proved that bounded source queries can restore unseen credential supply without changing scanner or keycheck authority: 18 DockerHub targets produced 24 genuinely new OpenAI credentials, all with explicit terminal API outcomes. None were usable, so increasing that broad query is not justified. Historical production attribution instead shows usable OpenAI outcomes behind narrower agent, chatbot, and conversation ecosystems, while source semantics differ enough that one shared keyword list is inefficient.
|
||||
|
||||
The existing exact-query override mechanism already validates and applies `pages`, `per_page`, and `max_targets`. This change can therefore remain configuration-only plus focused contract tests. Runtime configuration is immutable-authority covered, so deployment must use coordinated stop/start and must not modify persisted query state directly.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Test nine high-signal OpenAI integration and deployable-application queries across the sources where their search semantics fit.
|
||||
- Bound discovery admission as well as scan claims for every new query.
|
||||
- Run each new query once promptly after deployment without directly editing query state.
|
||||
- Attribute the canary through durable scans, candidates, provider results, and projections.
|
||||
- Keep or remove each query based on its own measured useful yield and operational cost.
|
||||
|
||||
**Non-Goals:**
|
||||
- Expanding broad generic terms such as `gpt`, `llm`, `chatgpt`, `ai`, or `model`.
|
||||
- Increasing global scan concurrency, source worker counts, or updated-target promotion limits.
|
||||
- Changing detector routing, OpenAI checker classification, known-credential caching, or projection semantics.
|
||||
- Re-enabling inactive broad sources or guaranteeing a valid funded credential.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Use source-specific first-wave queries
|
||||
|
||||
The first wave is:
|
||||
- GitHub: `OPENAI_API_KEY`, `api.openai.com`, `openai-agents`.
|
||||
- GitLab: `openai-api`, `openai-agents`, `librechat`.
|
||||
- DockerHub: `librechat`, `lobechat`, `openai-proxy`.
|
||||
|
||||
GitHub can search README content, so direct environment and endpoint signatures are appropriate. GitLab project search is metadata-oriented, so branded slug terms are used. DockerHub searches repository metadata and then resolves immutable image digests, so deployable project and proxy names are used.
|
||||
|
||||
Alternative: add the same list to all sources. Rejected because it lengthens every rotation and ignores source search semantics. Alternative: expand historically broad terms. Rejected because those cohorts produced volume without usable OpenAI outcomes.
|
||||
|
||||
### Bound every query independently
|
||||
|
||||
Bounds are:
|
||||
- GitHub `OPENAI_API_KEY`: `pages=1`, `per_page=25`, `max_targets=5`.
|
||||
- GitHub `api.openai.com`: `1/25/5`.
|
||||
- GitHub `openai-agents`: `1/50/5`.
|
||||
- GitLab `openai-api`: `1/50/5`.
|
||||
- GitLab `openai-agents`: `1/50/5`.
|
||||
- GitLab `librechat`: `1/25/5`.
|
||||
- DockerHub `librechat`, `lobechat`, and `openai-proxy`: each `2/10/10`.
|
||||
|
||||
Page and page-size limits bound fetched/admitted identities; `max_targets` separately bounds claims in the active cycle. The existing global three-slot limit, GitLab one-updated-target-per-cycle cap, 24-hour update cooldown, and Docker digest requirement remain unchanged.
|
||||
|
||||
### Place the wave at each stopped source's current rotation index
|
||||
|
||||
After a coordinated runtime stop, capture each source's persisted `query_index` and insert that source's three-query block at the same index. Restarting then exercises the block naturally. Successful discovery advances through the block; failures and backlog-only work retain the current query under existing semantics. State files are never edited.
|
||||
|
||||
Alternative: append and wait for a full rotation. Rejected because DockerHub cycles can be long and attribution would be delayed. Alternative: edit persisted state. Rejected because state is runtime authority and direct edits would weaken recovery evidence.
|
||||
|
||||
### Evaluate individual query funnels
|
||||
|
||||
The canary reports fetched, new/updated admissions, durable scan outcomes, findings, genuinely new OpenAI credentials, `api_check` versus `cached_status`, explicit provider outcomes, first-alive/usable counts, projection drain, and runtime health. Aggregate volume alone cannot justify retention.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [README signatures produce placeholders] -> Count genuinely new credentials and explicit API outcomes; remove queries with only placeholder/dead yield.
|
||||
- [A query admits more work than its claim cap] -> Keep `pages × per_page` small and inspect the exact query-attributed queue until terminal.
|
||||
- [Docker scans are expensive or inaccessible] -> Cap each query at 20 repositories and 10 claims, retain digest authority, and classify registry failures separately.
|
||||
- [Three added terms lengthen source rotations] -> Keep only terms that add distinct credential or usable yield after the canary.
|
||||
- [Config deployment triggers authority fail-close] -> Stop coordinately before editing and restart only through `start_runtime.ps1`.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add focused tests for exact source membership, uniqueness, and all nine bounds.
|
||||
2. Coordinately stop the runtime and capture persisted source indices.
|
||||
3. Insert each three-query block at its source's current index; do not edit state files.
|
||||
4. Run focused and full regression suites plus strict OpenSpec validation with bytecode writes disabled.
|
||||
5. Start through the authoritative runtime script and verify PostgreSQL, pipeline, core sources, keychecks, and recorder.
|
||||
6. Observe each exact query cohort through terminal scans and keycheck projection, then retain or remove each query based on measured evidence.
|
||||
7. Roll back any low-value query by removing that query and override during a coordinated stop/start. No schema or data rollback is required.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None. A second wave remains gated on this canary's per-query useful-yield evidence.
|
||||
@@ -0,0 +1,24 @@
|
||||
## Why
|
||||
|
||||
The bounded exact `openai` canary restored fresh credential supply but produced no usable credentials, while historical production evidence shows that narrower ecosystem and integration terms can reach different cohorts. A small source-specific keyword wave can test those higher-signal surfaces without expanding broad generic discovery or scan concurrency.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add three source-specific OpenAI ecosystem queries to each of GitHub, GitLab, and DockerHub.
|
||||
- Apply exact per-query page, page-size, and claim bounds so every new query is independently constrained.
|
||||
- Preserve existing query rotation, target deduplication, revision-aware rescan limits, Docker digest authority, scan concurrency, and keycheck behavior.
|
||||
- Run one controlled production canary per new query and measure discovery, scans, new OpenAI credentials, explicit API outcomes, and usable yield.
|
||||
- Retain, revise, or remove individual queries based on measured bounded evidence rather than fetched volume.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `openai-ecosystem-discovery`: Source-specific bounded discovery for OpenAI integration signatures and deployable ecosystem projects, with per-query canary evidence.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
The change affects `app/config.yaml`, focused query-configuration tests, authority-managed source rotation, and production canary operations. Existing allowlisted query override code is reused unchanged. There is no schema migration, new dependency, detector change, global concurrency increase, or credential recheck policy change.
|
||||
+73
@@ -0,0 +1,73 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Source-specific OpenAI ecosystem queries
|
||||
The system SHALL include the approved first-wave OpenAI ecosystem queries only in the source rotations whose search semantics match those queries.
|
||||
|
||||
#### Scenario: GitHub searches integration signatures
|
||||
- **WHEN** GitHub reaches the first-wave positions in its normal rotation
|
||||
- **THEN** it SHALL search `OPENAI_API_KEY`, `api.openai.com`, and `openai-agents`
|
||||
|
||||
#### Scenario: GitLab searches branded project metadata
|
||||
- **WHEN** GitLab reaches the first-wave positions in its normal rotation
|
||||
- **THEN** it SHALL search `openai-api`, `openai-agents`, and `librechat`
|
||||
|
||||
#### Scenario: DockerHub searches deployable ecosystems
|
||||
- **WHEN** DockerHub reaches the first-wave positions in its normal rotation
|
||||
- **THEN** it SHALL search `librechat`, `lobechat`, and `openai-proxy`
|
||||
|
||||
#### Scenario: Broad generic expansion is excluded
|
||||
- **WHEN** the first-wave configuration is evaluated
|
||||
- **THEN** it SHALL NOT add new broad variants of `gpt`, `llm`, `chatgpt`, `ai`, or `model`
|
||||
|
||||
### Requirement: Independent bounded query policies
|
||||
Each first-wave query SHALL have an exact allowlisted override that bounds both discovery volume and scan claims without changing source defaults or other query policies.
|
||||
|
||||
#### Scenario: GitHub signature queries are bounded
|
||||
- **WHEN** GitHub builds arguments for `OPENAI_API_KEY` or `api.openai.com`
|
||||
- **THEN** it SHALL use one page of 25 results and claim at most five targets
|
||||
|
||||
#### Scenario: GitHub agent query is bounded
|
||||
- **WHEN** GitHub builds arguments for `openai-agents`
|
||||
- **THEN** it SHALL use one page of 50 results and claim at most five targets
|
||||
|
||||
#### Scenario: GitLab queries are bounded
|
||||
- **WHEN** GitLab builds arguments for a first-wave query
|
||||
- **THEN** it SHALL use one page, the configured 25- or 50-result page size, and claim at most five targets
|
||||
|
||||
#### Scenario: DockerHub queries are bounded
|
||||
- **WHEN** DockerHub builds arguments for a first-wave query
|
||||
- **THEN** it SHALL use at most two pages of ten repositories and claim at most ten targets
|
||||
|
||||
#### Scenario: Existing authority limits remain unchanged
|
||||
- **WHEN** any first-wave query runs
|
||||
- **THEN** global scan concurrency, revision-aware promotion limits, cooldowns, and Docker digest requirements SHALL remain authoritative
|
||||
|
||||
### Requirement: Natural rotation and failure behavior
|
||||
The first-wave queries SHALL use the existing persisted source rotation without direct query-state modification.
|
||||
|
||||
#### Scenario: Successful query advances
|
||||
- **WHEN** a first-wave discovery cycle completes successfully
|
||||
- **THEN** the source SHALL advance through the existing persisted rotation semantics
|
||||
|
||||
#### Scenario: Failed query is retained
|
||||
- **WHEN** first-wave discovery fails before successful completion
|
||||
- **THEN** the source SHALL retain that query according to existing failure semantics
|
||||
|
||||
#### Scenario: Backlog work does not masquerade as discovery
|
||||
- **WHEN** a source drains existing backlog while a first-wave query is current
|
||||
- **THEN** canary attribution SHALL distinguish backlog-only cycles from the actual discovery cycle
|
||||
|
||||
### Requirement: Per-query end-to-end canary evidence
|
||||
Operators SHALL evaluate each first-wave query through durable source, scan, candidate, provider-result, and projection evidence without exposing targets or credential values.
|
||||
|
||||
#### Scenario: Query cohort reaches terminal accounting
|
||||
- **WHEN** a first-wave query admits new or updated targets
|
||||
- **THEN** operators SHALL verify queue dispositions, scan completion, candidate completion, and projection drain for that exact query cohort
|
||||
|
||||
#### Scenario: Useful yield is measured separately
|
||||
- **WHEN** a first-wave cohort creates OpenAI candidates
|
||||
- **THEN** genuinely new credentials and their explicit API outcomes SHALL be reported separately from cached-known occurrences
|
||||
|
||||
#### Scenario: Retention decision uses measured value
|
||||
- **WHEN** the bounded canary is complete
|
||||
- **THEN** each query SHALL be retained, revised, or removed using its distinct credential yield, usable outcomes, and operational error cost rather than fetched count alone
|
||||
@@ -0,0 +1,17 @@
|
||||
## 1. Source-specific configuration
|
||||
|
||||
- [x] 1.1 Coordinately stop the runtime, capture persisted source query indices, and insert each three-query block at its current index.
|
||||
- [x] 1.2 Add exact allowlisted bounds for all nine queries without changing source-wide defaults or existing safety policies.
|
||||
- [x] 1.3 Add focused configuration tests for source membership, uniqueness, exact bounds, and excluded broad expansion.
|
||||
|
||||
## 2. Verification
|
||||
|
||||
- [x] 2.1 Run focused and full regression suites with bytecode writes disabled.
|
||||
- [x] 2.2 Run strict OpenSpec validation and verify implementation against the artifacts.
|
||||
|
||||
## 3. Production canary
|
||||
|
||||
- [x] 3.1 Start the authority-managed runtime and verify PostgreSQL, pipeline, recorder, keychecks, and all core sources.
|
||||
- [x] 3.2 Observe one actual discovery cycle for each first-wave query and verify configured fetch and claim bounds.
|
||||
- [x] 3.3 Follow every exact query cohort through queue disposition, scan/candidate completion, provider outcomes, and projection drain.
|
||||
- [x] 3.4 Record per-query useful-yield evidence, retain or remove low-value terms, and update the parking lot.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-17
|
||||
@@ -0,0 +1,51 @@
|
||||
## Context
|
||||
|
||||
TruffleHog custom detector entries combine every regex in one detector through match permutation, so multiple regex fields are conjunctive rather than alternative. The Xai and ZaiGLM policies each place context-before and context-after patterns in one detector, making ordinary one-direction matches disappear. Existing tests compile each expression with Python and use `any(...)`, which does not exercise TruffleHog's actual configuration semantics.
|
||||
|
||||
The scanner already canonicalizes `CustomRegex` findings from `ExtraData.name`, and downstream candidate routing and keycheckers depend on the canonical names `Xai` and `ZaiGLM`.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Make both context directions independently executable for Xai and ZAI/GLM.
|
||||
- Preserve canonical persisted detector names and existing keycheck routing.
|
||||
- Exercise the complete policy through the real TruffleHog CLI without network verification.
|
||||
- Keep the policy compatible with the pinned fork on Windows and Linux.
|
||||
|
||||
**Non-Goals:**
|
||||
- Change provider verification endpoints or keycheck behavior.
|
||||
- Add new key formats or broaden the existing regex bounds.
|
||||
- Require TruffleHog to be installed for pure unit-test environments.
|
||||
- Perform live provider requests or use real credentials.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Split alternatives into uniquely named detectors
|
||||
|
||||
Keep the context-before expression under the existing canonical name and move the context-after expression into `XaiContextAfter` or `ZaiGLMContextAfter`. Duplicate names are not used because the custom detector registry can collapse same-name entries; a synthetic CLI probe confirmed unique names execute both alternatives.
|
||||
|
||||
Alternative considered: combine both alternatives into one regex. This was rejected because TruffleHog takes the first capture group as the secret, and Go regex does not support branch-reset groups needed to keep one capture position across both directions.
|
||||
|
||||
### Canonicalize compatibility aliases centrally
|
||||
|
||||
Extend custom detector normalization with a small alias map from the context-after names to `Xai` and `ZaiGLM`. This keeps candidate routing, keychecker gates, dashboards, deduplication, and persisted detector values unchanged.
|
||||
|
||||
Alternative considered: teach every downstream consumer the new names. This would widen the change and create divergent persisted identities for one provider.
|
||||
|
||||
### Add an optional real-CLI regression test
|
||||
|
||||
The test resolves TruffleHog from an explicit environment override, the existing Windows path, or `PATH`. When available, it scans deterministic high-entropy synthetic fixtures with the complete production YAML, `--no-verification`, and `--no-update`, then checks canonical finding names and candidate services. It skips only when no binary is available, while pure tests continue to validate policy structure and alias normalization everywhere.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] A CI environment without TruffleHog can skip the integration gate. -> Keep structural unit coverage and run the real-CLI test in Windows and Linux release jobs.
|
||||
- [Risk] Native Xai can duplicate the custom result for its exact 80-character format. -> Existing candidate identity deduplication remains authoritative; the regression fixture uses a shorter supported overlay form to isolate custom behavior.
|
||||
- [Risk] A future TruffleHog release can change custom detector output fields or flags. -> The real-CLI test asserts `CustomRegex`, `ExtraData.name`, normalization, and candidate routing as one contract.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
Deploy the YAML, scanner alias map, and tests together, rebuild the runtime image, then run the offline compatibility test against both worker binaries before enabling scans. Rollback restores the prior YAML and alias map; no persisted data migration is required.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None.
|
||||
@@ -0,0 +1,23 @@
|
||||
## Why
|
||||
|
||||
The Xai and ZaiGLM custom detector policies model alternative context directions as separate regex entries, but TruffleHog combines entries within one detector as an AND condition. This silently prevents the broader Xai overlay and ordinary ZAI/GLM source detection, while the current unit tests incorrectly model the entries as OR alternatives.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Express each alternative Xai and ZAI/GLM context direction as an independently executable custom detector while preserving the normalized provider names consumed by routing and keychecks.
|
||||
- Add an offline CLI compatibility test that runs the configured TruffleHog binary with the complete custom detector policy and synthetic high-entropy fixtures.
|
||||
- Verify that custom findings normalize and route to the expected Xai and ZAI keycheck services without performing provider verification requests.
|
||||
- Keep the existing native detector, result bundle, candidate, and keycheck contracts unchanged.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `custom-provider-detection-compatibility`: Defines executable compatibility requirements for external custom detector policies and their normalized keycheck routing.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
The change affects `app/trufflehog-custom-detectors.yaml`, scanner finding normalization/routing tests, and the provider detector compatibility test surface. It introduces no production API, schema, dependency, or persisted-data changes.
|
||||
+30
@@ -0,0 +1,30 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Alternative provider contexts execute independently
|
||||
The custom detector policy SHALL detect supported Xai and ZAI/GLM credentials when provider context appears either before or after the credential, without requiring both context directions in one input chunk.
|
||||
|
||||
#### Scenario: Provider context appears before the credential
|
||||
- **WHEN** an offline scan processes a bounded synthetic credential preceded by its supported provider context
|
||||
- **THEN** the policy emits one corresponding custom provider finding
|
||||
|
||||
#### Scenario: Provider context appears after the credential
|
||||
- **WHEN** an offline scan processes a bounded synthetic credential followed by its supported provider context
|
||||
- **THEN** the policy emits one corresponding custom provider finding
|
||||
|
||||
### Requirement: Alternative detector names normalize canonically
|
||||
The scanner MUST normalize all compatibility-only custom detector aliases to the existing canonical `Xai` or `ZaiGLM` detector identity before persistence and candidate extraction.
|
||||
|
||||
#### Scenario: Context-after alias is emitted
|
||||
- **WHEN** TruffleHog emits `CustomRegex` with a context-after compatibility name in `ExtraData.name`
|
||||
- **THEN** the scanner retains `CustomRegex` as the original detector and exposes the canonical provider detector name downstream
|
||||
|
||||
### Requirement: Compatibility is tested through the real CLI
|
||||
The compatibility suite SHALL run the complete configured custom detector policy through an available TruffleHog executable using deterministic synthetic credentials, disabled verification, and disabled update checks.
|
||||
|
||||
#### Scenario: Compatible executable is available
|
||||
- **WHEN** a configured Windows or Linux TruffleHog executable scans the compatibility fixtures
|
||||
- **THEN** both context directions produce canonical findings and route to the expected keycheck candidate services without network verification
|
||||
|
||||
#### Scenario: Executable is unavailable
|
||||
- **WHEN** no TruffleHog executable is available in a general unit-test environment
|
||||
- **THEN** the real-CLI test is explicitly skipped while policy-structure and normalization unit tests still execute
|
||||
@@ -0,0 +1,16 @@
|
||||
## 1. Detector Policy
|
||||
|
||||
- [x] 1.1 Split Xai context-before and context-after alternatives into uniquely named single-regex detector entries.
|
||||
- [x] 1.2 Split ZaiGLM context-before and context-after alternatives into uniquely named single-regex detector entries.
|
||||
- [x] 1.3 Canonicalize the compatibility-only detector names before persistence and candidate extraction.
|
||||
|
||||
## 2. Compatibility Coverage
|
||||
|
||||
- [x] 2.1 Add pure unit coverage for detector policy structure and canonical alias normalization.
|
||||
- [x] 2.2 Add an optional real-TruffleHog CLI regression test for both context directions and candidate routing.
|
||||
|
||||
## 3. Verification
|
||||
|
||||
- [x] 3.1 Run the focused provider and scanner unit tests.
|
||||
- [x] 3.2 Run the real CLI compatibility test with the current Windows binary and a Linux binary where available.
|
||||
- [x] 3.3 Validate the OpenSpec change and confirm the patch is formatting-clean.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-06-22
|
||||
@@ -0,0 +1,82 @@
|
||||
## Context
|
||||
|
||||
The scanner runs continuously and appends findings to `found_secrets.jsonl` and `scanner_active.db`. Keycheckers classify provider credentials into current-state status files under `runtime/keychecks/<service>/` and opportunistically write rows to `keycheck_results` for dashboard visibility.
|
||||
|
||||
The current behavior has several failure modes:
|
||||
|
||||
- Hourly keychecks can replay a multi-GB `found_secrets.jsonl`, delaying or blocking later services in the batch.
|
||||
- Some checkers implement custom input readers, so global tail behavior does not apply consistently.
|
||||
- Known keys are skipped before writing a new occurrence row, so repeated source/query hits for an already alive key are invisible in DB/dashboard history.
|
||||
- Per-result DB writes compete with scanner writes and can fail under SQLite locks, causing file state and DB observations to diverge.
|
||||
- Dashboard views mix file current-state, DB current-state, historical occurrences, and provider-specific usable access semantics.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make current-state status files remain the authoritative per-service status store.
|
||||
- Record historical occurrences for known keys without forcing API rechecks.
|
||||
- Make hourly keychecks bounded and incremental enough to keep up with continuous scanning.
|
||||
- Make DB writes resilient and explainable, with visible lag/error indicators.
|
||||
- Give operators dashboard presets for current usable keys, historical usable occurrences, no-quota/limited states, and pipeline health.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Replace provider status files with the SQLite DB.
|
||||
- Revalidate every known alive key on every hourly cycle.
|
||||
- Guarantee zero SQLite lock contention while scanner writers are active.
|
||||
- Redesign detector extraction or provider-specific validity semantics beyond accounting and visibility.
|
||||
|
||||
## Decisions
|
||||
|
||||
1. Keep status files as current-state truth.
|
||||
|
||||
Rationale: checkers already compact/move keys between `Alive`, `NoBalance`, `Dead`, `Network`, and related files. Replacing this would be riskier than making DB observations catch up.
|
||||
|
||||
Alternative considered: make `keycheck_results` the source of truth. Rejected because the active DB is large, frequently locked, and dashboard reads must not block scanner writes.
|
||||
|
||||
2. Introduce occurrence recording for skipped known keys.
|
||||
|
||||
When a checker sees a candidate that is already known in status files or checked files, it should write a lightweight occurrence row with the cached status, source line, and finding attribution. It must not call provider APIs unless retry flags or recheck flags require it.
|
||||
|
||||
Alternative considered: only record fresh API checks. Rejected because this hides repeated source/query yield for already alive keys.
|
||||
|
||||
3. Use a shared bounded/incremental input reader for all checkers.
|
||||
|
||||
Checkers should use common reader helpers rather than hand-rolled full-file loops. The minimum implementation can use a tail window; the target implementation should store per-service high-watermark offsets so hourly runs neither replay old data nor miss data outside a fixed tail window.
|
||||
|
||||
Alternative considered: keep reading full JSONL and rely on skip sets. Rejected because the input is already multi-GB and causes long stalls.
|
||||
|
||||
4. Centralize DB recording or make per-checker DB writes lock-tolerant.
|
||||
|
||||
The preferred direction is batching result/occurrence rows through `keycheck_runner` after each service completes. A smaller intermediate step is to avoid schema initialization on every single keycheck write and retry lock failures with bounded backoff.
|
||||
|
||||
Alternative considered: ignore DB write failures because files are authoritative. Rejected because dashboard and source attribution depend on DB visibility.
|
||||
|
||||
5. Distinguish access tiers from raw provider statuses.
|
||||
|
||||
Dashboard should classify provider statuses into operator-facing tiers such as `usable_llm`, `alive_unproven_llm`, `no_quota`, and `quota_limited`, while still allowing exact status filtering for values like `BEDROCK`, `VERTEX`, `VALID_RATE_LIMITED`, and `ALIVE`.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Cached occurrence rows could be mistaken for fresh provider rechecks -> Label occurrence rows with a result source such as `cached_status` versus `api_check`.
|
||||
- Tail windows can miss old-but-newly-unchecked lines after downtime or file rewrites -> Prefer high-watermark offsets and detect file truncation/rotation.
|
||||
- DB batching can still fail if SQLite is locked for extended periods -> Keep files authoritative and surface DB write lag/errors in dashboard.
|
||||
- Provider semantics differ: e.g. Gemini `VALID_RATE_LIMITED`, AWS `BEDROCK`, GCP `VERTEX` -> Keep exact statuses available and use access tiers only as an additional view.
|
||||
- Rechecking known alive keys too often can spend quota or trigger provider limits -> Occurrence recording must not imply revalidation.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add shared input high-watermark/tail behavior to all keycheckers, starting with hand-rolled readers.
|
||||
2. Add cached occurrence recording for known keys using existing status maps.
|
||||
3. Add batched DB write path or harden lock retry behavior.
|
||||
4. Update dashboard to show file current-state, DB observation freshness, and historical/current presets separately.
|
||||
5. Backfill/repair missing occurrence attribution from current status files and recent findings where safe.
|
||||
|
||||
Rollback: disable cached occurrence writes and fall back to existing status-file behavior; status files remain unchanged.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should high-watermark state live in `runtime/keychecks/<service>/state.json` or a shared `runtime/state/keycheck_offsets.json`?
|
||||
- Should cached occurrences be written for every repeated finding or deduped per service/key/finding/source per day?
|
||||
- Which dashboard panel should be considered the primary operator view: file current-state or DB latest-current-state?
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
Keycheck current-state files, DB observations, and dashboard views can diverge, making it unclear whether usable provider keys are still being found and which sources produced them. This is urgent because the scanner is running continuously, but large input files, known-key skipping, and SQLite lock behavior can hide fresh usable findings from operator-visible stats.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add reliable keycheck accounting for fresh checks and known-key occurrences.
|
||||
- Track current-state file counts separately from DB-observed validation rows.
|
||||
- Ensure hourly keychecks process recent findings efficiently without replaying multi-GB JSONL inputs from the beginning.
|
||||
- Preserve source/query/finding attribution even when a key was already classified as alive, dead, no-balance, or limited.
|
||||
- Surface keycheck pipeline health, write lag, skipped-known counts, and usable/no-quota status in dashboard views.
|
||||
- Reduce DB lock impact on keycheck result recording so file state and DB state remain explainably consistent.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `keycheck-accounting`: Defines reliable current-state, historical occurrence, and dashboard visibility behavior for keycheck results.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
## Impact
|
||||
|
||||
- `app/keycheck_runner.py` and `app/keycheckers/*`: keycheck execution, input reading, skip behavior, and result recording.
|
||||
- `app/scanner_db.py`: keycheck DB writes, lock handling, and occurrence recording.
|
||||
- `app/dashboard.py`: operator-facing keycheck and usable-key reporting.
|
||||
- Runtime files under `runtime/keychecks/`: current-state status files remain authoritative but gain clearer relationship to DB observations.
|
||||
- Runtime DB `scanner_active.db`: keycheck rows and attribution semantics become more complete and auditable.
|
||||
@@ -0,0 +1,76 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Current-state status files remain authoritative
|
||||
The system SHALL keep per-service keycheck status files as the authoritative current-state classification for keys.
|
||||
|
||||
#### Scenario: Key status changes after recheck
|
||||
- **WHEN** a checker revalidates a key and receives a new status
|
||||
- **THEN** the key MUST be removed from other status files for that service and written to the status file for the new status
|
||||
|
||||
#### Scenario: Dashboard compares files and DB
|
||||
- **WHEN** dashboard displays keycheck totals
|
||||
- **THEN** it MUST be clear whether each count comes from current-state status files or from DB observation rows
|
||||
|
||||
### Requirement: Known key occurrences are recorded
|
||||
The system SHALL record an occurrence when a checker sees a candidate that is already known in service status files or checked files.
|
||||
|
||||
#### Scenario: Known alive key appears in a new finding
|
||||
- **WHEN** a key already classified as alive appears in a new scanner finding
|
||||
- **THEN** the system MUST record the new source/query/finding occurrence without requiring a provider API recheck
|
||||
|
||||
#### Scenario: Known dead key appears in a new finding
|
||||
- **WHEN** a key already classified as dead appears in a new scanner finding
|
||||
- **THEN** the system MUST record the new source/query/finding occurrence with a cached dead status
|
||||
|
||||
#### Scenario: Cached occurrence is distinguishable from API recheck
|
||||
- **WHEN** an occurrence row is written without calling the provider API
|
||||
- **THEN** the row MUST indicate that the status came from cached current-state classification
|
||||
|
||||
### Requirement: Keycheck input processing is bounded and consistent
|
||||
The system SHALL avoid replaying the full scanner JSONL input on every hourly keycheck run.
|
||||
|
||||
#### Scenario: Hourly keychecks run on a multi-GB input file
|
||||
- **WHEN** `found_secrets.jsonl` is large
|
||||
- **THEN** each checker MUST process only a bounded recent range or an incremental range since its last processed offset
|
||||
|
||||
#### Scenario: Checker has a custom input loop
|
||||
- **WHEN** a checker reads scanner findings
|
||||
- **THEN** it MUST use shared keycheck input-reading behavior or implement equivalent high-watermark/tail semantics
|
||||
|
||||
#### Scenario: Input file rotates or shrinks
|
||||
- **WHEN** a stored high-watermark offset is larger than the current input file size
|
||||
- **THEN** the system MUST reset the offset safely and continue processing without crashing
|
||||
|
||||
### Requirement: Keycheck DB observation writes are resilient
|
||||
The system SHALL make keycheck DB observation writes resilient to active scanner DB contention.
|
||||
|
||||
#### Scenario: SQLite database is temporarily locked
|
||||
- **WHEN** a keycheck result or occurrence is ready to record and SQLite is locked
|
||||
- **THEN** the system MUST retry with bounded backoff before reporting a DB write failure
|
||||
|
||||
#### Scenario: DB write fails after retries
|
||||
- **WHEN** all DB write retries fail
|
||||
- **THEN** the status file write MUST remain intact and the failure MUST be visible in logs or dashboard health
|
||||
|
||||
#### Scenario: Schema initialization would contend with active writers
|
||||
- **WHEN** a checker records a single result row
|
||||
- **THEN** it MUST NOT run schema initialization or migration DDL as part of that per-result write path
|
||||
|
||||
### Requirement: Dashboard exposes keycheck pipeline health
|
||||
The dashboard SHALL expose keycheck pipeline health and freshness separately from provider status counts.
|
||||
|
||||
#### Scenario: DB observations lag behind status files
|
||||
- **WHEN** status files are newer than the latest DB keycheck row
|
||||
- **THEN** dashboard MUST show that DB observation data is stale relative to file current-state
|
||||
|
||||
#### Scenario: Keycheck run is stuck on a service
|
||||
- **WHEN** the keychecks process has not advanced past a service for longer than expected
|
||||
- **THEN** dashboard or supervisor-visible status MUST make the stuck service and elapsed time visible
|
||||
|
||||
#### Scenario: Operator wants current usable keys
|
||||
- **WHEN** an operator selects current usable key view
|
||||
- **THEN** dashboard MUST use provider-specific access tiers while retaining exact status filters such as `BEDROCK`, `VERTEX`, `ALIVE`, and `VALID_RATE_LIMITED`
|
||||
|
||||
#### Scenario: Operator wants historical source yield
|
||||
- **WHEN** an operator selects historical yield view
|
||||
- **THEN** dashboard MUST include cached known-key occurrences so source/query yield is not lost after rechecks or known-key skips
|
||||
@@ -0,0 +1,36 @@
|
||||
## 1. Input Processing
|
||||
|
||||
- [x] 1.1 Add shared keycheck input state storage for per-service file path, file size, inode/signature if available, and last processed byte offset.
|
||||
- [x] 1.2 Extend shared keycheck input reader to support high-watermark processing with safe reset on file truncation or rotation.
|
||||
- [x] 1.3 Migrate hand-rolled readers in OpenAI, OpenRouter, and Gemini to the shared input reader.
|
||||
- [x] 1.4 Keep a bounded tail fallback for first run or missing state, with clear logging of the active input mode.
|
||||
|
||||
## 2. Known-Key Occurrence Recording
|
||||
|
||||
- [x] 2.1 Add a shared helper that resolves cached status for a key from checked/status files without provider API calls.
|
||||
- [x] 2.2 Add occurrence recording for skipped known keys, including service, cached status, source line, finding payload, and detector.
|
||||
- [x] 2.3 Mark cached occurrence rows distinctly from API-check rows in metadata.
|
||||
- [x] 2.4 Deduplicate cached occurrences per service/key/finding/source to avoid unbounded repeated rows.
|
||||
|
||||
## 3. DB Write Reliability
|
||||
|
||||
- [x] 3.1 Ensure per-result keycheck DB writes do not run schema initialization or migration DDL.
|
||||
- [x] 3.2 Add bounded lock retry/backoff for keycheck result and occurrence writes.
|
||||
- [x] 3.3 Add keycheck runner counters for DB write success, retry, and failure counts per service.
|
||||
- [x] 3.4 Log DB write failures without preventing status-file updates.
|
||||
|
||||
## 4. Dashboard Visibility
|
||||
|
||||
- [x] 4.1 Add dashboard panel comparing status-file current-state counts with latest DB observation timestamps.
|
||||
- [x] 4.2 Add keycheck pipeline health panel showing last completed service, current/stuck service, run duration, and DB write errors.
|
||||
- [x] 4.3 Add current usable-key view based on provider-specific access tiers and exact statuses.
|
||||
- [x] 4.4 Add historical source-yield view that includes cached known-key occurrences.
|
||||
- [x] 4.5 Label cached-status rows separately from fresh API-check rows in validation tables.
|
||||
|
||||
## 5. Verification
|
||||
|
||||
- [x] 5.1 Add a smoke test or script that runs a checker over a small fixture and verifies status-file writes plus DB occurrence rows.
|
||||
- [x] 5.2 Verify OpenAI, OpenRouter, Gemini, AWS, and GCP checkers use bounded/incremental input processing.
|
||||
- [x] 5.3 Verify a known alive key found in a new source/query creates a cached occurrence without an API recheck.
|
||||
- [x] 5.4 Verify dashboard can show current-state file counts even when DB observations are stale.
|
||||
- [x] 5.5 Run `python -m py_compile` on changed Python modules.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-09
|
||||
@@ -0,0 +1,70 @@
|
||||
## Context
|
||||
|
||||
Managed DockerHub source cycles configure an explicit account pool, but repository search bypasses it and calls the Hub search endpoint anonymously. Docker Hub rejects anonymous result windows beyond 200 entries; an isolated authenticated probe using the existing Hub bearer flow returned 100 results for pages 3, 20, 21, and 30 at a page size of 100. The standard search path also schedules every page before learning the reported result count and currently accepts successful pages when another expected page fails, so the source cycle completes and advances its numeric query cursor with incomplete discovery.
|
||||
|
||||
The account manager already provides secret-safe accounts, cached Hub bearer tokens, account rotation, endpoint-keyed cooldowns, and persisted auth events. The implementation must reuse those controls without changing tag resolution, Registry access, or immutable scan behavior.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Authenticate standard and recent DockerHub repository searches through the configured account pool.
|
||||
- Fail closed when an explicit pool has no usable account.
|
||||
- Support at most 30 authenticated search pages and avoid requests beyond the result count reported by page one.
|
||||
- Give each page exactly one additional attempt for transient transport or server failures.
|
||||
- Return a complete expected page set or fail the source cycle without enqueueing partial results or advancing its query cursor.
|
||||
- Keep logs, exceptions, and auth events free of credentials and bearer values.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Adding discovery queries, changing active page sizes, changing repository refresh cadence, or altering scan budgets.
|
||||
- Changing Docker tag selection, Registry bearer authentication, layer scanning, immutable target identity, or scan retry policy.
|
||||
- Guaranteeing that Docker Hub will indefinitely support 30 pages; the local cap remains a safety bound, not an external SLA.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Reuse Hub bearer authentication with a search-specific endpoint identity
|
||||
|
||||
Add a repository-search response helper that follows the existing tag-response pattern but uses the `hub_search` endpoint key. It obtains cached or fresh Hub bearer tokens, sends `Authorization: Bearer ...`, refreshes once after a 401, and rotates across the configured accounts for 401, 403, or 429 responses. Search cooldowns remain separate from `hub_tags` and `registry` cooldowns. The shared token cache remains per account because the same Hub token is valid for both Hub endpoints.
|
||||
|
||||
When the manager represents an explicit pool and no account is usable, the helper fails before making an anonymous request. A non-explicit legacy invocation may retain anonymous behavior for direct CLI compatibility.
|
||||
|
||||
Alternative considered: add a second login mechanism or cookie session. Rejected because the existing `/v2/auth/token` bearer flow was verified against authenticated search through page 30 and avoids another credential path.
|
||||
|
||||
### Fetch page one before parallel remainder
|
||||
|
||||
Raise the code-level page cap from 20 to 30. Standard discovery fetches page one first, validates its payload, derives the expected page count from its reported `count`, and submits only pages 2 through `min(requested, expected, 30)` concurrently. This removes ambiguous post-hoc suppression of failed pages and avoids requesting pages objectively outside the reported result set.
|
||||
|
||||
Alternative considered: keep launching all pages concurrently and ignore failures above the largest successful count. Rejected because a failed early page can make the inferred boundary unreliable, and unnecessary out-of-range requests consume account budget.
|
||||
|
||||
### Bound page retries inside the request primitive
|
||||
|
||||
The authenticated search helper gives each page at most two search GETs across transient retry, bearer refresh, and account rotation. Each GET calls `api_request` with one total network attempt; the outer page budget supplies the single bounded retry for `408`, `500`, `502`, `503`, `504`, or account-specific failures. A legacy anonymous invocation has no account handling and calls `api_request` with two total attempts directly. This keeps the page-wide budget independent of global proxy retry settings and prevents it from resetting during account rotation.
|
||||
|
||||
### Treat incomplete pagination as a source-cycle transport failure
|
||||
|
||||
Introduce a DockerHub discovery transport error analogous to the existing GitLab error. Any expected page that remains unavailable after its bounded request/account handling aborts the complete search result before tag resolution or enqueue. The configured-source runner records a failed cycle and returns without advancing the query cursor; it does not crash the long-running source process. A later cycle retries the same query, and queue uniqueness keeps successful rediscovery idempotent.
|
||||
|
||||
### Keep discovery authentication isolated from scan behavior
|
||||
|
||||
The change only replaces repository-search HTTP calls and their error propagation. Tag/manifest resolution continues to use its existing endpoint identities, retries, caches, resolver states, and immutable digest constraints.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Thirty authenticated pages increase search request volume] -> Preserve the hard cap, configured worker bounds, and account rotation; this change does not raise active source page settings.
|
||||
- [Page-one count can change while later pages are fetched] -> Treat page one as the cycle snapshot boundary; all pages within that boundary must still succeed, and DB deduplication handles overlap caused by result movement.
|
||||
- [A single page outage now rejects otherwise usable pages] -> This is intentional complete-or-fail behavior; one retry limits transient loss, and the unchanged cursor retries the query later.
|
||||
- [Fetching page one serially adds one request latency before parallel work] -> It prevents unnecessary pages and provides an authoritative expected set, which is more valuable than the small latency saving.
|
||||
- [All accounts can be temporarily unavailable] -> Fail closed and persist endpoint-specific auth state rather than silently reverting to the anonymous 200-result window.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Canonically stop the supervisor and verify all managed children are down.
|
||||
2. Deploy the scanner, runner, and focused regression tests without changing source query configuration.
|
||||
3. Run focused DockerHub authentication/pagination/cursor tests and the broader relevant scanner suites.
|
||||
4. Canonically restart the supervisor and verify PostgreSQL, pipeline workers, DockerHub source worker, and restart counters.
|
||||
5. Roll back by restoring the previous code under a canonical stop/start if authenticated search causes an operational regression; no data migration is required.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None. Authenticated access through page 30 and the explicit-pool behavior have been verified or are covered by deterministic tests.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
DockerHub repository discovery is currently anonymous even when a managed account pool is configured, so searches are limited to the anonymous 200-result window and partial page failures can silently advance the query rotation. Authenticated probing confirms the configured Hub bearer flow can retrieve at least 30 pages of 100 results, making reliable deeper pagination available without new credentials or dependencies.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Authenticate every managed DockerHub repository-search mode through the configured account pool and existing Hub bearer-token flow.
|
||||
- **BREAKING**: when an explicit DockerHub account pool is configured, fail closed if no account can authenticate instead of falling back to anonymous search.
|
||||
- Support an authenticated search window of up to 30 pages while retaining a bounded code-level limit.
|
||||
- Retry transient page failures once, rotate accounts for account-specific failures, and reject an incomplete expected page set rather than enqueueing partial discovery results.
|
||||
- Preserve the current query cursor when pagination fails, while continuing to advance it after complete or objectively exhausted pagination.
|
||||
- Keep tag resolution, Registry authentication, immutable-digest deduplication, scan retries, and queue disposition unchanged.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `dockerhub-search-pagination`: Authenticated, bounded, complete-or-fail DockerHub repository-search pagination using the managed account pool.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects DockerHub search/authentication in `app/scanner.py`, source-cycle failure propagation in `app/console_runner.py`, and focused scanner/runner tests.
|
||||
- Reuses existing configured DockerHub accounts, Hub bearer tokens, cooldowns, and auth-event persistence; no new external dependency or credential format is introduced.
|
||||
- Active query lists, refresh cadence, scan budgets, cold/failed target policy, and layer-aware scanning are out of scope.
|
||||
+68
@@ -0,0 +1,68 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Managed repository search is authenticated
|
||||
The system SHALL authenticate every standard and recent DockerHub repository-search request through the configured DockerHub account pool using a Hub bearer token.
|
||||
|
||||
#### Scenario: Configured account performs search
|
||||
- **WHEN** a managed DockerHub source cycle searches for repositories with an available account
|
||||
- **THEN** the system sends the search request with that account's bearer authorization and records success under the `hub_search` endpoint identity
|
||||
|
||||
#### Scenario: Explicit pool has no usable account
|
||||
- **WHEN** an explicit DockerHub account pool is configured but no account can authenticate or leave cooldown
|
||||
- **THEN** the system fails the repository search without making an anonymous fallback request
|
||||
|
||||
#### Scenario: Search authorization is rejected
|
||||
- **WHEN** a search request receives an account-specific 401, 403, or 429 response
|
||||
- **THEN** the system refreshes a rejected bearer once where applicable and rotates to another usable account within the configured pool
|
||||
|
||||
### Requirement: Authenticated pagination is bounded
|
||||
The system SHALL support up to 30 DockerHub search pages per query and SHALL cap larger requested page counts at 30.
|
||||
|
||||
#### Scenario: Thirty-page authenticated search
|
||||
- **WHEN** a query requests 30 pages and the first page reports at least 30 pages of results
|
||||
- **THEN** the system requests the complete page range from 1 through 30
|
||||
|
||||
#### Scenario: Request exceeds safety cap
|
||||
- **WHEN** a query requests more than 30 pages
|
||||
- **THEN** the system limits the search to pages 1 through 30 and records that the requested range was capped
|
||||
|
||||
#### Scenario: Reported result set is shorter
|
||||
- **WHEN** page one reports fewer results than the requested page range would contain
|
||||
- **THEN** the system requests only the pages required by that reported count
|
||||
|
||||
### Requirement: Page acquisition is bounded and complete
|
||||
The system SHALL give each expected search page at most two transient transport/server attempts and SHALL not return partial repository results when any expected page remains unavailable.
|
||||
|
||||
#### Scenario: Transient failure recovers
|
||||
- **WHEN** an expected page receives a retryable transport error or transient HTTP status on its first attempt and succeeds on its second attempt
|
||||
- **THEN** the system includes that page and completes discovery without another transient attempt
|
||||
|
||||
#### Scenario: Expected page remains unavailable
|
||||
- **WHEN** an expected page still fails after bounded retry and account handling
|
||||
- **THEN** the system raises a DockerHub discovery transport failure before tag resolution or repository enqueue
|
||||
|
||||
#### Scenario: Every expected page succeeds
|
||||
- **WHEN** all expected pages return valid payloads
|
||||
- **THEN** the system combines their repositories in page order and proceeds with existing deduplication and resolution behavior
|
||||
|
||||
### Requirement: Failed pagination preserves query rotation
|
||||
The system SHALL record incomplete DockerHub pagination as a failed source cycle and SHALL keep the current query cursor unchanged.
|
||||
|
||||
#### Scenario: Source cycle receives pagination failure
|
||||
- **WHEN** repository discovery raises a DockerHub discovery transport failure
|
||||
- **THEN** the source cycle finishes with failed status, enqueues no partial search result, and selects the same query for the next cycle
|
||||
|
||||
#### Scenario: Complete source cycle succeeds
|
||||
- **WHEN** repository discovery and the remaining source cycle complete normally
|
||||
- **THEN** the existing query-advance policy remains unchanged
|
||||
|
||||
### Requirement: Search authentication is secret-safe and isolated
|
||||
The system MUST NOT expose account credentials or bearer tokens through search logs, errors, or auth events, and SHALL preserve existing tag, Registry, immutable-digest, and scan-retry behavior.
|
||||
|
||||
#### Scenario: Search request fails
|
||||
- **WHEN** an authenticated search request fails or exhausts the account pool
|
||||
- **THEN** emitted diagnostics identify only the safe endpoint/status category without including usernames, credentials, bearer values, or request authorization headers
|
||||
|
||||
#### Scenario: Repository search implementation changes
|
||||
- **WHEN** authenticated search pagination is deployed
|
||||
- **THEN** existing DockerHub tag resolution, Registry authentication, target deduplication, and scan retry contracts remain unchanged
|
||||
@@ -0,0 +1,21 @@
|
||||
## 1. Runtime Safety
|
||||
|
||||
- [x] 1.1 Canonically stop the live supervisor and verify all managed child processes are down before editing `app`
|
||||
|
||||
## 2. Authenticated Search Implementation
|
||||
|
||||
- [x] 2.1 Make Hub bearer acquisition endpoint-aware and add a `hub_search` response path with explicit-pool fail-closed behavior, account rotation, and two-attempt transient request bounds
|
||||
- [x] 2.2 Raise the search safety cap to 30 pages and make standard pagination fetch page one first, derive the expected range, and reject any incomplete expected page set
|
||||
- [x] 2.3 Route recent-mode repository search through the same authenticated bounded response path
|
||||
- [x] 2.4 Add a DockerHub discovery transport failure path that records a failed source cycle without advancing the query cursor
|
||||
|
||||
## 3. Regression Coverage
|
||||
|
||||
- [x] 3.1 Cover bearer authorization, token refresh, account rotation/cooldown, and explicit-pool no-fallback behavior without exposing secrets
|
||||
- [x] 3.2 Cover the 30-page cap, page-one count boundary, ordered complete results, one transient retry, and rejection of unresolved partial pagination
|
||||
- [x] 3.3 Cover failed-cycle cursor retention and successful-cycle compatibility
|
||||
- [x] 3.4 Run focused and broader relevant scanner/runner test suites
|
||||
|
||||
## 4. Runtime Verification
|
||||
|
||||
- [x] 4.1 Canonically restart the supervisor and verify PostgreSQL, pipeline workers, DockerHub source health, restart counters, and absence of app bytecode artifacts
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-08-28
|
||||
@@ -0,0 +1,120 @@
|
||||
## Context
|
||||
|
||||
DockerHub discovery currently inspects a bounded tag page but stops after the first eligible digest. Different tags often alias the same manifest or share the same ordered layers, so increasing the tag count without content-aware selection would mostly multiply duplicate work. The Hub tag response identifies platform digests but does not expose their layers; those come from the Docker Registry v2 manifest API.
|
||||
|
||||
GitHub and GitLab updated-target admission currently uses provider timestamps. A claimed scan then points TruffleHog at a mutable repository URL with a rolling age boundary and a depth cap. The timestamp is useful for deciding that something may have changed, but it neither identifies the exact ref nor proves which commit was scanned. Existing PostgreSQL reservations provide the fence under which an immutable plan can be attached.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Spend each repository's Docker budget on up to three distinct ordered layer graphs.
|
||||
- Keep Docker tag, manifest, request, cache, and emitted-target counts explicitly bounded.
|
||||
- Bind GitHub and GitLab scans to a provider-resolved ref and commit SHA after a fenced claim.
|
||||
- Use the last successfully covered SHA as the incremental boundary for the same ref.
|
||||
- Preserve exact plan and coverage evidence through retries, worker failure, and mid-scan updates.
|
||||
- Keep first-scan history bounded while making every later delta independent of rolling age and depth limits.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Persist a global historical inventory of Docker layers.
|
||||
- Guarantee discovery of a branch that appears and disappears between metadata polling cycles.
|
||||
- Replace repository metadata search with a global push-event feed.
|
||||
- Remove source API rate limits, target timeouts, result bounds, or queue admission bounds.
|
||||
- Claim that a bounded first scan covered history older than its configured baseline depth.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Resolve platform manifests before selecting Docker targets
|
||||
|
||||
For each bounded Hub tag candidate, the resolver will choose the requested platform child digest, obtain a short-lived pull token for the public repository, and fetch that child manifest from the Docker Registry v2 API. A candidate identity contains its tag, update time, immutable platform manifest digest, and ordered layer digest tuple.
|
||||
|
||||
The top-level multi-platform index digest will not be emitted when a matching child digest exists. Registry responses must be JSON manifests of bounded size with valid `sha256` layer digests. Token requests use a fixed Docker authentication origin and repository pull scope; credentials are never placed in cache records or logs.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- Comparing tag names or manifest digests alone, because different manifests can still contain the same layer chain.
|
||||
- Pulling every candidate image before selection, because manifest metadata is sufficient and much cheaper.
|
||||
- Persisting every layer immediately, because within-resolution novelty provides the requested bounded diversity without a new authoritative subsystem.
|
||||
|
||||
### Select up to three graphs deterministically
|
||||
|
||||
The source setting `docker_images_per_repository` is clamped to one through three. After exact ordered-tuple deduplication, selection uses stable tie breaking:
|
||||
|
||||
1. The newest resolved graph.
|
||||
2. The remaining graph that contributes the most layers not present in the selected union, with recency as the tie breaker.
|
||||
3. The oldest remaining distinct graph.
|
||||
|
||||
If fewer distinct graphs exist, fewer targets are emitted. Reordered layer tuples remain distinct because layer order changes the image filesystem. Selected targets retain the existing immutable `repository@sha256:...` identity, so queue deduplication and Docker scan execution do not change.
|
||||
|
||||
Only complete graph-selection results are stored in the disposable tag cache. A partial manifest-resolution failure can return successfully resolved targets for the current cycle, but it is not cached as complete and is reported as partial coverage.
|
||||
|
||||
### Resolve and bind a Git plan after claim
|
||||
|
||||
Repository timestamps remain coarse admission signals. Once PostgreSQL has fenced a queue row and result reservation, the worker resolves the source-provided ref hint or the repository's current default branch through the GitHub or GitLab API. The resolver returns a normalized ref and exact head SHA.
|
||||
|
||||
The reservation is then bound transactionally to an immutable plan containing ref, head SHA, optional covered base SHA, plan mode, and baseline bounds. Binding requires the active reservation and claim lease token. A missing or malformed revision is a retryable source failure; the worker does not silently scan a mutable URL and does not advance exact coverage.
|
||||
|
||||
Metadata search can identify only repository-level activity, so its exact scope is the provider-resolved default branch. Event-backed discovery may supply a more specific ref. Polling cannot guarantee capture of transient or unadvertised refs; that limitation remains explicit.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- Resolving one SHA for every search result before admission, because that would spend scarce API quota on records that are never claimed.
|
||||
- Encoding SHA in queue target identity, because it would create unbounded rows and bypass existing changed-target coalescing.
|
||||
- Copying a timestamp into a covered field at claim, because failed work would look complete.
|
||||
|
||||
### Execute pinned baselines and deltas
|
||||
|
||||
An initial or ref-changed plan scans the exact head with the configured first-scan depth bound. A same-ref plan with a different successfully covered head scans the pinned head with `--since-commit <covered SHA>` and omits rolling age and maximum-depth limits. A plan whose resolved head already equals the covered head produces an exact no-op result.
|
||||
|
||||
The installed TruffleHog binary's ability to accept a commit SHA as `--branch` is a deployment contract and will be covered by a local repository contract test. If the covered base is unavailable after a force push or ref recreation, execution falls back to a pinned bounded baseline and labels the result as such; it never reports an incremental range as covered when the base was not usable.
|
||||
|
||||
This delta means all commits reachable from the pinned head after the covered boundary, not merely the final filesystem diff. An add-then-delete sequence in separate new commits therefore remains visible.
|
||||
|
||||
### Advance covered SHA only during fenced successful ingestion
|
||||
|
||||
`target_queue` stores the last successfully covered ref and head SHA. `result_reservations` stores the immutable claimed plan, and normalized scan metadata stores the executed plan and whether execution remained pinned. On fenced ingestion with queue disposition `done`, the plan in metadata must match the reservation. Only a successful pinned baseline, successful incremental scan, or exact no-op advances or confirms the covered head. Failed, deferred, unbound, or mutable fallback work leaves coverage unchanged.
|
||||
|
||||
Because the claimed head is immutable, a newer provider update during execution remains discoverable after completion and can create a later plan from the just-covered head.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Registry manifest calls increase Docker API traffic] -> Keep tag candidates and emitted graphs bounded, reuse one scoped token per repository resolution, and do not cache partial results as complete.
|
||||
- [Many tags alias one graph] -> Deduplicate exact ordered layer tuples before queue insertion.
|
||||
- [A source API is unavailable after claim] -> Produce a retryable source failure and refund through the existing bounded lifecycle without changing coverage.
|
||||
- [TruffleHog SHA branch behavior differs by version] -> Add a local contract test against the configured binary and fail closed when immutable pinning is unsupported.
|
||||
- [Force-pushed base is no longer reachable] -> Retry as a pinned bounded baseline and label the loss of incremental continuity.
|
||||
- [Default-branch resolution misses non-default branch activity] -> Record the resolved scope honestly and allow event-backed ref hints; a complete ref-event feed remains future work.
|
||||
- [First-scan depth remains bounded] -> Treat the first head as the future delta baseline without claiming unbounded historical coverage.
|
||||
- [Three Docker graphs can triple downstream work] -> Clamp the per-repository setting to three and retain all existing source, queue, timeout, and output bounds.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add nullable Git coverage and reservation-plan columns through idempotent PostgreSQL schema initialization.
|
||||
2. Deploy graph resolution, plan binding, and tests with explicit configuration gates.
|
||||
3. Enable `docker_images_per_repository: 3` for DockerHub while retaining the existing twenty-tag candidate bound.
|
||||
4. Enable exact Git planning for GitHub and GitLab; legacy rows start with no covered SHA and receive a pinned bounded baseline on their next admitted scan.
|
||||
5. Observe partial manifest resolution, distinct graph counts, exact no-ops, baseline resets, delta scans, failures, and strict usable yield.
|
||||
|
||||
Rollback is configuration-first. Set Docker images per repository back to one and disable exact Git planning; nullable schema additions remain inert and require no destructive migration.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether production evidence supports increasing the Docker candidate tag page beyond twenty without exhausting Hub rate limits.
|
||||
- Whether a later change should snapshot all advertised refs or consume a dedicated push-event feed for complete non-default-branch coverage.
|
||||
|
||||
## Implementation Evidence
|
||||
|
||||
Implementation completed and locally validated on 2026-08-28:
|
||||
|
||||
- Docker resolution uses the bounded Registry v2 bearer flow, immutable platform-child digests, ordered-layer graph deduplication, and deterministic newest/novel/oldest selection. Partial graph resolution emits only proven targets, retains the repository for retry, and is not cached as complete.
|
||||
- PostgreSQL reservations bind one canonical exact Git plan under the active queue/reservation lease fence. Covered ref/head state advances in the same fenced transaction as successful result ingestion after exact plan and execution-evidence comparison.
|
||||
- GitHub and GitLab resolve either an explicit branch ref or the provider default branch to an exact commit. Baseline, delta, no-op, and continuity-reset execution remain pinned to the bound head.
|
||||
- The checked-in core configuration enables three Docker graphs and exact Git planning with a baseline depth of 100, two ref-resolution attempts, a 10-second shared timeout, and a 1 MiB response limit.
|
||||
- `python -B -m pytest -p no:cacheprovider -q tests/test_exact_git_scan_planning.py tests/test_validated_high_scanner_fixes.py::DockerTagIdentityTests tests/test_runtime_safety_layer.py tests/test_pipeline_cutover_invariants.py tests/test_migration_runtime_safety.py tests/test_pipeline_postgres_integration.py::PipelinePostgresIntegrationTests::test_exact_git_plan_binding_and_coverage_are_fenced` completed with `141 passed`.
|
||||
- The PostgreSQL integration scenario covers idempotent/conflicting binding, baseline to delta to no-op progression, durable exact metadata, a provider update observed during an active scan, successful fenced coverage advancement, and a post-bind worker refund that cannot advance coverage.
|
||||
- The configured local TruffleHog binary passed the exact-SHA `--branch` contract test included in the focused suite.
|
||||
- `python -B -m py_compile app/scanner.py app/scanner_db.py app/console_runner.py tests/test_exact_git_scan_planning.py tests/test_pipeline_postgres_integration.py` completed successfully.
|
||||
- `openspec validate improve-core-scan-coverage --strict` reported the change as valid.
|
||||
|
||||
Retained rollout limits are twenty Docker tag candidates, at most three emitted distinct graphs, 8 MiB per Registry manifest, 1,000 index descriptors, 2,048 layers, bounded source/queue/result limits, and PostgreSQL-only exact Git plan binding. No production migration, process restart, or rollout was performed as part of implementation. Deployment still requires the offline idempotent runtime-safety migration before restarting sources. Configuration-first rollback remains `docker_images_per_repository: 1` plus disabling `exact_git_planning_enabled`; nullable schema additions may remain in place.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The core DockerHub, GitHub, and GitLab sources spend most of their scan budget on repeated image contents or mutable repository snapshots, while recent strict-usable yield remains near zero. The scanner needs to cover materially different Docker layers and bind Git work to exact revisions so that additional work buys new evidence rather than another pass over the same surface.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Select up to three Docker images per repository whose ordered layer graphs are distinct, instead of stopping at the first eligible tag.
|
||||
- Prefer the newest graph, a graph adding the most not-yet-selected layers, and an older divergent graph while continuing to deduplicate immutable digest targets.
|
||||
- Resolve Git discovery observations into immutable ref and commit identities before scanning.
|
||||
- Scan all commits introduced since the last successfully covered commit for that ref, rather than relying on a moving repository URL, age cutoff, and depth cap.
|
||||
- Persist immutable Git scan plans and successfully covered heads; never advance exact coverage on a failed or unpinned fallback scan.
|
||||
- Bound API enumeration and Docker/Git expansion through explicit configuration and report partial or inexact coverage honestly.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `docker-layer-graph-selection`: Bounded selection of materially distinct platform-specific Docker image layer graphs.
|
||||
- `git-ref-delta-scanning`: Immutable, per-ref Git scan planning and successful incremental coverage tracking.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
The change affects Docker Hub tag and registry-manifest resolution, GitHub and GitLab metadata resolution, source-cycle configuration, PostgreSQL queue/reservation/scan state, TruffleHog command construction, and focused scanner/runtime integration tests. It adds bounded registry and source API requests but does not change external service APIs or credential output formats.
|
||||
+49
@@ -0,0 +1,49 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Platform-specific layer graphs are resolved
|
||||
The system SHALL resolve each bounded Docker tag candidate to the requested platform child manifest and SHALL represent its image contents as the ordered sequence of valid layer digests.
|
||||
|
||||
#### Scenario: Multi-platform tag contains the requested platform
|
||||
- **WHEN** a tag exposes a matching `linux/amd64` child manifest
|
||||
- **THEN** the system uses that child manifest digest and its ordered layers rather than the top-level index digest
|
||||
|
||||
#### Scenario: Candidate manifest is malformed
|
||||
- **WHEN** a registry response is oversized, malformed, or contains invalid layer identities
|
||||
- **THEN** the candidate is not represented as a resolved layer graph
|
||||
|
||||
### Requirement: Docker selection covers distinct graphs
|
||||
The system SHALL emit no more than the configured one-to-three image targets per repository and SHALL NOT emit two candidates with identical ordered layer digest sequences.
|
||||
|
||||
#### Scenario: Tags alias one graph
|
||||
- **WHEN** multiple tags resolve to the same ordered layer sequence
|
||||
- **THEN** only the newest alias remains eligible for selection
|
||||
|
||||
#### Scenario: Three or more graphs are available
|
||||
- **WHEN** at least three distinct graphs resolve successfully
|
||||
- **THEN** the system selects the newest graph, the remaining graph adding the most not-yet-selected layers, and the oldest remaining distinct graph
|
||||
|
||||
#### Scenario: Fewer graphs are available
|
||||
- **WHEN** fewer distinct graphs resolve successfully than the configured maximum
|
||||
- **THEN** the system emits only the distinct graphs that exist
|
||||
|
||||
#### Scenario: Layer order differs
|
||||
- **WHEN** two manifests contain the same layer identities in a different order
|
||||
- **THEN** the system treats them as distinct graphs
|
||||
|
||||
### Requirement: Selected Docker targets remain immutable
|
||||
The system SHALL emit each selected target as the canonical requested-platform manifest digest identity `repository@sha256:<digest>`.
|
||||
|
||||
#### Scenario: Selected tag moves later
|
||||
- **WHEN** a tag is republished after discovery
|
||||
- **THEN** the queued target continues to identify the originally selected platform manifest digest
|
||||
|
||||
### Requirement: Docker graph resolution is bounded and honest
|
||||
The system SHALL retain hard tag-candidate, selected-graph, response-size, retry, and timeout bounds and SHALL NOT cache partial graph resolution as complete.
|
||||
|
||||
#### Scenario: Some manifest requests fail
|
||||
- **WHEN** at least one candidate graph resolves and another candidate fails transiently
|
||||
- **THEN** the system may emit the resolved targets for the cycle but records partial resolution and does not write a complete positive cache entry
|
||||
|
||||
#### Scenario: Every manifest request fails transiently
|
||||
- **WHEN** no candidate graph can be resolved because registry metadata is unavailable
|
||||
- **THEN** the repository remains retryable through the existing deferred-resolution lifecycle
|
||||
@@ -0,0 +1,68 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Git scans are bound to an exact revision
|
||||
The system SHALL resolve a normalized ref and exact commit SHA for each claimed GitHub or GitLab repository before invoking TruffleHog and SHALL bind that immutable plan to the active reservation.
|
||||
|
||||
#### Scenario: Repository search supplies no ref hint
|
||||
- **WHEN** a claimed repository came from metadata search without an exact ref
|
||||
- **THEN** the system resolves the provider's current default branch and its exact head SHA
|
||||
|
||||
#### Scenario: Discovery supplies an exact ref hint
|
||||
- **WHEN** an event-backed target includes a valid branch ref
|
||||
- **THEN** the system resolves and binds that specific ref instead of substituting the default branch
|
||||
|
||||
#### Scenario: Revision lookup fails
|
||||
- **WHEN** the provider API cannot return a valid ref and commit SHA
|
||||
- **THEN** the claim receives a bounded retryable source failure and no exact coverage state advances
|
||||
|
||||
### Requirement: Git updates scan every newly introduced commit
|
||||
The system SHALL scan the exact claimed head after the last successfully covered head for the same ref and SHALL NOT apply rolling age or maximum-depth limits to that incremental range.
|
||||
|
||||
#### Scenario: Same ref advances
|
||||
- **WHEN** ref `R` was successfully covered at commit `A` and now resolves to descendant commit `D`
|
||||
- **THEN** the scan is pinned to `D` with `A` as its boundary and includes commits introduced between them
|
||||
|
||||
#### Scenario: Secret is added and then deleted in the delta
|
||||
- **WHEN** one newly introduced commit adds a secret and a later newly introduced commit removes it
|
||||
- **THEN** both commits remain in scan scope even though the final filesystem snapshot is clean
|
||||
|
||||
#### Scenario: Head is unchanged
|
||||
- **WHEN** the resolved head equals the successfully covered head for the same ref
|
||||
- **THEN** the system records an exact no-op without launching a redundant repository scan
|
||||
|
||||
### Requirement: Git baseline and discontinuity handling remain pinned
|
||||
The system SHALL use a pinned bounded baseline for a first-seen ref or an unusable incremental base and SHALL identify that mode without claiming unbounded historical coverage.
|
||||
|
||||
#### Scenario: Ref has no covered head
|
||||
- **WHEN** an exact ref is claimed without prior successful coverage
|
||||
- **THEN** the system scans its pinned head using the configured baseline depth bound and establishes that head as the future delta boundary on success
|
||||
|
||||
#### Scenario: Covered base is unavailable
|
||||
- **WHEN** force push, ref recreation, or remote history removal makes the covered SHA unusable
|
||||
- **THEN** the system falls back to a pinned bounded baseline and records the continuity reset
|
||||
|
||||
### Requirement: Git coverage advances only after successful fenced work
|
||||
The system SHALL update a queue row's covered ref and head only when successful ingestion applies a matching immutable reservation plan.
|
||||
|
||||
#### Scenario: Exact scan succeeds
|
||||
- **WHEN** a pinned baseline or delta result is ingested with queue disposition `done` and its plan matches the active reservation
|
||||
- **THEN** the queue's covered ref and head advance to the claimed head
|
||||
|
||||
#### Scenario: Exact scan fails or is deferred
|
||||
- **WHEN** execution fails, times out, loses its fence, or receives a deferred disposition
|
||||
- **THEN** the previously covered ref and head remain unchanged
|
||||
|
||||
#### Scenario: Remote advances during a scan
|
||||
- **WHEN** a newer commit appears after the worker binds its immutable head
|
||||
- **THEN** successful completion advances coverage only to the bound head and leaves the newer update eligible for later discovery
|
||||
|
||||
### Requirement: Exact Git scope is observable
|
||||
The system SHALL durably record the executed ref, head, base, scan mode, baseline bound, and whether immutable execution was preserved.
|
||||
|
||||
#### Scenario: Operator inspects an incremental scan
|
||||
- **WHEN** an exact delta result is committed
|
||||
- **THEN** its normalized scan metadata identifies the covered range without exposing source credentials
|
||||
|
||||
#### Scenario: Metadata discovery observes repository-level activity
|
||||
- **WHEN** no branch-specific event exists
|
||||
- **THEN** observability identifies the provider-resolved default-branch scope rather than implying coverage of every repository ref
|
||||
@@ -0,0 +1,20 @@
|
||||
## 1. Docker Layer Graph Selection
|
||||
|
||||
- [x] 1.1 Add and clamp `docker_images_per_repository` configuration through source argument construction and all Docker tag-resolution call sites.
|
||||
- [x] 1.2 Implement bounded Docker Registry token, platform-manifest, and ordered-layer resolution with strict response validation.
|
||||
- [x] 1.3 Implement deterministic distinct-graph selection, immutable child-digest targets, partial-result handling, and cache-version invalidation.
|
||||
- [x] 1.4 Add focused tests for aliases, platform children, novelty and age selection, malformed manifests, partial failures, and resolution plumbing.
|
||||
|
||||
## 2. Exact Git Revision Planning
|
||||
|
||||
- [x] 2.1 Add idempotent PostgreSQL fields and fenced methods for immutable reservation plans and last successfully covered Git ref/head.
|
||||
- [x] 2.2 Implement bounded GitHub and GitLab ref/head resolvers with default-branch and explicit-ref handling.
|
||||
- [x] 2.3 Extend Git scan execution for exact no-op, pinned baseline, and unbounded-by-age/depth delta modes, including continuity-reset fallback.
|
||||
- [x] 2.4 Bind plans after claims, include them in result metadata, and advance covered heads only during matching successful ingestion.
|
||||
- [x] 2.5 Add focused tests for plan resolution, command construction, unchanged heads, failure/refund fencing, successful coverage, force-push fallback, and mid-scan updates.
|
||||
|
||||
## 3. Configuration And Verification
|
||||
|
||||
- [x] 3.1 Enable three distinct Docker graphs and exact Git planning for the core sources with bounded documented defaults.
|
||||
- [x] 3.2 Run focused Docker, Git, queue, schema, lifecycle, and OpenSpec validation suites.
|
||||
- [x] 3.3 Record implementation evidence and any retained rollout limits in the change artifacts.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-30
|
||||
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The packaged CLI persists a private JSON configuration after `install`, but first installation accepts credentials only through `--token`. Windows has a root launcher, the Linux image relies on its entrypoint, and assembled Linux packages have no root launcher. `attach` already provides live coherent status, but operators reasonably look for a `watch` command. Active documentation mixes platform-specific syntax and server-capacity administration with local lifecycle commands.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make the same lifecycle vocabulary available from the package root on Windows and native Linux and through a short Docker helper.
|
||||
- Keep the device token out of command arguments and shell history during installation.
|
||||
- Verify start, clean stop, attach, status, and watch behavior before publishing copy-paste cheatsheets.
|
||||
- Remove server cap `0` from routine worker maintenance instructions.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Change the worker protocol, server API, assignment cancellation, or capacity semantics.
|
||||
- Replace the installed private JSON configuration or expose its token.
|
||||
- Add an updater, public artifact registry, or service-manager integration for native Linux.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Treat `watch` as the live human status alias
|
||||
|
||||
`watch` uses the same authenticated local control stream and coherent status refresh as `attach`. It accepts the same optional bounded `--follow-seconds` argument and detaches without stopping the worker. Reusing the control path avoids a second polling implementation and behaves consistently where an OS `watch` utility is absent.
|
||||
|
||||
### Accept a strict YAML installation document
|
||||
|
||||
`install --config PATH` reads an exact YAML mapping containing `server`, `token`, and `parallelism`. `PATH=-` reads a bounded document from standard input for Docker. A path must be a private regular file; YAML uses `safe_load`, rejects aliases/extra fields through exact shape validation, and never changes the durable installed JSON schema. Direct `--server` and `--token` remain supported for compatibility but cannot be combined with `--config`.
|
||||
|
||||
### Add only a Linux package-root launcher
|
||||
|
||||
Assembled non-Windows packages receive executable `truf-worker` and `run-worker` shell launchers equivalent to the Windows command files. Docker keeps its existing entrypoint; the launcher also enables release tooling to export the assembled package for native Linux without inventing another client implementation.
|
||||
|
||||
### Separate active instructions from evidence reports
|
||||
|
||||
The general quickstart links concise Windows, Linux, and Docker cheatsheets and retains conceptual guidance. Each platform sheet starts from its artifact or Compose root and contains installation plus start, stop, attach, status, and watch commands. Historical validation reports remain unchanged even where they record earlier cap-zero experiments.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [A YAML file leaves a plaintext token on disk] -> Require private permissions, label it one-time installation input, and instruct deletion after successful install; the durable token remains in the existing private state.
|
||||
- [Standard input cannot prove source-file permissions] -> Bound and validate its content and document `chmod 600` on the redirected host file.
|
||||
- [Watch and attach appear redundant] -> Document watch as a discoverable alias rather than maintaining distinct semantics.
|
||||
- [Native Linux lacks automatic daemon startup] -> Keep start/stop in the portable CLI and explicitly leave systemd installation out of scope.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add parser, YAML input, launcher, and tests without changing existing command forms.
|
||||
2. Build fresh Windows and Linux/Docker artifacts and run isolated lifecycle checks.
|
||||
3. Publish the updated general guide and platform sheets only after those checks pass.
|
||||
4. Roll back by restoring the previous artifact; installed schema-2 JSON remains compatible.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None.
|
||||
@@ -0,0 +1,23 @@
|
||||
## Why
|
||||
|
||||
Remote-worker instructions do not provide independently verified copy-paste workflows for Windows, native Linux, and Docker. They also document a nonexistent `watch` command, require a device token in the install command line, claim a native Linux launcher that is not packaged, and recommend server cap `0` for routine local lifecycle operations.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a bounded `watch` lifecycle command alongside `start`, `stop`, `attach`, and `status`.
|
||||
- Allow first-time installation to read the server origin, device token, and parallelism from a small YAML document instead of process arguments.
|
||||
- Package a native Linux launcher and verify the lifecycle command surface on Windows, native Linux, and Docker.
|
||||
- Keep the complete operator guide concise while adding separate copy-paste cheatsheets for each supported environment.
|
||||
- Remove routine cap `0` instructions; local graceful stop drains the selected worker without changing server scheduling for other devices.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `worker-operator-lifecycle`: Cross-platform worker lifecycle commands, private YAML installation input, and verified platform cheatsheets.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
## Impact
|
||||
|
||||
This affects `worker_cli.py`, worker package launchers, the Linux worker image, CLI/package tests, packaged lifecycle verification, and remote-worker operator documentation. The worker protocol, server API, stored private configuration schema, and assignment capacity model remain unchanged.
|
||||
+41
@@ -0,0 +1,41 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Cross-platform lifecycle command surface
|
||||
The packaged worker SHALL expose start, stop, attach, status, and watch operations with equivalent local-control semantics on Windows, native Linux, and Docker.
|
||||
|
||||
#### Scenario: Operator watches a running worker
|
||||
- **WHEN** an operator runs `watch` for a bounded interval
|
||||
- **THEN** the CLI emits coherent live status and detaches without stopping the worker
|
||||
|
||||
#### Scenario: Operator stops a worker
|
||||
- **WHEN** an operator requests a graceful stop while the worker can complete its local drain
|
||||
- **THEN** the CLI returns a clean shutdown receipt without requiring a server-side assignment-cap change
|
||||
|
||||
### Requirement: Installation credentials can come from YAML
|
||||
The worker SHALL accept a bounded strict YAML installation document containing the HTTPS server origin, device token, and positive local parallelism instead of requiring those values in process arguments.
|
||||
|
||||
#### Scenario: Install from a private YAML file
|
||||
- **WHEN** an operator invokes `install --config` with a private valid YAML file
|
||||
- **THEN** the worker verifies the package and persists the existing private installed configuration without exposing the token in argv
|
||||
|
||||
#### Scenario: Install from redirected standard input
|
||||
- **WHEN** a Docker operator redirects a valid YAML document to `install --config -`
|
||||
- **THEN** the worker performs the same installation without placing the token in the Compose or container command
|
||||
|
||||
#### Scenario: Reject ambiguous installation input
|
||||
- **WHEN** YAML input is malformed, has extra fields, exceeds its byte bound, or is combined with direct server/token arguments
|
||||
- **THEN** installation fails before writing worker configuration
|
||||
|
||||
### Requirement: Native Linux package is directly operable
|
||||
An assembled native Linux worker package SHALL include an executable package-root launcher for the same operator CLI used by Windows and Docker.
|
||||
|
||||
#### Scenario: Linux operator runs from the extracted package root
|
||||
- **WHEN** the operator invokes `./truf-worker status`
|
||||
- **THEN** the integrity-checking bootstrap runs with the package-local application and dependencies
|
||||
|
||||
### Requirement: Platform cheatsheets are executable and capacity-independent
|
||||
The active operator documentation SHALL provide separate Windows, native Linux, and Docker copy-paste cheatsheets whose lifecycle commands are verified against the corresponding packaged artifact and do not instruct routine use of server cap `0`.
|
||||
|
||||
#### Scenario: Operator follows one platform sheet
|
||||
- **WHEN** an operator starts from the documented artifact or Compose root
|
||||
- **THEN** installation, start, status, attach, watch, and clean stop require only the documented local files and commands
|
||||
@@ -0,0 +1,16 @@
|
||||
## 1. Operator CLI
|
||||
|
||||
- [x] 1.1 Add strict private YAML and standard-input configuration to worker installation
|
||||
- [x] 1.2 Add the bounded watch alias over the existing attach control stream
|
||||
- [x] 1.3 Add an executable native Linux package-root launcher
|
||||
|
||||
## 2. Platform Documentation
|
||||
|
||||
- [x] 2.1 Create separate copy-paste cheatsheets for Windows, native Linux, and Docker
|
||||
- [x] 2.2 Update the general quickstart and operations runbook and remove routine cap-zero guidance
|
||||
|
||||
## 3. Verification
|
||||
|
||||
- [x] 3.1 Add focused CLI, package, YAML-security, and documentation tests
|
||||
- [x] 3.2 Build fresh Windows and Linux/Docker artifacts and verify install, start, status, attach, watch, and clean stop
|
||||
- [x] 3.3 Run focused tests and strict OpenSpec validation and record the checked command matrix
|
||||
@@ -0,0 +1,79 @@
|
||||
# Worker Cheatsheet Validation - 2026-09-30
|
||||
|
||||
## Scope
|
||||
|
||||
Validation used isolated fake device tokens, private local state roots, a local
|
||||
TLS no-work fixture, unique Docker names/volumes, and fresh artifacts. It did not
|
||||
contact production, issue assignments, or use production credentials.
|
||||
|
||||
## Initial audit
|
||||
|
||||
- Windows package exposed install, start, stop, status, attach, logs, history,
|
||||
and doctor; `watch` was rejected as an invalid command.
|
||||
- The Linux image exposed the same command set, but an assembled native Linux
|
||||
package had no package-root launcher or preparation workflow.
|
||||
- First installation required token-bearing process arguments. Later lifecycle
|
||||
commands already read the private installed JSON configuration.
|
||||
- Active quickstart and operations instructions recommended server cap zero for
|
||||
routine local maintenance.
|
||||
- An unreachable-server negative check made `stop --timeout 120` return a
|
||||
non-drained receipt with exit code 2 instead of reporting false success.
|
||||
|
||||
## Fresh artifact identities
|
||||
|
||||
| Artifact | Identity |
|
||||
| --- | --- |
|
||||
| Final Windows ZIP SHA-256 | `210eb61e8d6c35b014b39b18b4dd7e28e1e057a11d57790188f7904487bac004` |
|
||||
| Windows package manifest | `4572e349cead890c4efdc113e4d41285981086db7c0783fb67493fdeb6bac04c` |
|
||||
| Linux/Docker image | `sha256:8bc99d9e7e5f5f364de9b7d2b30100942b5ce3d9170a64e5abf0068ffc3d02c4` |
|
||||
| Linux package manifest | `19907f29382bd2b5a1de13fd53a829e97a5626fbacbed8f4286071ba4d5a9bec` |
|
||||
|
||||
The final Windows ZIP was rebuilt after the last documentation correction. Its
|
||||
full lifecycle run used the same package-manifest identity; the rebuild changed
|
||||
only packaged operator-document bytes outside worker code authority.
|
||||
|
||||
## Checked command matrix
|
||||
|
||||
| Environment | YAML install | Doctor | Start | Status | Attach | Watch | Stop |
|
||||
| --- | --- | --- | --- | --- | --- | --- | --- |
|
||||
| Windows package | pass | pass | pass | pass | pass | pass | `drained=true`, `exit_code=0` |
|
||||
| Native Linux package under `/opt` | pass | pass | pass | pass | pass | pass | `drained=true`, `exit_code=0` |
|
||||
| Docker image with persistent volume | stdin pass | pass | pass | pass | pass | pass | `drained=true`, `exit_code=0` |
|
||||
|
||||
The stopped Docker container reported `status=exited`, `exit=0`, and
|
||||
`oom=false`. `attach` and `watch` detached without stopping the verified worker
|
||||
on every environment.
|
||||
|
||||
## Documentation result
|
||||
|
||||
The active general guide and operations runbook contain no cap-zero routine and
|
||||
no token-bearing install command. Separate Windows, native Linux, and Docker
|
||||
cheatsheets contain copy-paste install, start, stop, attach, status, and watch
|
||||
commands. Historical dated validation reports retain factual records of earlier
|
||||
cap-zero experiments and are not active instructions.
|
||||
|
||||
## Final gates
|
||||
|
||||
- Focused worker/package/documentation matrix: `157 passed, 3 skipped`.
|
||||
- Windows packaged `watch --help`: pass.
|
||||
- Linux image launcher, preparation script, packaged cheatsheets, and shell
|
||||
syntax smoke: pass.
|
||||
- Worker Compose rendering: pass.
|
||||
- Python compilation: pass.
|
||||
- Strict OpenSpec validation: pass.
|
||||
|
||||
## Release build
|
||||
|
||||
The final `dist/release-20260930-linux` release includes both native Linux and
|
||||
Docker Linux artifacts plus all platform sheets under `cheatsheets/`. All
|
||||
entries in `SHA256SUMS.txt` passed verification. The Docker bundle also contains
|
||||
the sheets and both `workerctl` helpers, and all eight internal checksums passed.
|
||||
The native archive contained no absolute paths, parent traversal, or links; its
|
||||
launchers retained mode 0755 and passed an extracted `/opt` preparation and CLI
|
||||
smoke test.
|
||||
|
||||
| Release artifact | SHA-256 |
|
||||
| --- | --- |
|
||||
| Native Linux package | `e44717c9e84fc73d1d0734189da5b809c1c2271648868c05caea954bf46f6ebc` |
|
||||
| Docker Linux image archive | `80a59f180e7c1e4e5861427d94b44605a54ba08ed12198c63c3ae98c35678f63` |
|
||||
| Trusted Linux package manifest | `a3e8b73855d3b0854c5891cb5a10ff892aa0929e24046d2ce6fd28a245317e82` |
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-09
|
||||
@@ -0,0 +1,91 @@
|
||||
## Context
|
||||
|
||||
Managed DockerHub repository search is authenticated and bounded to two GET attempts per page, but standard discovery still builds one all-pages result before database admission. An unavailable expected page therefore discards successful pages and leaves the numeric query cursor on the same keyword. Database deduplication happens only after pagination, so known repositories cannot currently stop deeper requests.
|
||||
|
||||
DockerHub search results are ordered but not snapshot-stable. A delayed retry of page N cannot reconstruct the exact earlier result window, so the design must combine idempotent page admission with periodic deep coverage rather than claim snapshot completeness. Repository anchors and immutable digest scan targets already have durable PostgreSQL identities, but none of the existing scan, resolver, projection, or source-cycle tables is a valid leaseable queue for failed discovery-page work.
|
||||
|
||||
The active source uses a single DockerHub writer, a numeric main query cursor, and additive JSON state. The new query list is appended to preserve that cursor. Runtime configuration and application files may only be changed while the canonical supervisor is stopped.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Persist each valid DockerHub page before requesting or acting on later pages.
|
||||
- Stop ordinary pagination after two consecutive nonempty pages containing only identities known before the current pass.
|
||||
- Use a 30-page, 100-result ceiling for every normal and deep query pass.
|
||||
- Give each exact query a deep pass that bypasses seen-page stopping when 72 hours have elapsed since its last durable dispatch.
|
||||
- Delegate unavailable page/query work to a durable, bounded, fenced retry queue without blocking main keyword rotation.
|
||||
- Add the confirmed 12 product/framework queries and disable only periodic re-resolution of completed repository anchors.
|
||||
- Preserve authenticated fail-closed requests, the existing two-GET page budget, cold/failed target exclusion, and immutable-digest error retries.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Treating DockerHub pagination as a stable snapshot or guaranteeing successful external acquisition every 72 hours.
|
||||
- Re-enabling cold/failed targets, rescanning successful immutable digests, or changing tag/layer selection, keychecks, scan workers, or scan budgets.
|
||||
- Reusing scan/resolver queues for HTTP discovery work.
|
||||
- Changing GitHub, GitLab, HuggingFace, or other source behavior.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Admit repository pages incrementally
|
||||
|
||||
The managed PostgreSQL DockerHub path will fetch page one first, validate `count`, and process the expected range sequentially. A narrow database operation will normalize and deduplicate the page, determine preexisting identities, insert bare repository anchors using the existing unresolved-anchor semantics, and return safe counts plus internal normalized identities needed for the current-pass knownness decision. It will not invoke resolver or scan claiming. The existing resolver gate runs once after pagination, not once per page.
|
||||
|
||||
Sequential acquisition is selected over the current parallel remainder because requests for page three and beyond must be avoidable after pages one and two prove fully known. It also bounds in-memory results and makes page-level durability explicit. The authenticated page helper and its two-total-GET budget remain unchanged.
|
||||
|
||||
Knownness is evaluated before page admission. A page increments the streak only when it is nonempty and every normalized repository existed before the pass. Identities inserted on an earlier page in the same pass do not count as preexisting if they appear again after result movement. Empty pages end the available range. Lookup uncertainty fails open by resetting the streak and continuing deeper.
|
||||
|
||||
Before each managed pass, source state records an incomplete marker keyed by exact query and effective policy hash. A matching marker forces deep behavior on the next attempt and is cleared only after a durable `completed`, `completed_with_retries`, or `query_invalid` outcome. This prevents a failed pass from treating pages it admitted before the failure as old-enough evidence for an early stop after restart.
|
||||
|
||||
### Delegate gaps to a PostgreSQL discovery retry queue
|
||||
|
||||
Add a dedicated `discovery_retry_queue`; existing target, resolver, projection, outbox, and source-cycle tables have incompatible lifecycle and foreign-key semantics. Work is coalesced by a deterministic key over source, exact query, effective policy hash, pass kind, and page/range. Rows move through `pending -> leased -> deleted`, with retryable failure returning to `pending`, expired leases reclaimable, and removed/mismatched policy work moved to `held`.
|
||||
|
||||
Claims use bounded `FOR UPDATE SKIP LOCKED` selection and random lease tokens. Completion and retry transitions require the exact row, owner, and token. A stale worker cannot acknowledge or replace newer work. The worker renews the same fenced lease before each page in a multi-page retry, so a bounded per-page lease cannot expire merely because an entire range takes longer than one lease interval. Retryable failures use exponential backoff with a cap and are never silently dropped at a maximum attempt count. Provider cooldown refunds the dispatch attempt only when the claim has made no remote request and uses the trusted retry time. Diagnostics store fixed safe categories, never authorization material or arbitrary response bodies.
|
||||
|
||||
If page one is unavailable, enqueue query-level work because the expected range is unknown. If a later page fails, enqueue page work and continue when account availability permits. If the account pool is exhausted before a remaining tail can be attempted, coalesce the tail into range work instead of creating one row per unattempted page. Successful pages remain admitted. The source cycle may advance its main query only after every observed gap is durably delegated; retry persistence failure keeps the old cursor and marks the cycle failed.
|
||||
|
||||
At most one due retry item is processed per normal source-loop iteration, independently of the main cursor. Main query state is saved before retry work can affect the next iteration. Retry success admits repositories before fenced acknowledgement; retry failure changes only retry state. This prevents a persistent page outage from head-of-line blocking the 61-keyword rotation.
|
||||
|
||||
### Represent partial success explicitly
|
||||
|
||||
Add `completed_with_retries` as a source-cycle outcome for a pass whose successful pages were admitted and whose gaps were durably delegated. The numeric query cursor advances for `completed`, `completed_with_retries`, and `query_invalid`, but not for `failed`, `source_failed`, or `backlog_only`. Invalid payloads, database admission failure, and retry-enqueue failure remain hard cycle failures rather than endlessly retryable transport work.
|
||||
|
||||
Alternative considered: keep the keyword pinned while retaining successful pages. Rejected because it preserves head-of-line blocking. Alternative considered: advance after logging a failed page without durable work. Rejected because a moving search window can make the gap unrecoverable.
|
||||
|
||||
### Schedule deep passes by exact query
|
||||
|
||||
DockerHub source state gains a versioned map keyed by exact query, not numeric list position. Each record contains the effective search policy hash and UTC deep-dispatch timestamp. A query is deep-due when the record is absent, malformed, policy-mismatched, or at least 72 hours old. Its next main rotation pass ignores only the two-known-page stop; page-one count, 30-page cap, 100-result size, authentication, retries, and incremental admission remain mandatory.
|
||||
|
||||
The dispatch timestamp is written only after successful page admission and durable delegation of any gaps. Scheduling from dispatch time avoids deep-pass storms during prolonged provider failure. Appended queries are immediately due without invalidating existing query timestamps. Removed queries are pruned from state and their retry work is held. This provides one deep dispatch per exact query on its first normal selection after the 72-hour boundary; it is not an external-success SLA.
|
||||
|
||||
State-file loss causes conservative early deep passes, not missed passes. PostgreSQL remains authoritative for retry work and repository identity.
|
||||
|
||||
### Apply the confirmed discovery policy
|
||||
|
||||
Set DockerHub defaults to 30 pages and 100 results per page and align existing per-query page overrides so every query has the same acquisition ceiling. Append exactly the 12 confirmed terms, preserving existing order and numeric cursor safety.
|
||||
|
||||
Set `docker_repository_refresh_max_per_cycle` to zero. This disables periodic claims of completed/resolved repository anchors only. `refresh_registry` remains enabled, so keyword search still runs; pending initial anchors and partial/error resolver rows remain eligible; new immutable digests discovered through those paths are scanned; immutable scan errors retain their existing retries.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [The first pass of 12 new terms can inspect up to 36,000 search rows and create a large resolver backlog] -> Keep existing resolver, scan-worker, admission, and pipeline bounds; append terms rather than force a manual pass; observe backlog after natural rotation.
|
||||
- [Two known pages do not prove all deeper pages are known] -> Treat stopping as an optimization and bypass it per exact query every 72 hours.
|
||||
- [A delayed page retry sees a moving result window] -> Admit retries idempotently and rely on recurring deep passes for eventual coverage rather than snapshot claims.
|
||||
- [A retry table can grow during a prolonged outage] -> Coalesce deterministic work keys, use range rows for unattempted tails, bound claims per loop, expose pending/oldest-age metrics, and never advance when durable delegation itself fails.
|
||||
- [Sequential pages increase latency for a truly deep pass] -> Normal passes usually stop early; deep passes deliberately trade latency for bounded coverage and remain capped at 30 pages.
|
||||
- [Disabling completed-anchor refresh can miss a new digest in a repository that no longer appears in search] -> This is the user's temporary policy choice; initial and failed resolution continue, and the switch can be restored independently.
|
||||
- [JSON deep-state corruption can trigger extra load] -> Validate schema strictly, key by exact query/policy, and fail toward an early deep pass rather than suppressing coverage.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Create and validate the OpenSpec artifacts and deterministic tests before touching the live runtime.
|
||||
2. Canonically stop the supervisor and verify all managed workers and PostgreSQL are down.
|
||||
3. Apply the additive PostgreSQL schema migration, scanner/runner orchestration, source-state validation, tests, and configuration changes.
|
||||
4. Run focused unit/SQL/migration/integration tests, then the broader relevant scanner and runtime-safety suites.
|
||||
5. Canonically start the supervisor; verify PostgreSQL migration authority, pipeline readiness, DockerHub authentication, worker restart counters, retry/deep-state health, and absence of bytecode artifacts.
|
||||
6. Roll back under another canonical stop by restoring code/config and leaving the additive retry table dormant; repository and immutable-target inserts are idempotent and require no destructive rollback.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None. Keyword scope, 30x100 policy, 72-hour deep cadence, disabled completed-anchor refresh, preserved scan retries, and independent retry backlog were explicitly confirmed.
|
||||
@@ -0,0 +1,31 @@
|
||||
## Why
|
||||
|
||||
DockerHub discovery currently downloads its full configured page range before database deduplication, and one exhausted page discards every successful page while blocking keyword rotation. The authenticated 30-page window makes deeper discovery possible, but it needs incremental persistence, seen-page stopping, and durable retry delegation to remain efficient and avoid silent gaps or head-of-line blocking.
|
||||
|
||||
## What Changes
|
||||
|
||||
- **BREAKING**: replace complete-or-fail DockerHub pagination with page-level durable persistence; successful pages remain admitted when another page exhausts its two request attempts.
|
||||
- Stop an ordinary query pass after two consecutive nonempty pages whose repository identities were already known before the pass.
|
||||
- Run every DockerHub query with an effective ceiling of 30 pages and 100 results per page.
|
||||
- Bypass seen-page stopping for a full deep pass of each exact query at least once per 72-hour scheduling interval.
|
||||
- Delegate failed page/query acquisition to a fenced PostgreSQL discovery retry backlog before allowing the main keyword rotation to advance.
|
||||
- Add 12 DockerHub product/framework queries: `open-webui`, `ragflow`, `dify`, `flowise`, `crewai`, `n8n`, `langflow`, `autogen`, `browser-use`, `openhands`, `anythingllm`, and `agent-zero`.
|
||||
- Temporarily disable periodic re-resolution of completed DockerHub repository anchors while preserving initial resolution, partial/error resolver retries, immutable-digest scan retries, and normal keyword discovery.
|
||||
- Preserve authenticated fail-closed search, the two-GET page budget, target uniqueness, cold/failed target policy, and secret-safe diagnostics.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `dockerhub-incremental-discovery`: Incremental DockerHub page admission, seen-page stopping, 72-hour deep passes, and durable failed-page retry work.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects DockerHub pagination and authentication integration in `app/scanner.py`, source-cycle orchestration/state in `app/console_runner.py`, PostgreSQL schema and retry claims in `app/scanner_db.py`, and DockerHub settings in `app/config.yaml`.
|
||||
- Adds a PostgreSQL discovery retry queue with bounded leases, fencing, backoff, and configured-query/policy validation.
|
||||
- Changes DockerHub query rotation from failure-blocking to durable retry delegation and adds an initial bounded backlog from 12 new deep searches.
|
||||
- Requires focused unit, SQL-shape, migration, PostgreSQL integration, source-state, configuration, and runtime health verification.
|
||||
- Does not add dependencies or credential formats and does not change Docker tag selection, Registry authentication, layer scanning, scan workers, keychecks, or immutable-digest retry policy.
|
||||
+164
@@ -0,0 +1,164 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: DockerHub pages are admitted incrementally
|
||||
The system SHALL validate and durably admit every successful DockerHub repository-search page before relying on later-page acquisition, while preserving target identity and cold/failed-target exclusion.
|
||||
|
||||
#### Scenario: Successful page precedes a later failure
|
||||
- **WHEN** an expected DockerHub page is valid and a later expected page exhausts its request budget
|
||||
- **THEN** repositories from the successful page remain idempotently admitted and the later failure cannot roll them back
|
||||
|
||||
#### Scenario: Page admission fails
|
||||
- **WHEN** repository normalization or durable page admission fails
|
||||
- **THEN** the source cycle fails, does not treat the page as complete, and does not advance the main query cursor
|
||||
|
||||
#### Scenario: Resolver processing follows pagination
|
||||
- **WHEN** one or more pages admit repository anchors
|
||||
- **THEN** the existing Docker resolver gate runs at most once after the pass rather than once per page
|
||||
|
||||
### Requirement: Ordinary discovery stops after consecutive known pages
|
||||
The system SHALL stop an ordinary DockerHub query before requesting another page after two consecutive nonempty pages contain only repository identities that existed before the current pass.
|
||||
|
||||
#### Scenario: First two pages were previously known
|
||||
- **WHEN** pages one and two are nonempty and every normalized repository identity existed before the pass
|
||||
- **THEN** the pass completes without requesting page three
|
||||
|
||||
#### Scenario: Page contains a new repository
|
||||
- **WHEN** either of the last two pages contains a repository identity not known before the pass
|
||||
- **THEN** the consecutive-known counter resets and discovery continues within the effective range
|
||||
|
||||
#### Scenario: Current-pass duplicate appears later
|
||||
- **WHEN** a repository first admitted earlier in the same pass appears on a later page
|
||||
- **THEN** that identity is not treated as preexisting evidence for the later page's known-page stop
|
||||
|
||||
#### Scenario: Knownness cannot be determined
|
||||
- **WHEN** the database cannot establish complete pre-pass knownness for a valid page
|
||||
- **THEN** discovery fails open by continuing deeper rather than stopping early
|
||||
|
||||
#### Scenario: Previous pass ended incompletely
|
||||
- **WHEN** an exact query and policy have a retained incomplete-pass marker
|
||||
- **THEN** the next selected pass bypasses known-page stopping until a durable pass outcome clears the marker
|
||||
|
||||
### Requirement: Deep discovery bypasses seen-page stopping every 72 hours
|
||||
The system SHALL make each exact configured DockerHub query deep-due no later than 72 hours after its last durable deep dispatch and SHALL bypass the consecutive-known-page stop on that query's next selected pass.
|
||||
|
||||
#### Scenario: Exact query reaches its deep interval
|
||||
- **WHEN** at least 72 hours have elapsed since the query's last durable deep dispatch
|
||||
- **THEN** its next selected pass processes the available expected range without stopping on known pages
|
||||
|
||||
#### Scenario: Query is new or policy changes
|
||||
- **WHEN** an exact configured query has no valid state or its effective search-policy hash changes
|
||||
- **THEN** that query is immediately deep-due without resetting unrelated query schedules
|
||||
|
||||
#### Scenario: Deep pass has durable gaps
|
||||
- **WHEN** valid pages are admitted and unavailable pages are durably delegated to retry work
|
||||
- **THEN** the deep dispatch timestamp advances while completion remains represented by the outstanding retry work
|
||||
|
||||
#### Scenario: Provider prevents acquisition
|
||||
- **WHEN** neither successful pages nor durable retry delegation can establish a deep dispatch
|
||||
- **THEN** the system does not advance that query's deep timestamp
|
||||
|
||||
### Requirement: DockerHub page gaps use a durable retry backlog
|
||||
The system SHALL persist unavailable DockerHub query/page work in a dedicated PostgreSQL retry backlog before allowing the main keyword rotation to advance.
|
||||
|
||||
#### Scenario: Page one is unavailable
|
||||
- **WHEN** page one exhausts its bounded request/account handling and the expected range is unknown
|
||||
- **THEN** the system durably enqueues query-level retry work
|
||||
|
||||
#### Scenario: Later page is unavailable
|
||||
- **WHEN** a later expected page exhausts its bounded request/account handling
|
||||
- **THEN** the system durably enqueues page-level work and retains every successfully admitted page
|
||||
|
||||
#### Scenario: Remaining tail cannot be attempted
|
||||
- **WHEN** provider/account exhaustion prevents remote attempts for a known remaining page range
|
||||
- **THEN** the system coalesces the unattempted range into bounded retry work instead of creating unbounded individual rows
|
||||
|
||||
#### Scenario: Durable delegation succeeds
|
||||
- **WHEN** every observed acquisition gap has durable retry work
|
||||
- **THEN** the cycle records `completed_with_retries` and advances the main keyword independently of retry processing
|
||||
|
||||
#### Scenario: Durable delegation fails
|
||||
- **WHEN** retry work cannot be persisted authoritatively
|
||||
- **THEN** the cycle fails and retains the current keyword cursor
|
||||
|
||||
### Requirement: Discovery retry claims are bounded and fenced
|
||||
The system SHALL coalesce retry work by source, exact query, effective policy, pass kind, and page/range; SHALL process bounded due work with expiring leases; and MUST reject stale acknowledgements.
|
||||
|
||||
#### Scenario: Concurrent workers claim due work
|
||||
- **WHEN** multiple workers attempt to claim the same due retry row
|
||||
- **THEN** at most one receives the active lease token
|
||||
|
||||
#### Scenario: Lease expires
|
||||
- **WHEN** a worker fails to complete work before its lease expires
|
||||
- **THEN** the row becomes reclaimable without deleting its attempt history
|
||||
|
||||
#### Scenario: Multi-page retry remains active
|
||||
- **WHEN** a leased query or range retry is about to request another page
|
||||
- **THEN** the worker renews the same owner-and-token fence before acquisition and stops if renewal is rejected
|
||||
|
||||
#### Scenario: Stale worker finishes
|
||||
- **WHEN** a worker presents an obsolete owner or lease token
|
||||
- **THEN** it cannot acknowledge, delete, defer, or hold the newer work
|
||||
|
||||
#### Scenario: Retryable acquisition fails again
|
||||
- **WHEN** leased work encounters another retryable failure
|
||||
- **THEN** it returns to pending with bounded exponential backoff and is not silently dropped at an attempt limit
|
||||
|
||||
#### Scenario: Provider cooldown follows partial remote progress
|
||||
- **WHEN** an earlier page in the same claim made a remote request before a later local provider cooldown
|
||||
- **THEN** the dispatch attempt is not refunded
|
||||
|
||||
#### Scenario: Query is removed or policy is obsolete
|
||||
- **WHEN** retry work no longer matches an exact configured query and effective policy
|
||||
- **THEN** it is held and cannot execute against stale discovery policy
|
||||
|
||||
### Requirement: Retry work does not block main keyword rotation
|
||||
The system SHALL process at most a bounded amount of due discovery retry work per source-loop iteration independently of the saved main query cursor.
|
||||
|
||||
#### Scenario: Persistent page failure exists
|
||||
- **WHEN** a retry row remains unavailable across multiple attempts
|
||||
- **THEN** ordinary configured keywords continue rotating while the row follows its own backoff
|
||||
|
||||
#### Scenario: Retry succeeds
|
||||
- **WHEN** leased page or range work returns valid repositories
|
||||
- **THEN** repositories are durably admitted before the fenced retry acknowledgement
|
||||
|
||||
### Requirement: DockerHub search policy uses the confirmed breadth
|
||||
The system SHALL use an effective ceiling of 30 pages and 100 results per page for every configured DockerHub query and SHALL include exactly the 12 confirmed new product/framework terms in addition to the existing ordered query set.
|
||||
|
||||
#### Scenario: Ordinary pass reaches known content
|
||||
- **WHEN** a normal 30-by-100 query pass reaches two consecutive preexisting-known pages
|
||||
- **THEN** it stops early despite the larger configured ceiling
|
||||
|
||||
#### Scenario: Deep pass remains novel
|
||||
- **WHEN** a due deep pass continues to return pages containing new repositories
|
||||
- **THEN** it processes up to the page-one result boundary or the 30-page safety cap
|
||||
|
||||
#### Scenario: Query list is expanded
|
||||
- **WHEN** the configuration is loaded after deployment
|
||||
- **THEN** `open-webui`, `ragflow`, `dify`, `flowise`, `crewai`, `n8n`, `langflow`, `autogen`, `browser-use`, `openhands`, `anythingllm`, and `agent-zero` each appear once at the ordered tail
|
||||
|
||||
### Requirement: Completed repository refresh is disabled without changing retries
|
||||
The system SHALL disable periodic re-resolution of successfully completed DockerHub repository anchors while retaining normal search, initial/partial resolver work, and immutable-digest error retries.
|
||||
|
||||
#### Scenario: Completed anchor becomes periodically due
|
||||
- **WHEN** a resolved repository anchor reaches its prior refresh interval
|
||||
- **THEN** periodic policy does not claim it solely for refresh
|
||||
|
||||
#### Scenario: Initial or partial anchor is due
|
||||
- **WHEN** an unresolved or retryable partial repository anchor is due
|
||||
- **THEN** the existing resolver remains eligible to process it
|
||||
|
||||
#### Scenario: Immutable digest scan fails retryably
|
||||
- **WHEN** an immutable Docker image scan meets the existing retry conditions
|
||||
- **THEN** its current bounded retry and backoff behavior remains unchanged
|
||||
|
||||
### Requirement: Incremental discovery remains authenticated and secret-safe
|
||||
The system MUST preserve explicit-pool fail-closed Hub bearer authentication and MUST NOT persist or emit credentials, bearer values, authorization headers, raw response bodies, arbitrary exception text, or repository targets through retry/deep-state diagnostics.
|
||||
|
||||
#### Scenario: Search or retry fails
|
||||
- **WHEN** DockerHub page acquisition or retry processing reports an error
|
||||
- **THEN** diagnostics contain only bounded status/category/count metadata required for operation
|
||||
|
||||
#### Scenario: Explicit account pool is unavailable
|
||||
- **WHEN** no configured account can perform repository search
|
||||
- **THEN** normal and retry acquisition fail closed without anonymous fallback
|
||||
@@ -0,0 +1,30 @@
|
||||
## 1. Runtime Safety
|
||||
|
||||
- [x] 1.1 Canonically stop the live supervisor and verify all managed workers and PostgreSQL are down before editing application or configuration files
|
||||
|
||||
## 2. Durable Discovery Storage
|
||||
|
||||
- [x] 2.1 Add and validate the additive PostgreSQL discovery retry queue schema, required indexes, migration metadata, and bounded lifecycle fields
|
||||
- [x] 2.2 Implement strict page admission plus retry enqueue, claim, reclaim, backoff, hold, and fenced completion database operations
|
||||
- [x] 2.3 Add SQL-shape and PostgreSQL migration/integration coverage for idempotency, concurrent claims, expired leases, and stale-token rejection
|
||||
|
||||
## 3. Incremental DockerHub Discovery
|
||||
|
||||
- [x] 3.1 Refactor managed DockerHub pagination to validate, deduplicate, and durably admit successful pages sequentially without invoking the resolver per page
|
||||
- [x] 3.2 Implement two-consecutive-preexisting-page stopping with current-pass duplicate protection and fail-open knownness handling
|
||||
- [x] 3.3 Implement per-query policy-hashed 72-hour deep scheduling that bypasses only seen-page stopping
|
||||
- [x] 3.4 Delegate page-one, later-page, and unavailable-tail failures to durable retry work and advance the main cursor only after authoritative delegation
|
||||
- [x] 3.5 Process bounded retry work independently of main rotation with lease fencing, backoff, policy/query validation, and safe diagnostics
|
||||
|
||||
## 4. Discovery Policy
|
||||
|
||||
- [x] 4.1 Configure every DockerHub query for a 30-page by 100-result ceiling and append the exact 12 confirmed product/framework queries once
|
||||
- [x] 4.2 Disable periodic completed-anchor digest refresh while preserving initial/partial resolver and immutable-digest scan retries
|
||||
- [x] 4.3 Add permanent configuration regressions for query order/count, effective breadth, disabled periodic refresh, and unchanged retry/resource policy
|
||||
|
||||
## 5. Verification And Deployment
|
||||
|
||||
- [x] 5.1 Run focused pagination, retry-queue, source-state, migration, and configuration tests without creating application bytecode
|
||||
- [x] 5.2 Run broader relevant scanner, queue, runtime-safety, and PostgreSQL integration suites and complete an independent read-only review
|
||||
- [x] 5.3 Strictly validate the OpenSpec change and verify no secret-bearing diagnostics or unrelated behavior changes
|
||||
- [x] 5.4 Canonically start the runtime and verify PostgreSQL readiness, migrations, pipeline workers, DockerHub authentication/worker health, retry/deep state, restart counters, and absence of application bytecode
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-05
|
||||
@@ -0,0 +1,75 @@
|
||||
## Context
|
||||
|
||||
The configured sources contain 543 query occurrences covering 133 normalized terms. Canonical PostgreSQL lineage links discovery scans to credentials through both candidate and result records, and the dashboard already defines the strict `usable_llm` tier. Applying that rule across every provider identified 38 non-operational terms with zero historical strict-usable linkage despite 117,492 scans and 1,890.6 cumulative scanner-hours. The same terms consumed 18,641 scans and 317.1 scanner-hours in the latest 30-day window.
|
||||
|
||||
The runtime configuration is code-authority protected. Query rotation state, target queues, scan history, and credential history are separate persisted authorities and must not be rewritten to deploy this change.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Remove only globally zero-yield terms with enough exposure to support a conservative decision.
|
||||
- Apply one decision consistently anywhere the exact normalized term is configured.
|
||||
- Preserve operational source sentinels and all terms with any demonstrated strict-usable linkage.
|
||||
- Remove stale per-query overrides and verify deterministic post-prune query sets.
|
||||
- Deploy through the coordinated authority lifecycle and verify normal source rotation.
|
||||
|
||||
**Non-Goals:**
|
||||
- Deleting, reprioritizing, or rewriting existing target backlog or historical records.
|
||||
- Optimizing for one provider, broad `alive`, raw findings, or candidate volume.
|
||||
- Changing detector routing, keycheck classification, source concurrency, scan limits, or query-state files.
|
||||
- Claiming that a retired term can never produce a useful credential in the future.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Use all-provider strict-usable evidence
|
||||
|
||||
A term is eligible only when no credential linked to that term has ever reached the canonical dashboard `usable_llm` tier and no earliest-origin credential attributed to it has reached that tier. Candidate/result scan links are unioned before attribution so migrated and resolver-routed credentials are not lost.
|
||||
|
||||
The decision is global by case-normalized exact term. If a term produced one strict-usable credential for any provider or source, it remains configured everywhere. This is more conservative than pruning source-term pairs independently and avoids removing cross-provider terms such as `groq`, `llm`, `chat`, `rag`, or `langchain`.
|
||||
|
||||
Alternative: use broad `status_group=alive` or OpenAI-only yield. Rejected because broad alive contains unproven and historically misclassified statuses, while provider-only analysis can remove terms that work for another provider.
|
||||
|
||||
### Require meaningful exposure
|
||||
|
||||
A zero-yield term qualifies when either it has at least 30 linked credential observations, or it has at least 200 completed scan events and 20 cumulative scanner-hours. The credential branch tests precision; the cost branch catches terms that repeatedly consume work without reaching candidate intake. The threshold is applied to all retained history, with the latest 30-day cost recorded as corroborating evidence.
|
||||
|
||||
Alternative: remove every zero-yield term. Rejected because recent and low-sample terms have insufficient evidence. Those terms remain canaries.
|
||||
|
||||
### Exempt source-operational sentinels
|
||||
|
||||
`gharchive`, `gharchive-files`, and `gists` are sole query tokens used to operate dedicated sources rather than interchangeable discovery keywords. They remain even though they have no strict-usable attribution. Emptying those lists would disable or invalidate source operation rather than merely prune a search term.
|
||||
|
||||
### Retire the approved cohort consistently
|
||||
|
||||
Remove these 32 terms from GitHub, GitLab, DockerHub, npm, PyPI, and package-git: `autonomous`, `benchmarks`, `claw`, `code-assistant`, `codegen`, `dspy`, `embedding`, `embeddings`, `eval`, `evals`, `gateway`, `grok`, `haystack`, `inference`, `inference-api`, `knowledge`, `llamaindex`, `model`, `model-router`, `models`, `ollama`, `orchestration`, `prompts`, `replicate`, `rerank`, `reranker`, `retrieval`, `router`, `tokenizer`, `tool-use`, `vector`, and `vllm`.
|
||||
|
||||
Remove `chatgpt`, `gpt`, and `moonshot` from those six sources and Postman. Remove `openai` from GitHub, GitLab, DockerHub, and Postman. Remove `dashscope-intl.aliyuncs.com` and `generativelanguage.googleapis.com` from Postman.
|
||||
|
||||
This removes 219 occurrences. Resulting list sizes are GitHub 64, GitLab 46, DockerHub 46, npm 43, PyPI 43, package-git 43, and Postman 33.
|
||||
|
||||
### Keep deployment configuration-only
|
||||
|
||||
Delete the three `openai` query overrides together with the query entries. Existing rotation reads `query_index` modulo the current list length, so no persisted state edit is needed. Runtime is stopped before editing and restarted only after tests and strict OpenSpec validation.
|
||||
|
||||
Alternative: rewrite query indices or purge queued targets attributed to removed terms. Rejected because both mutate independent durable authority and are unnecessary for preventing future discovery.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Historical zero yield may not predict future supply] -> Keep low-sample terms, retain all historical evidence, and make rollback a configuration-only restoration.
|
||||
- [Earliest-origin attribution can hide useful rediscovery] -> Require zero strict-usable linkage across every scan link in addition to zero origin yield.
|
||||
- [Large list reduction changes rotation cadence] -> Verify exact list sizes and allow normal modulo-based state handling; do not edit source state.
|
||||
- [Completed OpenAI rollout previously required the literal term] -> Record the requirement retirement explicitly and retain higher-signal bounded ecosystem queries.
|
||||
- [Authority drift during a live edit] -> Use coordinated stop, test, and canonical start rather than relying on fail-close shutdown.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add configuration contract tests for the exact retired set, retained sentinels, uniqueness, post-prune sizes, and absence of orphaned overrides.
|
||||
2. Stop the authenticated runtime coordinately.
|
||||
3. Remove the 219 query occurrences and three matching overrides from `app/config.yaml`; do not edit state or queue data.
|
||||
4. Run focused query tests, configuration/runtime safety tests as applicable, and strict OpenSpec validation.
|
||||
5. Restart through `start_runtime.ps1` and verify authenticated supervisor, PostgreSQL, pipeline readiness, source processes, and query-list loading.
|
||||
6. Roll back by restoring the configuration entries and overrides through the same coordinated lifecycle if source health regresses.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None.
|
||||
@@ -0,0 +1,23 @@
|
||||
## Why
|
||||
|
||||
Current discovery rotations spend substantial scanner time on query terms that have accumulated meaningful exposure without linking to a single strict-usable credential for any provider. Removing only this globally zero-yield cohort reduces avoidable discovery and scan work while preserving every query with demonstrated usable yield.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Remove 38 sufficiently exposed, globally zero-yield search terms from the configured GitHub, GitLab, DockerHub, npm, PyPI, package-git, and Postman rotations.
|
||||
- Remove query-scoped overrides whose corresponding query is retired.
|
||||
- Preserve source-operational sentinel queries and every term linked to at least one historical strict-usable credential for any provider.
|
||||
- Preserve historical queue rows, scan results, credential lineage, deduplication state, and all runtime concurrency limits.
|
||||
- Define a repeatable evidence rule for future pruning instead of using raw findings or provider-specific yield alone.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `discovery-keyword-pruning`: Evidence-based, all-provider retirement of sufficiently tested zero-yield discovery terms.
|
||||
|
||||
### Modified Capabilities
|
||||
- `openai-discovery-coverage`: Retire the literal `openai` core query after its bounded rollout produced no strict-usable credential for any provider.
|
||||
|
||||
## Impact
|
||||
|
||||
The change affects `app/config.yaml`, focused query-configuration tests, and authority-managed source rotation after a coordinated restart. It removes 219 configured query occurrences but introduces no schema migration, dependency, queue rewrite, credential recheck, detector change, or scan-concurrency change.
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: All-provider strict-yield pruning rule
|
||||
The system SHALL retire a discovery term only when canonical lineage shows zero historical strict-usable credential linkage for every provider and the term has meaningful measured exposure.
|
||||
|
||||
#### Scenario: Any strict-usable linkage preserves a term
|
||||
- **WHEN** any credential linked to a configured term has ever met the canonical `usable_llm` rule
|
||||
- **THEN** that term SHALL remain in every configured source rotation
|
||||
|
||||
#### Scenario: Credential exposure qualifies a zero-yield term
|
||||
- **WHEN** a term has zero strict-usable linkage and at least 30 linked credential observations
|
||||
- **THEN** the term SHALL qualify for retirement
|
||||
|
||||
#### Scenario: Scanner-cost exposure qualifies a zero-yield term
|
||||
- **WHEN** a term has zero strict-usable linkage, at least 200 scan events, and at least 20 cumulative scanner-hours
|
||||
- **THEN** the term SHALL qualify for retirement
|
||||
|
||||
#### Scenario: Low-exposure zero-yield term remains a canary
|
||||
- **WHEN** a zero-yield term satisfies neither exposure condition
|
||||
- **THEN** it SHALL remain configured until more evidence is available
|
||||
|
||||
### Requirement: Approved global retirement cohort
|
||||
The system SHALL omit the approved 38-term zero-yield cohort from every source rotation where each exact term was configured.
|
||||
|
||||
#### Scenario: Shared broad-source cohort is removed
|
||||
- **WHEN** GitHub, GitLab, DockerHub, npm, PyPI, or package-git loads its query rotation
|
||||
- **THEN** it SHALL omit `autonomous`, `benchmarks`, `claw`, `code-assistant`, `codegen`, `dspy`, `embedding`, `embeddings`, `eval`, `evals`, `gateway`, `grok`, `haystack`, `inference`, `inference-api`, `knowledge`, `llamaindex`, `model`, `model-router`, `models`, `ollama`, `orchestration`, `prompts`, `replicate`, `rerank`, `reranker`, `retrieval`, `router`, `tokenizer`, `tool-use`, `vector`, and `vllm`
|
||||
|
||||
#### Scenario: Cross-source zero-yield terms are removed
|
||||
- **WHEN** an affected rotation is loaded
|
||||
- **THEN** `chatgpt`, `gpt`, and `moonshot` SHALL be absent from GitHub, GitLab, DockerHub, npm, PyPI, package-git, and Postman, and `openai` SHALL be absent from GitHub, GitLab, DockerHub, and Postman
|
||||
|
||||
#### Scenario: Zero-yield Postman signatures are removed
|
||||
- **WHEN** Postman loads its query rotation
|
||||
- **THEN** `dashscope-intl.aliyuncs.com` and `generativelanguage.googleapis.com` SHALL be absent
|
||||
|
||||
#### Scenario: Post-prune list sizes are deterministic
|
||||
- **WHEN** canonical configuration is loaded
|
||||
- **THEN** query counts SHALL be GitHub 64, GitLab 46, DockerHub 46, npm 43, PyPI 43, package-git 43, and Postman 33
|
||||
|
||||
### Requirement: Operational and historical authority is preserved
|
||||
Keyword retirement SHALL stop future discovery for the retired terms without deleting or rewriting source state, target queues, scans, findings, credentials, or results.
|
||||
|
||||
#### Scenario: Dedicated source sentinels remain
|
||||
- **WHEN** archive and gist source rotations are loaded
|
||||
- **THEN** `gharchive`, `gharchive-files`, and `gists` SHALL remain as their sole configured query tokens
|
||||
|
||||
#### Scenario: Persisted rotation index remains valid
|
||||
- **WHEN** an existing query index exceeds a shortened query list
|
||||
- **THEN** normal modulo-based rotation SHALL select a valid configured query without a state-file edit
|
||||
|
||||
#### Scenario: Existing backlog remains intact
|
||||
- **WHEN** the pruned configuration is deployed
|
||||
- **THEN** previously admitted targets and all historical attribution records SHALL remain unchanged
|
||||
|
||||
### Requirement: Retired query overrides are removed
|
||||
The canonical configuration SHALL NOT retain a query override for a retired query.
|
||||
|
||||
#### Scenario: Literal OpenAI overrides are absent
|
||||
- **WHEN** GitHub, GitLab, and DockerHub configuration is loaded
|
||||
- **THEN** each source SHALL omit the `openai` query override while preserving overrides for retained bounded queries
|
||||
+37
@@ -0,0 +1,37 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Query-scoped safety bounds
|
||||
The system SHALL support exact-query overrides for configured queries only, limited to `pages`, `per_page`, and `max_targets`, without changing source-wide defaults for other queries.
|
||||
|
||||
#### Scenario: Configured query receives bounded arguments
|
||||
- **WHEN** a source builds arguments for a configured query with an exact override
|
||||
- **THEN** it SHALL apply that query's configured page, page-size, and target bounds
|
||||
|
||||
#### Scenario: Retired query has no override
|
||||
- **WHEN** a query is removed from a source rotation
|
||||
- **THEN** the source SHALL NOT retain an override for that query
|
||||
|
||||
#### Scenario: Ordinary query retains source defaults
|
||||
- **WHEN** the same source builds arguments for any query without an override
|
||||
- **THEN** it SHALL retain the source-wide page, page-size, and target values
|
||||
|
||||
#### Scenario: Invalid override fails closed
|
||||
- **WHEN** a query override is not a mapping or contains a key outside the allowlist
|
||||
- **THEN** argument construction SHALL fail before discovery or queue mutation
|
||||
|
||||
## REMOVED Requirements
|
||||
|
||||
### Requirement: Exact OpenAI core discovery
|
||||
**Reason**: The completed bounded rollout produced 43 linked origin credentials and no strict-usable credential for any provider, meeting the approved global retirement rule.
|
||||
|
||||
**Migration**: Remove `openai` from GitHub, GitLab, and DockerHub rotations and allow normal modulo-based query rotation to continue without editing persisted source state.
|
||||
|
||||
### Requirement: Source-specific rollout limits
|
||||
**Reason**: The exact-query rollout is complete and its query is being retired, so source-specific `openai` execution bounds are no longer active policy.
|
||||
|
||||
**Migration**: Remove the three matching `openai` overrides while retaining the generic exact-query override mechanism and all overrides for configured ecosystem queries.
|
||||
|
||||
### Requirement: End-to-end canary evidence
|
||||
**Reason**: The exact-query canary reached terminal evidence and its measured all-provider strict yield is captured by the pruning decision.
|
||||
|
||||
**Migration**: Evaluate future keyword retirement under `discovery-keyword-pruning` using canonical all-provider lineage and measured exposure.
|
||||
@@ -0,0 +1,19 @@
|
||||
## 1. Configuration Contract
|
||||
|
||||
- [x] 1.1 Update focused query tests to assert the exact retired cohort, retained sentinels, retained productive terms, and deterministic list sizes.
|
||||
- [x] 1.2 Assert that retired queries have no orphaned query overrides while retained bounded ecosystem queries remain unchanged.
|
||||
|
||||
## 2. Runtime Configuration
|
||||
|
||||
- [x] 2.1 Stop the authenticated runtime through the coordinated lifecycle before changing authority-covered configuration.
|
||||
- [x] 2.2 Remove the approved 219 query occurrences and three `openai` overrides from `app/config.yaml` without modifying persisted state or backlog data.
|
||||
|
||||
## 3. Verification
|
||||
|
||||
- [x] 3.1 Run focused query/configuration tests and verify canonical configuration loads with the required query sets and counts.
|
||||
- [x] 3.2 Run strict OpenSpec validation and relevant supervisor/runtime safety tests.
|
||||
|
||||
## 4. Deployment
|
||||
|
||||
- [x] 4.1 Start the runtime through `start_runtime.ps1` and verify authenticated supervisor and managed PostgreSQL readiness.
|
||||
- [x] 4.2 Verify pipeline readiness, normal source processes, keychecks, and post-prune query loading without queue mutations.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-08-25
|
||||
@@ -0,0 +1,87 @@
|
||||
## Context
|
||||
|
||||
GitHub, GitLab, and HuggingFace discovery return stable target identities together with remote update timestamps. The runner currently discards those timestamps and the PostgreSQL queue permanently deduplicates targets by `(source, normalized_target)`. A completed repository or Space is therefore never scanned again even when its content changes.
|
||||
|
||||
Production discovery is already saturated: less than one percent of fetched records are new target identities and the core queue is often empty. The change must restore changed-content coverage without turning every rediscovery into a rescan, growing one queue row per revision, or creating an unbounded backlog.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Preserve a bounded remote content-update signal for GitHub, GitLab, and HuggingFace targets.
|
||||
- Rescan a completed target only when discovery observes content newer than the revision covered by its latest claim.
|
||||
- Coalesce multiple remote updates into one mutable queue row and one pending follow-up.
|
||||
- Bound changed-target admission and expose it separately from new-target admission.
|
||||
- Roll out without treating every legacy row as changed.
|
||||
|
||||
**Non-Goals:**
|
||||
- Requeue terminal failures or actively leased/deferred targets.
|
||||
- Rescan every known target on a timer.
|
||||
- Change DockerHub's digest-based identity and refresh behavior.
|
||||
- Guarantee that provider activity timestamps always represent content changes.
|
||||
- Enable inactive historical sources or enlarge global scan concurrency.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Carry one normalized discovery record
|
||||
|
||||
The runner will represent eligible discovery results internally as a target plus an optional UTC `remote_modified_at`. GitHub uses `pushed_at`, GitLab uses `last_activity_at`, and HuggingFace uses `lastModified`. Missing, malformed, or non-monotonic timestamps remain valid discovery results but cannot trigger an updated-target rescan.
|
||||
|
||||
HuggingFace discovery will use a newest-modified feed so old Spaces changed recently are observable. Identity-only known-page stopping will be disabled when updated-target rescans are enabled; hard page and result limits remain the discovery bound.
|
||||
|
||||
Alternatives rejected:
|
||||
- Repository `updated_at` on GitHub, because metadata-only edits are not content pushes.
|
||||
- A HEAD-SHA request per repository, because it multiplies API traffic and rate-limit exposure.
|
||||
- Revision-aware early stopping, because a known first page does not prove later pages contain no changed targets.
|
||||
|
||||
### Keep one queue row and two remote timestamps
|
||||
|
||||
`target_queue` will gain nullable `remote_modified_at` and `scan_remote_modified_at` columns. Discovery monotonically advances `remote_modified_at`. Claiming atomically copies the currently observed value into `scan_remote_modified_at`, recording what that scan covers.
|
||||
|
||||
The separate claim snapshot is required because a remote update can arrive while a scan is running. On completion, a newer observed timestamp remains ahead of the claimed timestamp and is eligible for exactly one later scan. Encoding revisions into `normalized_target` was rejected because it would grow queue rows and weaken queue authority.
|
||||
|
||||
For a legacy completed row with no claim snapshot, `completed_at` is the rollout baseline. It is eligible only when the first valid remote timestamp observed is newer than that completion. This prevents a migration surge while still admitting updates that occurred after the historical scan.
|
||||
|
||||
### Observe and requeue atomically under a hard budget
|
||||
|
||||
A PostgreSQL transaction will upsert discovery observations and requeue at most `updated_rescan_max_per_cycle` eligible rows for one source. Eligibility requires:
|
||||
- `status='done'`;
|
||||
- a strictly newer observed timestamp than `scan_remote_modified_at`, or than `completed_at` for a legacy row;
|
||||
- elapsed `updated_rescan_cooldown_seconds` since completion;
|
||||
- no lease, reservation, claim, or resolver authority.
|
||||
|
||||
Eligible rows are locked with `FOR UPDATE SKIP LOCKED`. Requeue resets only retry/completion scheduling fields required for a normal pending claim. Failed, pending, deferred, in-progress, unresolved, unchanged, and invalid-timestamp rows are never promoted by this path.
|
||||
|
||||
The initial production setting is one updated target per source cycle. Existing backlog-first behavior remains enabled, so a source drains its admitted work before discovery can admit more; `refresh_registry` is not enabled by this change.
|
||||
|
||||
Alternatives rejected:
|
||||
- Global `requeue_done=True`, because it requeues unchanged rows on every cycle.
|
||||
- Comparing only remote time with local completion time forever, because provider and host clocks differ and a mid-scan update can be lost.
|
||||
- A separate maintenance queue, because it duplicates existing lease, reservation, and completion authority.
|
||||
|
||||
### Account for updated targets separately
|
||||
|
||||
`source_cycles` will gain `queued_updated_count`. `queued_new_count` keeps its current meaning. Source logs and dashboard aggregation will report changed-target admissions separately so rollout volume and yield can be audited.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [GitLab activity can change without a repository push] -> Use a one-target-per-cycle cap and cooldown; report updated admissions separately.
|
||||
- [Provider clock skew] -> Require strict monotonicity and use the claimed remote timestamp after the first revision-aware scan.
|
||||
- [More discovery API traffic after disabling identity-only early stop] -> Keep existing hard page/per-page limits and source intervals.
|
||||
- [Changed targets consume capacity without useful findings] -> Start at one per cycle and compare updated-target yield before increasing the cap.
|
||||
- [Schema rollout while runtime is active] -> Stop the authority-managed runtime, apply schema through the normal initialization path, run tests, then restart through `start_runtime.ps1`.
|
||||
- [Rollback leaves nullable columns] -> Disable the feature in source configuration; nullable columns and metrics are backward-compatible and can remain.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add nullable queue columns, the source-cycle counter, and a partial eligibility index through idempotent schema initialization.
|
||||
2. Deploy code and tests with updated-target rescans disabled by default.
|
||||
3. Enable GitHub, GitLab, and HuggingFace with a cap of one and a conservative cooldown; leave DockerHub unchanged.
|
||||
4. Restart the authority-managed runtime and verify discovery, queue authority, projection, and source health.
|
||||
5. Observe changed-target admission, completion, findings, and worker occupancy before changing any cap.
|
||||
|
||||
Rollback is configuration-first: disable updated-target rescans and restart the managed runtime. Existing pending work completes under normal queue semantics; no destructive data migration is required.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether production evidence supports different cooldowns per source after the initial canary.
|
||||
- Whether a later change should add provider-specific immutable revisions when APIs can supply them without extra requests.
|
||||
@@ -0,0 +1,28 @@
|
||||
## Why
|
||||
|
||||
GitHub, GitLab, and HuggingFace discovery continuously observe recently changed repositories and Spaces, but the queue permanently suppresses every URL or Space ID after its first scan. Production consequently fetched roughly 290,000 discovery results in the latest 24 hours while admitting less than one percent as targets, leaving scan capacity underused and ignoring secrets added to already-known projects.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Preserve the remote content-update timestamp supplied by GitHub, GitLab, and HuggingFace discovery.
|
||||
- Requeue a completed target only when the observed remote content timestamp is newer than its last completed scan.
|
||||
- Bound changed-target admission per source cycle and enforce a per-target cooldown.
|
||||
- Keep active, deferred, failed, and unchanged targets untouched.
|
||||
- Expose changed-target requeue counts separately from newly discovered target counts.
|
||||
- Leave DockerHub digest-based refresh behavior unchanged.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `updated-target-rescan`: Safely detect and rescan remotely changed core repository and Space targets under bounded rollout controls.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Discovery metadata and target preparation in `app/scanner.py` and `app/console_runner.py`.
|
||||
- PostgreSQL target queue state, migrations, cycle accounting, and dashboard observability in `app/scanner_db.py` and `app/dashboard.py`.
|
||||
- Source configuration for GitHub, GitLab, and HuggingFace.
|
||||
- Focused queue, discovery, and PostgreSQL integration tests.
|
||||
@@ -0,0 +1,87 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Discovery preserves remote content recency
|
||||
The system SHALL preserve a normalized remote content-update timestamp for GitHub, GitLab, and HuggingFace discovery records and SHALL keep DockerHub digest-based identity behavior unchanged.
|
||||
|
||||
#### Scenario: Source-specific remote timestamp is retained
|
||||
- **WHEN** GitHub supplies `pushed_at`, GitLab supplies `last_activity_at`, or HuggingFace supplies `lastModified`
|
||||
- **THEN** the queue observation stores the valid UTC timestamp with the target identity
|
||||
|
||||
#### Scenario: Missing or malformed remote timestamp
|
||||
- **WHEN** a discovery result has no valid remote content-update timestamp
|
||||
- **THEN** the target remains eligible for normal new-target admission but MUST NOT trigger an updated-target rescan
|
||||
|
||||
#### Scenario: DockerHub discovery
|
||||
- **WHEN** DockerHub resolves an image tag
|
||||
- **THEN** the existing digest identity and refresh behavior remain authoritative without updated-target promotion
|
||||
|
||||
### Requirement: Only changed completed targets are promoted
|
||||
The system SHALL promote a known target only when it is completed, its observed remote timestamp is strictly newer than the remote timestamp covered by its latest scan, and its configured cooldown has elapsed.
|
||||
|
||||
#### Scenario: Completed target changed after its covered revision
|
||||
- **WHEN** discovery observes a newer valid remote timestamp for a completed target after cooldown
|
||||
- **THEN** the target becomes pending for one normal fenced scan
|
||||
|
||||
#### Scenario: Unchanged target is rediscovered
|
||||
- **WHEN** discovery observes the same or an older remote timestamp
|
||||
- **THEN** the completed target remains unchanged and unclaimable
|
||||
|
||||
#### Scenario: Legacy completed target is first observed
|
||||
- **WHEN** a completed target has no claimed remote timestamp
|
||||
- **THEN** it is promoted only if the observed remote timestamp is strictly newer than its last completion time
|
||||
|
||||
#### Scenario: Non-completed target is rediscovered
|
||||
- **WHEN** the target is failed, pending, deferred, in progress, unresolved, leased, claimed, or reserved
|
||||
- **THEN** updated-target discovery MUST NOT alter its lifecycle or authority fields
|
||||
|
||||
#### Scenario: Update arrives during a scan
|
||||
- **WHEN** discovery records a newer remote timestamp after the active claim captured its scan timestamp
|
||||
- **THEN** completion preserves the newer observation and permits one bounded follow-up scan after cooldown
|
||||
|
||||
### Requirement: Updated-target admission is bounded
|
||||
The system SHALL enforce a source-configured hard maximum of updated-target promotions per discovery cycle and a per-target cooldown.
|
||||
|
||||
#### Scenario: Eligible changes exceed the cycle budget
|
||||
- **WHEN** more completed changed targets are eligible than the configured maximum
|
||||
- **THEN** at most the configured maximum are promoted and the remainder stay eligible for later cycles
|
||||
|
||||
#### Scenario: Cooldown has not elapsed
|
||||
- **WHEN** a changed completed target was scanned within the configured cooldown
|
||||
- **THEN** it remains completed until a later eligible cycle
|
||||
|
||||
#### Scenario: Concurrent discovery cycles
|
||||
- **WHEN** multiple workers observe the same changed target concurrently
|
||||
- **THEN** transactional row fencing permits at most one promotion for the covered revision
|
||||
|
||||
### Requirement: Updated discovery remains capable of seeing changed known targets
|
||||
The system SHALL NOT use identity-only known-page stopping for a source while updated-target rescans are enabled and SHALL keep discovery bounded by explicit page and result limits.
|
||||
|
||||
#### Scenario: Known identities appear on an early page
|
||||
- **WHEN** an early discovery page contains only known target identities
|
||||
- **THEN** discovery continues within its configured hard page limit so later changed targets can be observed
|
||||
|
||||
#### Scenario: HuggingFace Space was created long ago and recently updated
|
||||
- **WHEN** a known Space has a recent `lastModified` value
|
||||
- **THEN** update-sorted HuggingFace discovery can observe it independently of creation time
|
||||
|
||||
### Requirement: Claims record the covered remote revision
|
||||
The system SHALL atomically snapshot the newest observed remote timestamp when a target is claimed.
|
||||
|
||||
#### Scenario: Revision-aware target is claimed
|
||||
- **WHEN** a pending target receives a valid lease and reservation
|
||||
- **THEN** its scan-covered timestamp equals the newest remote timestamp observed before that claim
|
||||
|
||||
#### Scenario: Claim fails before authority is committed
|
||||
- **WHEN** capacity or fencing prevents the claim
|
||||
- **THEN** the scan-covered timestamp MUST NOT advance
|
||||
|
||||
### Requirement: Updated-target activity is separately observable
|
||||
The system SHALL report updated-target promotions separately from newly discovered target admissions in durable cycle metrics, source logs, and dashboard summaries.
|
||||
|
||||
#### Scenario: Cycle admits new and updated targets
|
||||
- **WHEN** a discovery cycle inserts new identities and promotes changed completed identities
|
||||
- **THEN** `queued_new_count` and `queued_updated_count` record the respective counts without overlap
|
||||
|
||||
#### Scenario: No changed targets are promoted
|
||||
- **WHEN** a cycle observes only unchanged or ineligible known targets
|
||||
- **THEN** `queued_updated_count` is zero
|
||||
@@ -0,0 +1,24 @@
|
||||
## 1. Queue State And Atomic Admission
|
||||
|
||||
- [x] 1.1 Add idempotent PostgreSQL schema support for observed and scan-covered remote timestamps, updated-target cycle counts, and eligibility indexing.
|
||||
- [x] 1.2 Implement atomic discovery observation and bounded promotion for changed completed targets while preserving all non-eligible lifecycle authority.
|
||||
- [x] 1.3 Snapshot the observed remote timestamp only when a legacy or slot-first target claim commits.
|
||||
|
||||
## 2. Discovery And Runner Integration
|
||||
|
||||
- [x] 2.1 Preserve GitHub `pushed_at`, GitLab `last_activity_at`, and HuggingFace `lastModified` through target preparation.
|
||||
- [x] 2.2 Make HuggingFace discovery newest-modified and disable identity-only page stopping when updated rescans are enabled.
|
||||
- [x] 2.3 Wire source-specific enablement, per-cycle caps, and cooldowns without changing DockerHub or enabling refresh-while-backlogged discovery.
|
||||
|
||||
## 3. Observability
|
||||
|
||||
- [x] 3.1 Persist and log `queued_updated_count` separately from new-target admission.
|
||||
- [x] 3.2 Add updated-target admission to dashboard source-cycle summaries.
|
||||
|
||||
## 4. Verification And Rollout
|
||||
|
||||
- [x] 4.1 Add focused discovery and runner tests for timestamp preservation, invalid timestamps, known-page behavior, and DockerHub isolation.
|
||||
- [x] 4.2 Add PostgreSQL integration tests for legacy baselines, changed and unchanged targets, cooldown and cap enforcement, concurrent fencing, claim snapshots, and mid-scan updates.
|
||||
- [x] 4.3 Run focused and full regression suites with bytecode writes disabled and run strict OpenSpec validation.
|
||||
- [x] 4.4 Restart the authority-managed runtime and verify PostgreSQL, pipeline, source, queue, and recorder health.
|
||||
- [x] 4.5 Observe a one-per-cycle production canary and compare changed-target admission, completion, findings, and worker occupancy before increasing any cap.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-08-16
|
||||
@@ -0,0 +1,78 @@
|
||||
## Context
|
||||
|
||||
The scanner already persists exact and ambiguous routing hints for overlapping Qwen, DeepSeek, and Kimi `sk-...` findings. Provider workers, however, claim service-specific candidates and reject any hint that is not exactly their own service, so an ambiguous candidate can be repeatedly left unconsumed and quarantined. A ZAI detector already exists, but candidate extraction and a ZAI keychecker do not.
|
||||
|
||||
The authoritative runtime uses PostgreSQL candidate leases and transactional result completion. Compatibility JSONL and status files are projections and must not control routing.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Resolve one ambiguous credential sequentially across only its compatible providers.
|
||||
- Stop at the first response that proves the credential belongs to a provider.
|
||||
- Persist the successful result under the provider that recognized the credential.
|
||||
- Add direct ZAI extraction plus model-list authentication and a minimal generation/billing probe through the global and China APIs.
|
||||
- Preserve fenced candidate completion, capacity accounting, and restart safety.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Redesign credential storage or secret-retention policy.
|
||||
- Probe every supported provider for every unknown string.
|
||||
- Run provider probes for one ambiguous credential in parallel.
|
||||
- Automatically retry credentials whose current status is configured as terminal.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Use one virtual resolver candidate
|
||||
|
||||
Findings with a single strong provider hint continue to produce that provider's existing candidate. Findings with an ambiguous generic-key hint produce one `provider_resolver` candidate, deduplicated by credential within a staged scan bundle. The resolver owns one normal PostgreSQL lease and invokes compatible provider adapters in order, avoiding sibling candidates and cross-worker races.
|
||||
|
||||
This is preferred to enqueueing one active candidate per provider because the latter requires new coordination state, can spend quota concurrently, and complicates exact queue-capacity release.
|
||||
|
||||
### Keep route selection bounded and deterministic
|
||||
|
||||
Persisted provider evidence limits the compatible set. The default fallback order is `deepseek,zai,qwen,kimi`; a provider identified by the originating detector is moved to the front when it belongs to the compatible set. The order is configurable, deduplicated, and never expanded beyond the supported generic-key provider set.
|
||||
|
||||
Provider-specific formats such as `sk-sp-...`, `zai-...`, and ZAI's dotted key form remain direct routes when the finding evidence is unambiguous.
|
||||
|
||||
### Normalize adapter outcomes
|
||||
|
||||
Each adapter returns its existing detailed status plus a resolver outcome:
|
||||
|
||||
- `match`: a successful authenticated response or provider-specific account/quota response proves ownership.
|
||||
- `no_match`: the provider definitively rejects the credential as invalid.
|
||||
- `retry`: network, server, generic rate-limit, malformed, or otherwise inconclusive responses.
|
||||
|
||||
The resolver continues past `no_match` and may continue past `retry` to find a later positive match. If no provider matches, any retryable attempt keeps the result unresolved; only an all-`no_match` route is exhausted.
|
||||
|
||||
### Reassign a matched candidate during fenced completion
|
||||
|
||||
When a resolver result names a matched provider, `complete_keycheck_candidate` obtains or creates the canonical credential row for that provider, reassigns the leased candidate to it, and writes the result/current state under the matched service in the same transaction. The existing provider-key fingerprint, event fence, projection reservation, and capacity accounting remain unchanged. No schema migration is required.
|
||||
|
||||
Legacy Qwen, DeepSeek, or Kimi candidates carrying an ambiguous persisted hint delegate to the same resolver so explicitly retried old candidates do not return to the unconsumed quarantine loop.
|
||||
|
||||
### Authenticate ZAI through model listing and prove usability
|
||||
|
||||
The ZAI adapter first calls authenticated `GET /models` on `https://api.z.ai/api/paas/v4` and `https://open.bigmodel.cn/api/paas/v4`, then sends a one-token `POST /chat/completions` probe to the fixed `glm-5.2` target. A key is `VALID` only when that probe succeeds. Model-list authentication still proves provider ownership when the probe reports quota, balance, permission, model access, or transient failures, but those outcomes are persisted outside the alive set. HTTP, documented ZAI business codes, and bounded message markers distinguish invalid authentication, recognized quota/balance restrictions, rate limits, permission restrictions, and transient failures.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [A transient response from an early provider could hide a later match if probing stopped] -> Continue through the bounded compatible set while retaining the transient outcome if nobody matches.
|
||||
- [The same text could theoretically be valid at more than one compatible gateway] -> Deterministic first-match ordering is explicit and recorded with all preceding attempts.
|
||||
- [All-provider fallback increases requests for weak-context findings] -> Restrict it to detector-qualified generic-key formats and one sequential resolver candidate.
|
||||
- [Existing quarantined candidates are not silently mutated] -> Make legacy candidates resolver-aware; operators can explicitly retry affected quarantine records through the existing review path.
|
||||
- [Provider API behavior may change] -> Keep ZAI endpoints configurable and cover response classification with mocked regression tests.
|
||||
- [The usability probe consumes provider resources] -> Request one output token from one deterministic chat model and stop after the first conclusive authenticated endpoint.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Deploy extraction, resolver, ZAI adapter, runner registration, and transactional service reassignment together.
|
||||
2. Restart the supervised runtime so the lifecycle code manifest and service registry are rebuilt atomically.
|
||||
3. Verify new ambiguous candidates are owned by `provider_resolver` and matched rows are projected under the actual provider.
|
||||
4. Explicitly retry only relevant legacy provider-routing quarantine records after the new behavior is active.
|
||||
|
||||
Rollback requires stopping the runtime and restoring the previous code/config manifest. No database schema rollback is needed.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
Generic `sk-...` credentials can match several supported providers, while the current exact-hint routing leaves ambiguous findings unconsumed and can eventually quarantine them without testing a compatible provider. The keycheck pipeline needs ordered provider resolution and ZAI coverage so a credential is attributed to the first provider that positively recognizes it.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add durable, sequential resolution for credentials whose format or finding context permits multiple providers.
|
||||
- Distinguish provider mismatch from authenticated match and retryable probe failure.
|
||||
- Stop remaining provider attempts after the first positive match while retaining auditable attempt outcomes.
|
||||
- Add a ZAI provider checker with a minimal generation/billing probe and include ZAI in compatible generic-key routing.
|
||||
- Keep explicit single-provider findings on their existing direct validation path.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `ambiguous-provider-resolution`: Ordered, durable validation of one ambiguous credential across compatible providers until one positively matches or all definitive routes are exhausted.
|
||||
- `zai-key-validation`: Extraction, probing, classification, persistence, and runtime registration for ZAI API credentials.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects scanner provider hints, keycheck candidate extraction, PostgreSQL queue/schema operations, provider checker orchestration, status projection, runtime configuration, and lifecycle authority manifests.
|
||||
- Adds a ZAI checker module using the existing HTTP and PostgreSQL keycheck infrastructure; model-list authentication alone does not qualify a key as alive.
|
||||
- Requires regression coverage for route ordering, retry behavior, atomic match resolution, candidate deduplication, and ZAI response classification; the existing free-form service and credential tables require no schema migration.
|
||||
+52
@@ -0,0 +1,52 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Ambiguous credentials use one sequential resolver
|
||||
The system SHALL represent one ambiguous credential occurrence as one leased resolver candidate and SHALL probe only the compatible provider set in deterministic order.
|
||||
|
||||
#### Scenario: Weak generic-key context
|
||||
- **WHEN** a detector-qualified generic `sk-...` finding has no single strong provider attribution
|
||||
- **THEN** the system SHALL enqueue one resolver candidate rather than independently active candidates for every compatible provider
|
||||
|
||||
#### Scenario: Strong provider attribution
|
||||
- **WHEN** a finding has one strong provider-specific format or context signal
|
||||
- **THEN** the system SHALL retain the direct provider route without invoking unrelated provider adapters
|
||||
|
||||
### Requirement: Resolver outcomes control progression
|
||||
The resolver SHALL distinguish positive match, definitive provider mismatch, and retryable uncertainty.
|
||||
|
||||
#### Scenario: Provider rejects credential
|
||||
- **WHEN** a provider definitively reports invalid authentication for an ambiguous credential
|
||||
- **THEN** the resolver SHALL record that attempt and continue to the next compatible provider
|
||||
|
||||
#### Scenario: Provider response is inconclusive
|
||||
- **WHEN** a provider attempt fails because of a network error, server error, or otherwise inconclusive response
|
||||
- **THEN** the resolver SHALL NOT classify that attempt as a definitive provider mismatch
|
||||
|
||||
#### Scenario: All providers reject credential
|
||||
- **WHEN** every compatible provider definitively rejects the credential
|
||||
- **THEN** the resolver SHALL record an exhausted unresolved result after the final attempt
|
||||
|
||||
### Requirement: First positive match terminates resolution
|
||||
The resolver SHALL stop after the first response that proves the credential belongs to a provider.
|
||||
|
||||
#### Scenario: Later provider recognizes credential
|
||||
- **WHEN** earlier providers reject a credential and a later provider positively recognizes it
|
||||
- **THEN** the resolver SHALL stop without calling subsequent providers and SHALL retain the ordered attempt evidence
|
||||
|
||||
### Requirement: Matched result uses actual provider authority
|
||||
A positively resolved candidate SHALL be completed transactionally under the provider that recognized it.
|
||||
|
||||
#### Scenario: Resolver candidate matches another service
|
||||
- **WHEN** a leased resolver candidate receives a positive ZAI result
|
||||
- **THEN** the same fenced transaction SHALL associate the candidate and current state with the canonical ZAI credential and project the result as service `zai`
|
||||
|
||||
#### Scenario: Completion loses its lease fence
|
||||
- **WHEN** the candidate lease no longer matches during provider reassignment
|
||||
- **THEN** no result, current-state update, or partial service reassignment SHALL be committed
|
||||
|
||||
### Requirement: Legacy ambiguous candidates remain recoverable
|
||||
Existing generic-provider candidates with persisted ambiguous routing evidence SHALL use the resolver when explicitly retried.
|
||||
|
||||
#### Scenario: Retried legacy Qwen candidate
|
||||
- **WHEN** an old Qwen candidate carries an ambiguous generic-provider hint and is retried
|
||||
- **THEN** the Qwen worker SHALL delegate it to the shared resolver instead of leaving it unconsumed again
|
||||
@@ -0,0 +1,49 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: ZAI findings produce keycheck candidates
|
||||
The scanner SHALL create ZAI keycheck candidates for detector-qualified `zai-...`, compatible `sk-...`, and bounded dotted ZAI/Zhipu key forms.
|
||||
|
||||
#### Scenario: Existing ZaiGLM detector finding
|
||||
- **WHEN** the `ZaiGLM` custom detector emits a bounded API credential
|
||||
- **THEN** candidate extraction SHALL preserve its finding attribution and route it to ZAI or the ambiguous resolver according to persisted provider evidence
|
||||
|
||||
### Requirement: ZAI validation proves generation availability
|
||||
The ZAI checker SHALL authenticate with a model-list request and SHALL require a bounded one-token `glm-5.2` generation probe before classifying a credential as valid and alive.
|
||||
|
||||
#### Scenario: Global ZAI credential
|
||||
- **WHEN** the global ZAI `/models` endpoint accepts the credential and the bounded generation probe succeeds
|
||||
- **THEN** the checker SHALL record a valid authenticated ZAI result with bounded model and probe metadata
|
||||
|
||||
#### Scenario: China Zhipu credential
|
||||
- **WHEN** the global endpoint rejects a credential but the configured China endpoint accepts it and its bounded generation probe succeeds
|
||||
- **THEN** the checker SHALL record the credential as ZAI with the successful endpoint region
|
||||
|
||||
#### Scenario: Model listing succeeds but generation is unavailable
|
||||
- **WHEN** `/models` authenticates the credential but the generation probe reports quota, balance, permission, transient, or inconclusive failure
|
||||
- **THEN** the checker SHALL preserve the authenticated ZAI match but SHALL NOT classify the credential as valid or write it to the alive set
|
||||
|
||||
#### Scenario: Model listing omits GLM 5.2
|
||||
- **WHEN** `/models` authenticates the credential but does not advertise `glm-5.2`
|
||||
- **THEN** the checker SHALL still probe the fixed `glm-5.2` target and SHALL NOT substitute another model
|
||||
|
||||
### Requirement: ZAI responses are classified by protocol evidence
|
||||
The checker SHALL classify HTTP status and documented ZAI business error codes without treating inconclusive failures as invalid credentials.
|
||||
|
||||
#### Scenario: Authentication rejected
|
||||
- **WHEN** ZAI returns HTTP 401 or an authentication-failure business code
|
||||
- **THEN** the attempt SHALL be classified as a definitive provider mismatch or dead direct credential
|
||||
|
||||
#### Scenario: Authenticated balance or plan restriction
|
||||
- **WHEN** ZAI returns a provider-specific balance, usage-plan, or permission response during model listing or the generation probe
|
||||
- **THEN** the attempt SHALL be marked as belonging to ZAI with the corresponding limited, no-balance, or restricted status
|
||||
|
||||
#### Scenario: Network or server failure
|
||||
- **WHEN** the request fails in transit or ZAI returns a server error
|
||||
- **THEN** the checker SHALL classify the attempt as retryable rather than dead
|
||||
|
||||
### Requirement: ZAI participates in normal runtime accounting
|
||||
The ZAI checker SHALL use the existing PostgreSQL lease, result, current-state, projection, and summary infrastructure.
|
||||
|
||||
#### Scenario: Scheduled ZAI work exists
|
||||
- **WHEN** the unified keycheck scheduler detects claimable ZAI candidates
|
||||
- **THEN** it SHALL launch the ZAI checker with the same authority and bounded-slice controls used for other providers
|
||||
@@ -0,0 +1,27 @@
|
||||
## 1. Routing And Candidate Extraction
|
||||
|
||||
- [x] 1.1 Extend provider evidence and persisted hints to include ZAI and weak-context generic-key ambiguity.
|
||||
- [x] 1.2 Route ambiguous findings to one deduplicated `provider_resolver` candidate while preserving direct strong-provider candidates.
|
||||
- [x] 1.3 Add bounded ZAI key formats and `ZaiGLM` service extraction.
|
||||
|
||||
## 2. Resolver And ZAI Checker
|
||||
|
||||
- [x] 2.1 Implement shared deterministic provider resolution with match, no-match, and retry outcomes.
|
||||
- [x] 2.2 Add the PostgreSQL `provider_resolver` checker and legacy ambiguous-candidate delegation.
|
||||
- [x] 2.3 Add the ZAI `/models` authentication checker with global/China endpoint and business-code classification.
|
||||
- [x] 2.4 Require a bounded ZAI generation/billing probe before assigning `VALID`, while preserving authenticated non-alive outcomes.
|
||||
- [x] 2.5 Pin the ZAI usability probe to `glm-5.2` without model-list fallback.
|
||||
|
||||
## 3. Transactional Persistence And Runtime
|
||||
|
||||
- [x] 3.1 Reassign a positively matched resolver candidate to the canonical provider credential during fenced completion.
|
||||
- [x] 3.2 Register resolver and ZAI services, capabilities, status projections, configuration, and lifecycle runtime behavior.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Add extraction, route ordering, short-circuit, retry, ZAI classification, and service-reassignment regression tests.
|
||||
- [x] 4.2 Run targeted and existing keycheck/scanner test suites with bytecode writes disabled.
|
||||
- [x] 4.3 Restart the supervised runtime and verify READY status plus live resolver/ZAI queue behavior.
|
||||
- [x] 4.4 Add regression coverage for successful generation, no-balance, limited, restricted, and inconclusive ZAI probes.
|
||||
- [x] 4.5 Run the affected suites and verify the supervised runtime plus one live ZAI recheck.
|
||||
- [x] 4.6 Verify the fixed `glm-5.2` target with regression tests and one live recheck.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user