Initial server source import
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-01
|
||||
@@ -0,0 +1,254 @@
|
||||
## Context
|
||||
|
||||
DockerHub currently queues immutable `repository@sha256:<platform-manifest>` targets and gives each target to TruffleHog's Docker source as one indivisible operation. The process must fetch and inspect every layer before the external 600-second deadline. Large model images contain tens of compressed GiB, commonly dominated by one binary/model-weight layer, so two such targets can occupy both Docker workers while still producing only partial findings.
|
||||
|
||||
Measured over 48 hours, 196 deadline terminations consumed about 32.7 Docker worker-hours. Every timeout emitted findings, but all timeout findings routed to only 14 provider keys and none was usable. The queue currently resets target attempts after a timeout and the timeout disposition bypasses the maximum-attempt check, so the same immutable target can restart from byte zero indefinitely.
|
||||
|
||||
Registry manifest resolution already obtains the exact platform child manifest and ordered layer digests. The missing information is each descriptor's compressed size, media type, image configuration descriptor, durable per-layer coverage, and a scanner capable of processing one bounded content blob independently.
|
||||
|
||||
The runtime must retain its authenticated Docker pool, immutable digest identities, slot-first PostgreSQL admission, result-bundle fencing, Windows Job containment, output bounds, private temporary storage, and fail-closed behavior. Bearer tokens and provider keys must remain memory-only or in their existing protected stores.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- End unbounded timeout retries while preserving findings emitted before a deadline.
|
||||
- Scan image configuration and useful application layers without downloading giant model/data layers.
|
||||
- Resume image coverage after failure at layer granularity rather than restarting completed work.
|
||||
- Scan each immutable content digest once globally and reuse its durable coverage across images.
|
||||
- Keep download, disk, archive expansion, command execution, and database state bounded and fenced.
|
||||
- Report exact selected, covered, skipped, failed, and shared-pending scope for every image.
|
||||
- Prove the strategy against full-image control results before broad production enablement.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Reconstruct a runnable merged container root filesystem.
|
||||
- Claim complete image coverage when configured size or format bounds skip content.
|
||||
- Download giant layers in a separate long-running production lane in the initial change.
|
||||
- Implement arbitrary Registry hosts, private registries, cross-host credential forwarding, or resumable CDN downloads.
|
||||
- Change global guaranteed scan slots, Docker worker count, keycheck classification, or non-Docker scanners.
|
||||
- Delete the existing full-image implementation or schema on rollback.
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. Keep the image target, bind a fenced layer plan after claim
|
||||
|
||||
The existing immutable image target remains the queue and reporting identity. After slot-first admission claims an image, a dedicated PostgreSQL connection resolves its exact manifest and binds a canonical `docker_layer_plan` to that result reservation under the queue lease/event fences, analogous to exact Git plan binding.
|
||||
|
||||
The plan contains a version, repository, platform manifest digest, configuration descriptor, ordered layer descriptors, configured byte limits, selection reason for each descriptor, and a canonical SHA-256. The reservation stores the canonical JSON and hash. Replay is idempotent; changed replay or manifest mismatch is a conflict.
|
||||
|
||||
This avoids multiplying normal target-queue rows and keeps one authoritative parent disposition while still permitting durable child coverage.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- A separate full-image heavy lane still redownloads all content and cannot resume.
|
||||
- One target-queue row per layer complicates parent completion, finding attribution, and alternate repository fetch sources.
|
||||
- Encoding mutable coverage metadata into queue identity would break immutable image deduplication.
|
||||
|
||||
### 2. Add durable content and image-coverage tables
|
||||
|
||||
Add PostgreSQL-compatible tables through the additive runtime-safety migration:
|
||||
|
||||
- `docker_content_blobs`: one row per valid SHA-256 content digest, descriptor kind (`config` or `layer`), declared compressed bytes, media type, state, bounded attempts, active reservation/lease fence, completion metadata, and last bounded error.
|
||||
- `docker_image_blob_coverage`: one row per platform manifest and content digest with ordered position, selected/skipped reason, plan hash, and covered timestamp.
|
||||
|
||||
Successful content state is global because the digest authenticates the bytes. The current image repository is retained only as the bounded fetch source in the reservation plan. If one reservation owns a blob, another image records it as shared-pending and defers without duplicating the download. Expired/refunded reservation leases are reclaimable.
|
||||
|
||||
Blob completion changes only during ingestion of a matching reservation plan and matching per-blob execution record. Refund/recovery releases matching blob leases. A stale token cannot mark coverage or insert findings.
|
||||
|
||||
### 3. Select content by configurable byte budget, top layers first
|
||||
|
||||
The image configuration is always selected within a small independent hard bound because `Env`, labels and history are high-value and cheap.
|
||||
|
||||
Layer selection walks the manifest from highest layer to base layer. Already-covered digests require no byte budget. New layers are selected while both the per-layer compressed-byte cap and remaining per-image compressed-byte budget permit them. Non-selected descriptors are recorded with explicit reasons such as `layer_too_large`, `image_budget_exhausted`, or `unsupported_media_type`.
|
||||
|
||||
Initial numeric defaults are chosen only after the controlled spike. Configuration always has hard upper bounds; invalid values fail closed to full-image mode while the production gate is disabled.
|
||||
|
||||
This is a deliberate partial-coverage policy. It targets application/config layers and globally amortizes common base layers without pretending that skipped model weights were scanned.
|
||||
|
||||
Alternatives rejected:
|
||||
|
||||
- A whole-image size cutoff can discard a small valuable application layer sitting above giant weights.
|
||||
- Selecting oldest/base layers first spends budget on widely shared dependencies before application content.
|
||||
- Inferring binary content without downloading a compressed tar stream is not reliable.
|
||||
|
||||
### 4. Stream, verify, scan, and delete one blob at a time
|
||||
|
||||
For each blob leased by the reservation:
|
||||
|
||||
1. Obtain an account-scoped Registry bearer through the existing Docker account manager.
|
||||
2. Request the exact Registry blob with a bounded custom redirect policy. Redirects must remain HTTPS, must not contain userinfo, must reject local/private destinations, and must never receive the Registry Authorization header on another host.
|
||||
3. Stream into a private bounded work file while computing SHA-256. Reject excess bytes, short bodies, digest mismatch, unsupported media types, disk-floor violations, and response deadline exhaustion.
|
||||
4. Scan configuration JSON directly. Scan supported layer archives with TruffleHog `filesystem` under the existing OwnedProcess Job, timeout, output cap, configured detector policy, archive size/depth/time bounds, and normal process priority.
|
||||
5. Delete the blob work file before releasing the scan slot. No layer bearer, provider key, or raw result is written outside existing protected result artifacts.
|
||||
|
||||
Each layer runs as its own bounded TruffleHog command. This adds small startup overhead but gives exact completion, global deduplication, and restart from the first unfinished layer. Findings are enriched with immutable image, blob digest, kind, and layer position before normal bundle staging.
|
||||
|
||||
The controlled spike must prove that the installed TruffleHog build correctly scans supported real layer archive media types. Unsupported formats remain explicit uncovered scope rather than being silently accepted.
|
||||
|
||||
### 5. Parent disposition derives from explicit child execution
|
||||
|
||||
The result bundle carries the exact bound plan plus one execution record per claimed blob. Ingestion verifies plan hash and lease ownership before updating blob state.
|
||||
|
||||
- Successful blob command and digest verification marks that blob covered globally.
|
||||
- Incomplete blob execution preserves its findings but does not mark it covered.
|
||||
- Retryable blob failures release it to bounded retry; terminal failures remain explicit uncovered scope.
|
||||
- A blob active under another valid reservation causes a short parent deferral without charging a content attempt.
|
||||
- An image whose selected blobs are covered and whose remaining blobs are intentionally skipped completes with `coverage_complete=false` and detailed reasons.
|
||||
- An image with retryable selected work remains deferred; exhausted selected work becomes terminal failed/degraded according to the bound policy.
|
||||
|
||||
The existing full-image path remains available as a rollback/control path.
|
||||
|
||||
### 6. Fix timeout accounting before layer rollout
|
||||
|
||||
Both production-v2 and legacy completion paths stop resetting attempts for target-scoped timeouts. Timeout disposition checks the configured maximum before returning deferred. Source-wide infrastructure failures may retain their existing attempt-refund semantics.
|
||||
|
||||
A stopped-runtime repair reconciles only unfenced Docker queue rows whose latest durable result is a command timeout. It derives prior immutable-target attempt count from durable scans, sets the queue attempt count up to the configured maximum, and terminally closes already exhausted rows. It does not requeue or mutate active leases, reservations, findings, or successful targets.
|
||||
|
||||
### 7. Roll out through deterministic modes and evidence gates
|
||||
|
||||
Configuration exposes `full`, `canary`, and `layer` modes. Canary eligibility requires a durable previous full-image command timeout, and membership within that eligible set is a stable hash of immutable manifest digest. New images and normally completed controls therefore remain on the full path during canary rollout. Defaults remain `full` until migration and spike criteria pass.
|
||||
|
||||
The offline spike compares 20-50 timeout-heavy images and completed controls without printing findings or keys. It records bytes transferred, wall/slot time, peak resource use, distinct detector identities, and routed key identity recall. Production canary additionally tracks coverage reasons, blob reuse, timeout rate, keycheck candidate yield, and strict usable yield.
|
||||
|
||||
Broad enablement requires:
|
||||
|
||||
- no credential/token persistence regression;
|
||||
- no stale-fence or duplicate-blob completion;
|
||||
- exact digest verification for every covered blob;
|
||||
- material byte and slot-hour reduction on heavy images;
|
||||
- all routed key identities from completed control images retained, unless an explicitly reviewed coverage bound explains the difference;
|
||||
- no quarantine growth, projection regression, or source failure increase.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Secrets can exist in skipped giant layers] -> Record explicit incomplete scope, keep configurable budgets, compare controls, and retain full mode for targeted replay.
|
||||
- [Scanning individual layers can report files deleted by later whiteouts] -> Preserve layer provenance and treat this as historical image-content evidence rather than merged-root truth.
|
||||
- [Archive support differs by media type] -> Verify installed binary in the spike and mark unsupported formats uncovered.
|
||||
- [Global deduplication can be poisoned by stale completion] -> Require streamed digest verification plus reservation/lease/plan fences in the same ingestion transaction.
|
||||
- [Registry blob redirects introduce SSRF or credential-forwarding risk] -> Use a bounded validated redirect implementation and strip authorization across hosts.
|
||||
- [Per-layer process startup adds overhead for tiny layers] -> Batch measurement first; skip already-covered layers and permit a bounded future batching optimization only if needed.
|
||||
- [A crash can strand blob leases] -> Tie leases to result reservations, release on refund, and permit exact expiry reclamation.
|
||||
- [Database state grows with layer relationships] -> Enforce descriptor count bounds, compact metadata, indexed identities, and retention metrics.
|
||||
- [Layer mode can reduce broad secret coverage] -> Report coverage honestly and retain deterministic full-mode controls and rollback.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Ship and test bounded timeout accounting independently; repair exhausted historical timeout rows with sources stopped.
|
||||
2. Add nullable reservation plan columns, content/coverage tables, indexes, runtime validation, and additive migration marker. Keep mode `full`.
|
||||
3. Run a local controlled spike against retained immutable targets and select conservative byte/archive defaults from evidence.
|
||||
4. Deploy code with `full` mode, migrate offline, restart, and verify no behavior change.
|
||||
5. Enable deterministic low-percentage canary only for DockerHub images whose latest durable full-image result timed out. Monitor at least one full repository-refresh interval and sufficient heavy-image samples.
|
||||
6. Increase the timeout-fallback canary only after acceptance gates pass. Keep broad `layer` mode disabled until completed-control routed recall becomes adequate under revised bounds.
|
||||
7. Roll back by returning mode to `full`. Durable layer tables and nullable columns remain for audit and future resume; no destructive migration is required.
|
||||
|
||||
## Capability Spike Evidence
|
||||
|
||||
The installed `C:\Tools\trufflehog.exe` development build was exercised through 39 bounded
|
||||
`filesystem` invocations over synthetic direct tar, gzip-tar, zstd-tar, Docker outer-tar, and OCI
|
||||
outer-tar fixtures. All invocations exited successfully and all 24 expected synthetic findings were
|
||||
preserved. Extensionless gzip and zstd blobs were content-sniffed successfully.
|
||||
|
||||
The minimum archive depth was two for a direct tar, three for direct compressed layers, and four
|
||||
for an OCI/Docker archive containing a compressed layer. Shallower bounds produced an explicit
|
||||
non-fatal `max archive depth reached` diagnostic. `--archive-max-size` was proven to be a per-member
|
||||
bound rather than a cumulative compressed-image bound, so application-side per-blob and aggregate
|
||||
byte limits remain mandatory.
|
||||
|
||||
Initial conservative implementation bounds are therefore archive depth four, archive member size
|
||||
256 MiB, archive timeout 30 seconds, one-GiB aggregate selected compressed bytes per image, 256 MiB
|
||||
per selected layer, and filesystem concurrency two. These are implementation starting points, not
|
||||
broad-rollout acceptance evidence. The required aggregate timeout-heavy/completed-control image
|
||||
comparison remains an explicit gate before canary expansion.
|
||||
|
||||
The controlled aggregate comparison then ran against ten repeatedly timed-out immutable images and
|
||||
ten longest completed controls. No target names, findings, keys, or credentials were printed or
|
||||
persisted. Exact manifest resolution succeeded for all 20 images. The timeout-heavy manifests
|
||||
contained 341.68 GB of compressed descriptors; the bounded layer policy selected 944.02 MB (0.28%)
|
||||
and processed 84 blobs in 313.03 worker-seconds versus 6,015.31 historical full-image seconds
|
||||
(5.2%). All selected timeout-heavy blobs completed. The bounded path produced three routed
|
||||
credential identities, none overlapping the two identities in the historical incomplete full
|
||||
results, so it added useful scope while avoiding another byte-zero full-image retry.
|
||||
|
||||
The completed controls contained 98.05 GB; the policy selected 1.61 GB (1.64%) and processed 77
|
||||
blobs in 357.83 worker-seconds versus 5,956.38 historical seconds (6.0%). Five blobs reported
|
||||
bounded incomplete chunk processing. Only four of 19 historical routed identities were retained
|
||||
(21.1% recall), and distinct detector-identity recall was 19 of 326 (5.8%). System sampling across
|
||||
the run observed average CPU 20.27%, peak CPU 44.48%, minimum available physical memory 17.38 GB,
|
||||
and peak committed memory 30.00 GB; no resource-limit failure occurred.
|
||||
|
||||
This evidence accepts the existing 1 MiB config, 256 MiB per-layer, 1 GiB aggregate, eight-layer,
|
||||
archive-depth-four, 256 MiB archive-member, 30-second archive, 600-second blob, and filesystem
|
||||
concurrency-two bounds only for a deterministic timeout-fallback canary. It rejects broad random
|
||||
canary or broad layer mode because completed-control routed recall failed the acceptance gate. Full
|
||||
mode remains authoritative for new and normally completed images.
|
||||
|
||||
The post-refinement regression gate passed 329 related tests, including bounded transfer and content
|
||||
validation, parent/blob timeout accounting, PostgreSQL migration atomicity, policy-scoped global
|
||||
deduplication, stale fences, reclaim/quarantine, multi-checkpoint resume, and the durable full-timeout
|
||||
canary-eligibility transition. Strict OpenSpec validation passed, and no application bytecode was
|
||||
present after the run.
|
||||
|
||||
Rejected approaches from the capability spike are `--force-skip-archives` (it suppresses expected
|
||||
archive findings), relying on media-type labels without content verification, and relying on
|
||||
TruffleHog's `--archive-max-size` as an outer download or aggregate expansion bound. The exact
|
||||
development binary must remain fingerprinted by a capability contract because its reported version
|
||||
does not identify a stable release.
|
||||
|
||||
## Initial Production Canary Evidence
|
||||
|
||||
The additive migration and guarded historical timeout repair were applied offline before the
|
||||
timeout-only canary. Runtime first restarted in full mode with no layer rows, and the bounded canary
|
||||
was then enabled at 2,500 basis points only for images whose latest durable full-image result was a
|
||||
command timeout. New and normally completed images remained on the full scanner.
|
||||
|
||||
The first selected production image completed nine unique content checkpoints plus one final
|
||||
no-work completion checkpoint. The nine bounded blobs transferred 327,512 bytes and completed in
|
||||
33.176 seconds of aggregate parent duration, including 9.358 seconds of transfer and 23.496 seconds
|
||||
of contained filesystem scanning. All nine blobs reached policy-scoped global coverage, the parent
|
||||
finished `done`, no selected continuation remained, and quarantine stayed unchanged at 192 items /
|
||||
337,349,428 bytes. The prior full-image path for this eligibility class reached the 600-second
|
||||
deadline.
|
||||
|
||||
The initial run exposed and then verified a selection-continuity invariant: per-image byte/layer
|
||||
bounds must apply to the first immutable selection set, not be recomputed after each covered blob.
|
||||
The binder now reuses the earliest exact `(queue, manifest, coverage policy, position)` selection map
|
||||
on every later reservation. A PostgreSQL regression proves that layers skipped by the original
|
||||
count budget remain skipped after selected layers become globally covered. Production replay then
|
||||
completed without leasing content outside the original config-plus-eight-layer selection.
|
||||
|
||||
This single successful image proves the end-to-end checkpoint, resume, bounded selection and final
|
||||
completion paths, but is not enough evidence to increase the 25% timeout-only canary. Expansion
|
||||
still requires a longer observation window and more naturally eligible timeout samples.
|
||||
|
||||
The following overnight window added a second timeout-only image before a host reboot. Across both
|
||||
images, twelve unique blobs reached policy-scoped coverage with no retryable or terminal blob
|
||||
failure. The layer path used 53.192 seconds and transferred 79,614,382 bytes, compared with 1,201.330
|
||||
seconds consumed by the immediately preceding full-image timeout attempts. One parent completed;
|
||||
the second retained six exact selected checkpoints for durable resume. No layer findings or routed
|
||||
candidate identities were produced in this small sample, and quarantine remained unchanged.
|
||||
|
||||
Because this installation is an experimental rather than production service, the operator approved
|
||||
expanding the stable timeout-only cohort from 2,500 to 10,000 basis points. This does not enable broad
|
||||
layer mode: every new or normally completing image still uses the full scanner, and only an image
|
||||
with a durable prior full-image timeout may enter the bounded layer fallback. Broad layer mode
|
||||
remains rejected by the completed-control recall result. Task 7.5 remains open until the expanded
|
||||
cohort produces additional completed parents and a stable runtime observation window.
|
||||
|
||||
The expanded cohort exposed one additional metadata-normalization defect: identical content digests
|
||||
and sizes can be referenced through equivalent Docker and OCI media-type labels. Treating the label
|
||||
text itself as immutable metadata caused the Docker source to stop fail-closed before handoff. The
|
||||
binder now compares the validated semantic content class (`config-json`, `layer-tar`, `layer-gzip`,
|
||||
or `layer-zstd`) while still rejecting kind, byte-size, and compression-class conflicts. A real
|
||||
PostgreSQL concurrency regression and the related layer/runtime suites passed 126 tests. After the
|
||||
restart, 116 layer reservations were acknowledged in the first ten minutes, global covered blobs
|
||||
grew from 9 to 113, no blob entered a failed state, the source remained running, and quarantine was
|
||||
unchanged. The immediate digest queue remained intentionally thin (six due and eighteen delayed),
|
||||
while 12,679 repository anchors remained available to refill it after claimable digest work drains.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Which revised per-layer and per-image bounds can improve completed-control routed recall beyond the measured 21.1% without losing the measured slot-hour advantage?
|
||||
- Which OCI/Docker layer compression media types does the installed TruffleHog filesystem source handle reliably?
|
||||
- Is one TruffleHog process per selected layer sufficiently efficient, or is a later bounded multi-layer archive batch warranted?
|
||||
- What short deferral is appropriate when all remaining selected blobs are actively leased by other reservations?
|
||||
@@ -0,0 +1,31 @@
|
||||
## Why
|
||||
|
||||
Large Docker images currently monopolize both Docker scan workers until the 600-second deadline and can retry indefinitely because timeout completion resets the target attempt counter. Over the measured 48-hour window, hard timeouts consumed about 32.7 worker-hours while repeated partial scans produced no usable LLM access, so full-image retries are reducing useful throughput without providing proportional coverage.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Enforce the existing bounded target-attempt policy for Docker timeouts while preserving findings emitted before termination.
|
||||
- Resolve immutable image manifests into image configuration and ordered content-addressed layers with bounded size metadata.
|
||||
- Scan image configuration and selected layer content under an explicit per-image byte budget instead of treating every image as an indivisible download.
|
||||
- Deduplicate successful layer scans globally by immutable layer digest so shared base layers are not downloaded and scanned repeatedly.
|
||||
- Prioritize upper application layers and small layers; record oversized or out-of-budget layers as explicit uncovered scope rather than silently claiming complete image coverage.
|
||||
- Preserve the existing full-image path behind a rollout gate for controlled comparison and rollback.
|
||||
- Repair currently deferred Docker targets whose timeout attempts were incorrectly reset.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `docker-layer-content-scanning`: Bounded, content-addressed Docker config and layer scanning with global deduplication, explicit coverage, safe retry limits, and controlled rollout against the existing full-image scanner.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects Docker Registry manifest/blob access, immutable Docker target planning, scan queue state, result metadata, and Docker source configuration.
|
||||
- Adds durable PostgreSQL state for layer identities, leases, coverage, attempts, and image-to-layer plans.
|
||||
- Reuses the existing authenticated Docker account pool, scan-slot limiter, Windows Job containment, bundle ingestion, findings projection, and keycheck pipeline.
|
||||
- Requires an offline additive runtime-safety migration before enabling production layer scanning.
|
||||
- Does not change Git, Hugging Face, keycheck classification, global guaranteed scan-slot capacity, or secret persistence boundaries.
|
||||
+190
@@ -0,0 +1,190 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Docker timeout retries are bounded
|
||||
The system SHALL count target-scoped Docker command timeouts against the configured target-attempt maximum and SHALL preserve partial findings without creating an unbounded retry loop.
|
||||
|
||||
#### Scenario: Timeout before attempt limit
|
||||
- **WHEN** a Docker command times out before the configured maximum attempt
|
||||
- **THEN** its emitted findings remain durable and the immutable target is deferred using the configured timeout delay without resetting its attempt count
|
||||
|
||||
#### Scenario: Timeout reaches attempt limit
|
||||
- **WHEN** a Docker command times out at the configured maximum attempt
|
||||
- **THEN** its emitted findings remain durable and the queue records a terminal target disposition
|
||||
|
||||
#### Scenario: Source-wide outage
|
||||
- **WHEN** Docker execution is prevented by a source-wide infrastructure failure rather than target-scoped work
|
||||
- **THEN** the existing fenced source-failure recovery policy remains applicable
|
||||
|
||||
### Requirement: Immutable image content plans are bound after claim
|
||||
The system SHALL resolve a claimed Docker image's exact platform manifest into a canonical bounded plan containing its configuration and ordered layer descriptors, and SHALL bind that plan to the active result reservation before content execution.
|
||||
|
||||
#### Scenario: Valid immutable manifest
|
||||
- **WHEN** the claimed `repository@sha256:<digest>` resolves to a valid requested-platform manifest
|
||||
- **THEN** the bound plan identifies the same manifest digest and contains only bounded valid SHA-256 content descriptors, sizes, media types, and order
|
||||
|
||||
#### Scenario: Changed replay
|
||||
- **WHEN** the same reservation attempts to bind a different content plan
|
||||
- **THEN** the system rejects the replay as a fenced conflict and executes neither plan
|
||||
|
||||
#### Scenario: Invalid or oversized manifest
|
||||
- **WHEN** the Registry manifest is malformed, exceeds descriptor bounds, or disagrees with the immutable target
|
||||
- **THEN** the system fails closed without claiming complete content coverage
|
||||
|
||||
### Requirement: Content selection is bounded and application-first
|
||||
The system SHALL always select bounded image configuration and SHALL select new layers from highest to lowest under configured per-layer and per-image compressed-byte limits.
|
||||
|
||||
#### Scenario: Giant base or model layer
|
||||
- **WHEN** a layer exceeds the configured per-layer limit
|
||||
- **THEN** the layer is not downloaded by the normal layer scanner and coverage records `layer_too_large`
|
||||
|
||||
#### Scenario: Image byte budget is exhausted
|
||||
- **WHEN** another unscanned layer would exceed the remaining per-image budget
|
||||
- **THEN** the layer remains unselected and coverage records `image_budget_exhausted`
|
||||
|
||||
#### Scenario: Shared layer is already covered
|
||||
- **WHEN** a layer digest has successful global coverage
|
||||
- **THEN** the image reuses that coverage without consuming its transfer budget or launching another scan
|
||||
|
||||
#### Scenario: Upper and base layers both fit
|
||||
- **WHEN** multiple unscanned layers fit within all configured bounds
|
||||
- **THEN** the system selects them in highest-to-lowest manifest order
|
||||
|
||||
### Requirement: Layer coverage is globally deduplicated and fenced
|
||||
The system SHALL maintain one authoritative content-scan state per immutable digest and SHALL change successful coverage only through matching reservation, lease, plan, and ingestion fences.
|
||||
|
||||
#### Scenario: Concurrent images share a layer
|
||||
- **WHEN** two image plans reference the same unscanned digest concurrently
|
||||
- **THEN** at most one reservation owns its active scan and the other image records shared pending work without duplicate execution
|
||||
|
||||
#### Scenario: Successful matching ingestion
|
||||
- **WHEN** a result bundle contains a successful execution for a blob leased by its exact bound plan
|
||||
- **THEN** ingestion marks that digest globally covered in the same durable transaction
|
||||
|
||||
#### Scenario: Stale completion
|
||||
- **WHEN** a bundle or worker presents an expired, refunded, or mismatched blob lease
|
||||
- **THEN** it cannot mark the digest covered or advance image coverage
|
||||
|
||||
#### Scenario: Reservation is refunded
|
||||
- **WHEN** an image reservation is durably refunded before handoff
|
||||
- **THEN** only blob leases owned by that reservation are released for bounded reclamation
|
||||
|
||||
### Requirement: Registry blob transfer is authenticated, bounded, and verified
|
||||
The system SHALL fetch selected content from the trusted Docker Registry using the existing account pool, bounded streaming, private storage, safe redirect handling, and exact digest verification.
|
||||
|
||||
#### Scenario: Valid content download
|
||||
- **WHEN** the Registry returns exactly the declared bounded blob bytes whose SHA-256 matches the descriptor
|
||||
- **THEN** the private work artifact becomes eligible for scanning
|
||||
|
||||
#### Scenario: Cross-host redirect
|
||||
- **WHEN** the trusted Registry redirects a blob request to an allowed public HTTPS content host
|
||||
- **THEN** the system follows only the bounded validated redirect and does not forward Registry authorization to the other host
|
||||
|
||||
#### Scenario: Unsafe redirect
|
||||
- **WHEN** a blob redirect uses HTTP, userinfo, a local/private destination, or exceeds redirect bounds
|
||||
- **THEN** the transfer fails closed without exposing authentication material
|
||||
|
||||
#### Scenario: Size or digest mismatch
|
||||
- **WHEN** streamed bytes exceed bounds, end short, or do not match the expected digest
|
||||
- **THEN** the system deletes the work artifact and does not record successful coverage
|
||||
|
||||
#### Scenario: Insufficient private storage
|
||||
- **WHEN** the configured private work volume cannot retain its required free-space floor
|
||||
- **THEN** no blob download begins and the failure receives bounded retry disposition
|
||||
|
||||
### Requirement: Configuration and layers are scanned independently
|
||||
The system SHALL scan bounded image configuration and each newly leased supported layer as independent contained commands while preserving image and layer provenance on findings.
|
||||
|
||||
#### Scenario: Configuration contains candidate material
|
||||
- **WHEN** bounded configuration JSON contains detector-matching data
|
||||
- **THEN** findings identify the image and configuration digest and enter the normal result and keycheck pipeline
|
||||
|
||||
#### Scenario: Supported layer completes
|
||||
- **WHEN** TruffleHog filesystem scanning of a verified layer archive completes successfully
|
||||
- **THEN** its findings retain image, layer digest, kind, and position provenance and the layer becomes globally covered after fenced ingestion
|
||||
|
||||
#### Scenario: Layer scan is incomplete
|
||||
- **WHEN** a layer command times out or exits without confirmed completion
|
||||
- **THEN** emitted findings remain durable but that digest does not become covered
|
||||
|
||||
#### Scenario: Unsupported media type
|
||||
- **WHEN** a layer compression or media type is not supported by the validated scanner path
|
||||
- **THEN** no unsafe fallback executes and image coverage records `unsupported_media_type`
|
||||
|
||||
### Requirement: Image coverage is explicit and honest
|
||||
The system SHALL persist and expose selected, covered, shared-pending, failed, and intentionally skipped content for each immutable image plan.
|
||||
|
||||
#### Scenario: Every descriptor is covered
|
||||
- **WHEN** configuration and all image layers have successful global coverage
|
||||
- **THEN** the image records complete content coverage
|
||||
|
||||
#### Scenario: Bounds skip content
|
||||
- **WHEN** one or more descriptors are excluded by configured size, budget, or format bounds
|
||||
- **THEN** the image may finish as bounded partial coverage but SHALL NOT report complete content coverage
|
||||
|
||||
#### Scenario: Selected content remains retryable
|
||||
- **WHEN** at least one selected blob failed retryably or is actively covered by another reservation
|
||||
- **THEN** the image remains deferred without claiming complete coverage
|
||||
|
||||
#### Scenario: Selected content exhausts retries
|
||||
- **WHEN** required selected content reaches its terminal attempt limit
|
||||
- **THEN** the image receives terminal incomplete disposition with durable coverage detail
|
||||
|
||||
### Requirement: Layer work resumes without repeating completed content
|
||||
The system SHALL resume an incomplete image from selected content that lacks successful global coverage and SHALL NOT relaunch completed content digests.
|
||||
|
||||
#### Scenario: Parent image retries
|
||||
- **WHEN** an image retry follows partial layer completion
|
||||
- **THEN** the new plan reuses covered digests and leases only remaining eligible content
|
||||
|
||||
#### Scenario: Process crashes after one layer
|
||||
- **WHEN** one layer was durably ingested before a later layer or parent process failed
|
||||
- **THEN** recovery preserves the completed layer and reclaims only unfinished leased content
|
||||
|
||||
### Requirement: Full-image compatibility and deterministic canary are retained
|
||||
The system SHALL retain the existing full-image scanner behind configuration. Canary layer execution SHALL require a durable previous full-image command timeout and SHALL be selected deterministically from immutable manifest identity within that eligible set.
|
||||
|
||||
#### Scenario: Full mode
|
||||
- **WHEN** Docker layer mode is disabled or set to `full`
|
||||
- **THEN** the existing immutable full-image execution path remains authoritative
|
||||
|
||||
#### Scenario: Canary retry
|
||||
- **WHEN** a canary image is retried
|
||||
- **THEN** its immutable manifest digest selects the same scanner mode as its previous attempt
|
||||
|
||||
#### Scenario: Non-timeout image during canary rollout
|
||||
- **WHEN** an image has no durable previous full-image command timeout
|
||||
- **THEN** canary configuration keeps that image on the full-image execution path
|
||||
|
||||
#### Scenario: Layer mode rollback
|
||||
- **WHEN** operators return configuration from `layer` or `canary` to `full`
|
||||
- **THEN** new claims use full-image execution without deleting durable layer audit state
|
||||
|
||||
### Requirement: Controlled evidence gates production rollout
|
||||
The system SHALL compare layer scanning with completed full-image controls and SHALL keep broad production layer mode disabled until security, coverage, and throughput gates pass.
|
||||
|
||||
#### Scenario: Controlled spike
|
||||
- **WHEN** the offline spike runs against bounded heavy and completed-control samples
|
||||
- **THEN** it records aggregate bytes, wall time, slot time, coverage, distinct detector identities, and routed key recall without printing findings or keys
|
||||
|
||||
#### Scenario: Acceptance criteria fail
|
||||
- **WHEN** digest integrity, fence safety, routed-key recall, resource bounds, or throughput criteria fail
|
||||
- **THEN** production remains in full mode
|
||||
|
||||
#### Scenario: Completed-control recall fails but timeout fallback passes
|
||||
- **WHEN** bounded layer scanning materially reduces timeout-heavy work but does not retain completed-control routed-key recall
|
||||
- **THEN** operators may canary only prior full-image timeout retries and SHALL NOT enable broad layer mode
|
||||
|
||||
#### Scenario: Acceptance criteria pass
|
||||
- **WHEN** the controlled spike and deterministic production canary satisfy all defined gates
|
||||
- **THEN** operators may increase canary coverage or enable layer mode through configuration
|
||||
|
||||
### Requirement: Historical timeout state is repaired safely
|
||||
The system SHALL reconcile incorrectly reset Docker timeout attempts only while sources are stopped and only for unfenced immutable targets backed by durable timeout results.
|
||||
|
||||
#### Scenario: Exhausted historical timeout target
|
||||
- **WHEN** an unfenced deferred Docker target has durable timeout executions at or above the configured maximum
|
||||
- **THEN** repair marks it terminal without deleting its existing findings
|
||||
|
||||
#### Scenario: Active or ambiguous target
|
||||
- **WHEN** a Docker target has an active lease, reservation, event fence, or ambiguous latest result
|
||||
- **THEN** repair leaves it unchanged
|
||||
@@ -0,0 +1,45 @@
|
||||
## 1. Restore Bounded Timeout Semantics
|
||||
|
||||
- [x] 1.1 Make Docker timeout disposition terminal at the configured target-attempt maximum in production-v2 and legacy paths without resetting attempts
|
||||
- [x] 1.2 Add regression tests proving partial findings survive and timeout attempts stop at the configured limit
|
||||
- [x] 1.3 Add a stopped-runtime guarded repair for unfenced historical Docker targets whose durable timeout attempts were reset
|
||||
|
||||
## 2. Validate Layer Scanning
|
||||
|
||||
- [x] 2.1 Prove the installed TruffleHog filesystem source safely scans bounded Docker gzip and supported OCI layer archives with preserved findings
|
||||
- [x] 2.2 Run an aggregate-only spike across timeout-heavy and completed-control images and select conservative config, layer, image, archive, and deadline bounds
|
||||
- [x] 2.3 Record spike acceptance evidence and rejected formats/approaches in the design
|
||||
|
||||
## 3. Add Durable Layer State
|
||||
|
||||
- [x] 3.1 Add reservation plan columns plus Docker content-blob and image-coverage tables, indexes, migration marker, and exact runtime validation
|
||||
- [x] 3.2 Implement bounded canonical Docker layer-plan validation, hashing, idempotent fenced binding, and deterministic canary selection
|
||||
- [x] 3.3 Implement content lease claim, expiry, refund, retry, terminal failure, and globally successful coverage transitions
|
||||
- [x] 3.4 Wire matching blob execution and image coverage updates into the authoritative fenced result-ingestion transaction
|
||||
|
||||
## 4. Implement Bounded Registry Content Access
|
||||
|
||||
- [x] 4.1 Extend exact platform manifest resolution with bounded configuration and ordered layer size/media descriptors
|
||||
- [x] 4.2 Implement authenticated Registry blob streaming with safe redirects, byte/disk/deadline bounds, private artifacts, and SHA-256 verification
|
||||
- [x] 4.3 Implement contained configuration and layer archive scans with immutable image/blob provenance and deterministic cleanup
|
||||
|
||||
## 5. Integrate Layer-Aware Execution
|
||||
|
||||
- [x] 5.1 Select configuration and highest-first layers under per-layer and per-image byte budgets while reusing globally covered digests
|
||||
- [x] 5.2 Resolve and bind the Docker layer plan after the fenced parent claim using a dedicated database connection
|
||||
- [x] 5.3 Execute only leased blobs, preserve partial findings, and emit exact plan/execution/coverage metadata through result bundles
|
||||
- [x] 5.4 Resume deferred images without rerunning covered blobs and defer shared active content without charging duplicate attempts
|
||||
|
||||
## 6. Add Controlled Rollout and Observability
|
||||
|
||||
- [x] 6.1 Add bounded `full`, deterministic `canary`, and `layer` configuration with full-image rollback
|
||||
- [x] 6.2 Expose aggregate selected/covered/skipped/shared/failed bytes, blob reuse, transfer duration, timeout, and image coverage metrics without secret material
|
||||
- [x] 6.3 Keep existing scan-slot, Windows Job, output, keycheck, and credential-isolation invariants unchanged
|
||||
|
||||
## 7. Verify and Deploy
|
||||
|
||||
- [x] 7.1 Add unit tests for plan bounds, selection order, downloader security, digest verification, provenance, timeout policy, and canary stability
|
||||
- [x] 7.2 Add PostgreSQL integration tests for concurrent global deduplication, stale fences, refund/reclaim, partial ingestion, resume, and image coverage
|
||||
- [x] 7.3 Run related regression suites, strict OpenSpec validation, and aggregate control comparison; document evidence
|
||||
- [x] 7.4 Apply the additive migration offline, repair historical attempts, restart in full mode, and verify no behavior regression
|
||||
- [ ] 7.5 Enable a bounded deterministic canary, monitor throughput/coverage/keycheck/quarantine gates, and expand only if acceptance criteria pass
|
||||
Reference in New Issue
Block a user