Files
2026-09-30 20:30:56 +03:00

22 KiB

Context

Docker scanning currently has two materially different paths. The full-image TruffleHog source is authoritative for new and normally completing images, but it treats an immutable image as one indivisible operation. The bounded layer path can resume and globally reuse content digests, but its fixed highest-eight-layer selector was admitted only for prior full-image timeouts. In the completed control cohort it retained 21.1% of routed identities and 5.8% of detector identities, so enabling that selector broadly would trade too much useful coverage for speed.

The layer path already provides the expensive safety primitives this change needs: exact manifest resolution, authenticated bounded Registry transfer, digest verification, private artifacts, contained TruffleHog filesystem execution, canonical reservation-bound plans, policy-scoped global blob leases, fenced result ingestion, and explicit per-image coverage. The missing pieces are a content-aware selector, a reusable execution-policy identity that is independent of selector budgets, and trustworthy evidence for deciding whether the adaptive path is safe to broaden.

Docker/OCI configuration contains an ordered history list that can help identify COPY, ADD, application setup, package installation, generic RUN, and likely bulk-data layers. That history is untrusted and may itself contain secret material. It is therefore only a bounded selection hint; it must never become an authority for content identity or successful coverage and its raw commands must never be persisted or logged.

The production database already contains durable version-one layer plans and covered blob rows. Changing their interpretation in place would invalidate audit and lease fences. The migration must be additive, retain version-one validation, and allow only explicitly proven compatible successful coverage to enter the new execution namespace.

Goals / Non-Goals

Goals:

  • Scan all supported unique content for images that fit conservative bounds.
  • Prioritize likely application, configuration, and source-bearing layers when a large image cannot fit those bounds.
  • Reuse successful immutable blob coverage across images and selector revisions when execution semantics are unchanged.
  • Freeze the first deterministic selection for an image and selector policy across every checkpoint, retry, and execution-policy transition.
  • Reduce per-image checkpoint overhead with bounded multi-blob leases without weakening per-blob execution and ingestion fences.
  • Preserve exact selected, reused, skipped, failed, and partial coverage semantics.
  • Produce private, aggregate, non-authoritative shadow evidence against 50-100 completed full-image controls before any broad adaptive rollout.
  • Retain full-image scanning and the existing timeout-only layer canary as immediate rollback paths.

Non-Goals:

  • Infer arbitrary file contents without downloading a compressed layer.
  • Claim complete coverage for an image with unsupported, oversized, failed, or budget-excluded descriptors.
  • Persist raw Docker config history, shadow findings, provider keys, target names, or Registry bearer tokens in rollout evidence.
  • Change detector classification, keycheck routing, Docker account ownership, repository resolver scheduling, guaranteed scan-slot capacity, worker count, or non-Docker scanners.
  • Reconstruct a merged container filesystem or remove historical whiteout content from individual layer evidence.
  • Destructively rewrite or delete version-one plans and coverage rows.

Decisions

1. Introduce a version-two immutable content plan

Adaptive execution uses a canonical version-two plan. It retains the version-one image, repository, manifest, platform, media, limits, scan-policy, descriptor, reservation, and plan-hash fences and adds:

  • a versioned selector algorithm and selection_policy_sha256;
  • an execution_policy_sha256 for reusable successful blob evidence;
  • one bounded classifier value per descriptor;
  • exact selection and omission reasons; and
  • bounded checkpoint lease limits that do not alter the frozen selection set.

Version-one plan and execution validators remain exact because their JSON is already durable. Version-two validation dispatches by the exact integer version and rejects unknown fields, unknown classes, malformed hashes, descriptor reordering, and oversized canonical JSON. A reservation can bind one canonical plan only; idempotent replay must be byte-identical.

After exact manifest resolution and before plan binding, the scanner fetches the configuration blob through the existing bounded, authenticated, digest-verifying Registry path. It parses the JSON in bounded private memory/storage and maps non-empty_layer history entries from base to top onto the ordered manifest layers. The optional rootfs.diff_ids count and the non-empty history count must agree with the layer count. A missing field, malformed value, excess entry count, or alignment mismatch classifies every layer as unknown; it does not change descriptor identity or fail a valid immutable manifest.

The classifier uses a small versioned allow-list and emits only these bounded classes: config, copy_add, app_config_run, package_run, other_run, bulk_data, and unknown. Raw created_by values are discarded before plan construction and are excluded from logs, result metadata, errors, and database rows.

Alternatives rejected:

  • Persisting normalized command text would retain unnecessary secret-bearing input.
  • Treating history as authoritative would let malformed or adversarial metadata hide content.
  • Mutating version-one plans would break exact replay and auditability.

2. Separate scan, execution, and selection identities

Three hashes have distinct responsibilities:

  • scan_policy_sha256 remains the installed scanner/detector/config fingerprint.
  • execution_policy_sha256 hashes the scan policy plus versioned content validation and all archive semantics that can change which bytes a successful command examines. It excludes image/layer selection budgets, classifier weights, retry counts, lease duration, checkpoint size, and delay.
  • selection_policy_sha256 hashes the selector version, class ordering, deterministic tie-breaks, supported descriptor classes, and all limits that determine the initial selected set. It excludes mutable global coverage and execution scheduling.

For version-two rows, the existing docker_content_blobs.coverage_policy_sha256 key stores the execution-policy hash. Selector changes therefore do not force an identical successfully scanned digest through TruffleHog again, while archive or detector semantic changes still create a separate coverage namespace.

docker_image_blob_coverage receives a non-null selection_policy_sha256. Existing rows are backfilled from the exact bound plan on their linked reservation and indexed by (queue_id, manifest_digest, selection_policy_sha256, position, reservation_id). The earliest full position map under that key is the immutable selection baseline. Coverage-policy changes may require new execution but cannot expand or contract that baseline silently.

Successfully covered version-one evidence may be copied lazily into the version-two execution namespace only in the same transaction that validates all of the following:

  • the old row is durably covered, not pending, leased, submitted, failed, or ambiguous;
  • a linked covered image row and reservation contain an exact valid version-one plan;
  • the descriptor digest, kind, declared bytes, and semantic media class match;
  • the old coverage key is exactly the legacy hash derived from that plan; and
  • the scan fingerprint and every scan-affecting archive semantic equal the requested version-two execution policy.

The alias keeps the original successful reservation, plan, byte count, and completion provenance. Legacy rows are never rekeyed or deleted. If any compatibility proof is absent, the new namespace starts uncovered and normal fenced execution is required.

Alternatives rejected:

  • Keeping selector limits in the coverage hash defeats global reuse whenever budgets are tuned.
  • Reusing every covered digest across scanner versions can suppress required rescans.
  • Bulk-rekeying legacy rows destroys provenance and races active leases.

3. Select all-fit images and rank large-image payload deterministically

Configuration remains independently eligible under its hard configuration-byte bound. Supported unique layer digests are evaluated under the hard per-layer bound. Covered digests in the matching execution namespace and duplicate positions in the same image are selected at zero new transfer bytes and zero new execution count.

If every supported unique descriptor fits the configured aggregate bytes and unique-layer count, the selector selects all of them regardless of history class. This is the complete bounded path for small images and avoids reducing their coverage merely because history hints are absent.

When the complete set does not fit, new unique layer candidates are sorted by:

  1. class priority: copy_add, app_config_run, package_run, unknown, other_run, bulk_data;
  2. highest manifest position first;
  3. smallest compressed descriptor first; and
  4. lexical digest as the final stable tie-break.

The selector greedily admits candidates while both aggregate compressed-byte and unique-layer-count limits permit them. Hard per-descriptor limits are never exceeded. Each descriptor records one exact reason, including selected class, already_covered, duplicate_digest, unsupported_media_type, config_too_large, layer_too_large, image_budget_exhausted, or layer_limit_exhausted. Changing any class order, classifier rule, supported-media rule, or selection bound changes the selector hash.

The selection algorithm receives a transactionally consistent coverage snapshot, but mutable coverage is not part of its identity. The first complete descriptor-position map is written before any new lease and reused exactly on later checkpoints. Consequently a layer skipped by the original budget never becomes newly selected merely because an earlier selected layer became globally covered.

Alternatives rejected:

  • Fixed highest-first selection has already failed the completed-control recall gate.
  • A whole-image byte cutoff loses small application layers above giant data layers.
  • Selecting only recognized commands lets missing or unusual history hide useful payload.
  • Selecting globally covered content only when it still fits the current budget wastes verified immutable evidence.

4. Lease bounded multi-blob checkpoints

The current executor and ingestion format already support more than one leased descriptor, but the binder leases one new digest and then defers the parent for 60 seconds. Adaptive execution leases a deterministic bounded batch from the frozen selected set under configurable maximum blob count and compressed bytes. A first eligible blob larger than the checkpoint-byte target but within its hard per-layer bound may be leased alone so it cannot starve indefinitely.

Every digest still has its own advisory lock, lease token, attempt count, execution record, digest verification, and final state. The executor processes the batch sequentially inside the same owned slot and bundle. Ingestion may cover successful earlier blobs while returning a later retryable blob to pending. A crash before durable handoff covers none of the un-ingested batch and normal exact lease expiry/recovery applies.

Checkpoint count, byte target, retry count, lease duration, and continuation delay are scheduling controls. They do not enter execution or selector hashes because they cannot turn an incomplete blob into successful coverage or change the frozen selected set.

5. Keep image coverage explicit and policy-specific

An image is complete only when its configuration and every manifest layer position are successfully covered under the requested execution policy. Reused and duplicate digests count as covered only after exact policy-compatible evidence exists. Any unsupported, oversized, budget-excluded, failed, or otherwise unselected descriptor makes the image bounded partial coverage.

Selected retryable work keeps the parent deferred. Shared active work does not consume another blob attempt. Exhausted selected work produces terminal incomplete disposition. Findings from completed selected blobs retain image, digest, kind, class, and position provenance and use the existing authoritative ingestion, projection, and keycheck paths.

Config history classification affects only selection order. It never changes detector output, finding authority, digest identity, or completion criteria.

6. Add adaptive modes without changing legacy rollout semantics

Existing full, timeout-only canary, and legacy layer meanings remain available for durable version-one work and rollback. Two explicit version-two modes are added:

  • adaptive-canary assigns a configured basis-point cohort across all immutable Docker manifests by a stable versioned hash; cohort members use adaptive plans and non-members use full-image scanning.
  • adaptive uses adaptive plans for every eligible immutable Docker claim.

Neither mode depends on a previous full-image timeout. Retry and checkpoint attempts for the same manifest and selector retain the same assignment. Invalid mode, policy, migration, or gate state fails closed to full-image execution before any adaptive plan is bound. Returning configuration to full changes only new claims and leaves adaptive plans and audit rows intact.

The existing timeout-only canary remains independent and may continue while adaptive shadow evidence is gathered. The broad legacy layer mode remains operationally disabled because its selector did not pass recall gates.

7. Gate rollout with non-authoritative aggregate shadow evidence

An operator-invoked bounded shadow evaluator selects 50-100 exact immutable images whose authoritative full-image scans completed successfully under one scan fingerprint. It executes the candidate adaptive policy using the same downloader, validators, process containment, deadlines, and scanner fingerprint, but under a shadow authority that cannot call normal result ingestion or mutate target status, result reservations, global blob coverage, findings, keycheck candidates, projections, or source counters.

For each paired control, full routed identities are read from the protected database as (service, provider_key_hash) and adaptive routed identities are derived in private memory through the same candidate normalization. Detector identities use detector_secret_hash. Identity sets, raw findings, commands, provider material, image names, and bearer tokens are discarded after intersection counts are computed.

The durable report contains only policy hashes, cohort and completion counts, aggregate full, adaptive, and intersection counts, aggregate slot milliseconds, bounded failure counts, threshold results, and timestamps. It records no per-image row or identity. Slot timing uses the same outer monotonic boundary from admitted work through durable shadow sink completion for both paths; scan subprocess duration remains a diagnostic, not the gate denominator.

A report passes only when:

  • 50-100 controls completed both paths without integrity, containment, or fence failure;
  • aggregate routed-identity recall, intersection / full, is at least 85%;
  • aggregate adaptive/full slot-time ratio is at most 40%;
  • every adaptive omission is represented in coverage counts; and
  • no credential persistence, quarantine, projection, source-failure, or resource-bound regression is observed.

Reports are bound to exact selector, execution, and scan policy hashes. Stale or incomplete reports cannot authorize another policy. Adaptive canary remains fail-closed until a matching report passes. Broad adaptive enablement additionally requires a stable low-percentage production canary over at least one repository-refresh interval. Operators change rollout configuration explicitly; shadow evidence never changes execution mode by itself.

Alternatives rejected:

  • Routing shadow candidates through keycheck would make the experiment authoritative and consume external capacity.
  • Persisting per-image shadow identities creates unnecessary sensitive correlation data.
  • Comparing only detector counts does not measure the routed identities the scanner is intended to produce.
  • Automatically enabling adaptive mode from a report removes the operational rollback checkpoint.

8. Preserve incomplete warning semantics without redundant retries

The first completed 50-control production shadow report failed closed. Routed recall was 12 of 18 identities (66.7%), adaptive/full slot time was 46.6%, and the report recorded 43 aggregate failures. Its selection evidence showed 230 descriptors omitted by the eight-layer limit and 14 oversized descriptors, so selector recall remains the primary rollout blocker.

The same evidence exposed a separate execution defect. TruffleHog diagnostics such as chunk_processing and detector_timeout are explicitly deterministic, non-retryable warnings. Their findings must be retained, but the affected blob cannot establish complete coverage. The diagnostic adapter previously discarded the non-retryable bit, causing the layer executor to download and scan the same incomplete blob up to three times before reaching the same terminal state. The adapter now preserves aggregate warning retryability and the layer executor terminates that blob after the first deterministic warning. It does not mark the blob covered or remove the report failure.

The historical report schema retained only a total failure count, so its 43 failures cannot be decomposed exactly after the fact. Future shadow runs keep a fixed allow-list of aggregate-only failure categories in protected memory and print their totals in the final operator summary without changing report authority or persisting target-level evidence.

Shadow execution also suppresses target labels in finding-filter logs. Production scans retain their existing target logging, while both private full and private layer paths emit only aggregate filter counts. This closes a privacy gap found in the first report log without weakening normal operational diagnostics.

Risks / Trade-offs

  • [History is malformed, misleading, or secret-bearing] -> Bound and validate it, persist only an enum, fall back to unknown, and keep identity/coverage independent of classification.
  • [Application secrets exist in a low-priority or giant layer] -> Scan all-fit images, keep unknown ahead of generic/bulk classes, record partial scope, retain full controls, and enforce the 85% routed-recall gate.
  • [Unsafe policy reuse suppresses a required rescan] -> Separate hashes and permit legacy aliasing only from exact successful compatible evidence in one fenced transaction.
  • [Selector changes expand coverage during retry] -> Freeze the earliest full position map under the selector hash before leases are issued.
  • [Multi-blob checkpoints increase work lost on crash] -> Bound count/bytes and retain independent per-blob leases and ingestion records; no pre-handoff result becomes covered.
  • [Config prefetch adds Registry traffic] -> Reuse the already bounded authenticated downloader and avoid a second fetch when the config is leased in the same plan.
  • [Individual layer scans expose whiteouted historical files] -> Preserve position provenance and describe evidence as image-content coverage, not merged-root state.
  • [Shadow evaluation consumes slots] -> Keep it operator-invoked, bounded, deterministic, and subject to the existing slot/resource controls.
  • [Aggregate reports hide individual anomalies] -> Fail the whole report on incomplete paired work and retain bounded failure counts without persisting target identity.
  • [Deterministic scanner warnings consume repeated transfer and slot time] -> Preserve their non-retryable policy, keep findings and incomplete coverage, and terminate the blob on its first bounded attempt.

Migration Plan

  1. Ship version-one and version-two validators, new modes, and migration code while production stays on its existing timeout-only canary.
  2. Stop authoritative runtime and verify no active result, queue, or blob leases remain.
  3. Add the selector-policy coverage column, aggregate shadow-report state, required timing/fence columns, indexes, and a new migration marker.
  4. Backfill every historical image-coverage row from its exact valid reservation plan. Abort and roll back the migration transaction on an orphan, malformed plan, or invalid hash; then make the column non-null and run exact schema validation.
  5. Restart with unchanged mode and verify legacy claims, projection, keycheck, resolver, quarantine, and source health before creating version-two work.
  6. Run the private shadow evaluator for 50-100 completed controls. Keep adaptive modes fail-closed if the matching report misses recall, timing, safety, or completion gates.
  7. Enable a low deterministic adaptive-canary, monitor at least one repository-refresh interval, and compare source failures, coverage reasons, routed yield, slot time, and quarantine.
  8. Increase canary basis points and finally enable adaptive only after every gate remains satisfied.
  9. Roll back immediately by setting mode to full or the existing timeout-only canary. Keep all version-two plans, reports, and coverage rows for audit and exact future resume.

Open Questions

  • Which initial bounded checkpoint count and byte target provide the best reduction in continuation delay without increasing crash rework materially?
  • Which classifier allow-list revisions improve routed recall in the first 50-100 controls? Every revision will receive a new selector-policy hash rather than changing an existing policy.
  • What adaptive-canary basis-point sequence should operators use after the shadow gate passes?