Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,352 @@
## Context
Docker scanning currently has two materially different paths. The full-image TruffleHog source is
authoritative for new and normally completing images, but it treats an immutable image as one
indivisible operation. The bounded layer path can resume and globally reuse content digests, but its
fixed highest-eight-layer selector was admitted only for prior full-image timeouts. In the completed
control cohort it retained 21.1% of routed identities and 5.8% of detector identities, so enabling
that selector broadly would trade too much useful coverage for speed.
The layer path already provides the expensive safety primitives this change needs: exact manifest
resolution, authenticated bounded Registry transfer, digest verification, private artifacts,
contained TruffleHog filesystem execution, canonical reservation-bound plans, policy-scoped global
blob leases, fenced result ingestion, and explicit per-image coverage. The missing pieces are a
content-aware selector, a reusable execution-policy identity that is independent of selector
budgets, and trustworthy evidence for deciding whether the adaptive path is safe to broaden.
Docker/OCI configuration contains an ordered `history` list that can help identify `COPY`, `ADD`,
application setup, package installation, generic `RUN`, and likely bulk-data layers. That history is
untrusted and may itself contain secret material. It is therefore only a bounded selection hint; it
must never become an authority for content identity or successful coverage and its raw commands
must never be persisted or logged.
The production database already contains durable version-one layer plans and covered blob rows.
Changing their interpretation in place would invalidate audit and lease fences. The migration must
be additive, retain version-one validation, and allow only explicitly proven compatible successful
coverage to enter the new execution namespace.
## Goals / Non-Goals
**Goals:**
- Scan all supported unique content for images that fit conservative bounds.
- Prioritize likely application, configuration, and source-bearing layers when a large image cannot
fit those bounds.
- Reuse successful immutable blob coverage across images and selector revisions when execution
semantics are unchanged.
- Freeze the first deterministic selection for an image and selector policy across every checkpoint,
retry, and execution-policy transition.
- Reduce per-image checkpoint overhead with bounded multi-blob leases without weakening per-blob
execution and ingestion fences.
- Preserve exact selected, reused, skipped, failed, and partial coverage semantics.
- Produce private, aggregate, non-authoritative shadow evidence against 50-100 completed full-image
controls before any broad adaptive rollout.
- Retain full-image scanning and the existing timeout-only layer canary as immediate rollback paths.
**Non-Goals:**
- Infer arbitrary file contents without downloading a compressed layer.
- Claim complete coverage for an image with unsupported, oversized, failed, or budget-excluded
descriptors.
- Persist raw Docker config history, shadow findings, provider keys, target names, or Registry bearer
tokens in rollout evidence.
- Change detector classification, keycheck routing, Docker account ownership, repository resolver
scheduling, guaranteed scan-slot capacity, worker count, or non-Docker scanners.
- Reconstruct a merged container filesystem or remove historical whiteout content from individual
layer evidence.
- Destructively rewrite or delete version-one plans and coverage rows.
## Decisions
### 1. Introduce a version-two immutable content plan
Adaptive execution uses a canonical version-two plan. It retains the version-one image, repository,
manifest, platform, media, limits, scan-policy, descriptor, reservation, and plan-hash fences and
adds:
- a versioned selector algorithm and `selection_policy_sha256`;
- an `execution_policy_sha256` for reusable successful blob evidence;
- one bounded classifier value per descriptor;
- exact selection and omission reasons; and
- bounded checkpoint lease limits that do not alter the frozen selection set.
Version-one plan and execution validators remain exact because their JSON is already durable.
Version-two validation dispatches by the exact integer version and rejects unknown fields, unknown
classes, malformed hashes, descriptor reordering, and oversized canonical JSON. A reservation can
bind one canonical plan only; idempotent replay must be byte-identical.
After exact manifest resolution and before plan binding, the scanner fetches the configuration blob
through the existing bounded, authenticated, digest-verifying Registry path. It parses the JSON in
bounded private memory/storage and maps non-`empty_layer` history entries from base to top onto the
ordered manifest layers. The optional `rootfs.diff_ids` count and the non-empty history count must
agree with the layer count. A missing field, malformed value, excess entry count, or alignment
mismatch classifies every layer as `unknown`; it does not change descriptor identity or fail a
valid immutable manifest.
The classifier uses a small versioned allow-list and emits only these bounded classes:
`config`, `copy_add`, `app_config_run`, `package_run`, `other_run`, `bulk_data`, and `unknown`.
Raw `created_by` values are discarded before plan construction and are excluded from logs, result
metadata, errors, and database rows.
Alternatives rejected:
- Persisting normalized command text would retain unnecessary secret-bearing input.
- Treating history as authoritative would let malformed or adversarial metadata hide content.
- Mutating version-one plans would break exact replay and auditability.
### 2. Separate scan, execution, and selection identities
Three hashes have distinct responsibilities:
- `scan_policy_sha256` remains the installed scanner/detector/config fingerprint.
- `execution_policy_sha256` hashes the scan policy plus versioned content validation and all archive
semantics that can change which bytes a successful command examines. It excludes image/layer
selection budgets, classifier weights, retry counts, lease duration, checkpoint size, and delay.
- `selection_policy_sha256` hashes the selector version, class ordering, deterministic tie-breaks,
supported descriptor classes, and all limits that determine the initial selected set. It excludes
mutable global coverage and execution scheduling.
For version-two rows, the existing `docker_content_blobs.coverage_policy_sha256` key stores the
execution-policy hash. Selector changes therefore do not force an identical successfully scanned
digest through TruffleHog again, while archive or detector semantic changes still create a separate
coverage namespace.
`docker_image_blob_coverage` receives a non-null `selection_policy_sha256`. Existing rows are
backfilled from the exact bound plan on their linked reservation and indexed by
`(queue_id, manifest_digest, selection_policy_sha256, position, reservation_id)`. The earliest full
position map under that key is the immutable selection baseline. Coverage-policy changes may require
new execution but cannot expand or contract that baseline silently.
Successfully covered version-one evidence may be copied lazily into the version-two execution
namespace only in the same transaction that validates all of the following:
- the old row is durably `covered`, not pending, leased, submitted, failed, or ambiguous;
- a linked covered image row and reservation contain an exact valid version-one plan;
- the descriptor digest, kind, declared bytes, and semantic media class match;
- the old coverage key is exactly the legacy hash derived from that plan; and
- the scan fingerprint and every scan-affecting archive semantic equal the requested version-two
execution policy.
The alias keeps the original successful reservation, plan, byte count, and completion provenance.
Legacy rows are never rekeyed or deleted. If any compatibility proof is absent, the new namespace
starts uncovered and normal fenced execution is required.
Alternatives rejected:
- Keeping selector limits in the coverage hash defeats global reuse whenever budgets are tuned.
- Reusing every covered digest across scanner versions can suppress required rescans.
- Bulk-rekeying legacy rows destroys provenance and races active leases.
### 3. Select all-fit images and rank large-image payload deterministically
Configuration remains independently eligible under its hard configuration-byte bound. Supported
unique layer digests are evaluated under the hard per-layer bound. Covered digests in the matching
execution namespace and duplicate positions in the same image are selected at zero new transfer
bytes and zero new execution count.
If every supported unique descriptor fits the configured aggregate bytes and unique-layer count,
the selector selects all of them regardless of history class. This is the complete bounded path for
small images and avoids reducing their coverage merely because history hints are absent.
When the complete set does not fit, new unique layer candidates are sorted by:
1. class priority: `copy_add`, `app_config_run`, `package_run`, `unknown`, `other_run`, `bulk_data`;
2. highest manifest position first;
3. smallest compressed descriptor first; and
4. lexical digest as the final stable tie-break.
The selector greedily admits candidates while both aggregate compressed-byte and unique-layer-count
limits permit them. Hard per-descriptor limits are never exceeded. Each descriptor records one exact
reason, including selected class, `already_covered`, `duplicate_digest`, `unsupported_media_type`,
`config_too_large`, `layer_too_large`, `image_budget_exhausted`, or `layer_limit_exhausted`.
Changing any class order, classifier rule, supported-media rule, or selection bound changes the
selector hash.
The selection algorithm receives a transactionally consistent coverage snapshot, but mutable
coverage is not part of its identity. The first complete descriptor-position map is written before
any new lease and reused exactly on later checkpoints. Consequently a layer skipped by the original
budget never becomes newly selected merely because an earlier selected layer became globally
covered.
Alternatives rejected:
- Fixed highest-first selection has already failed the completed-control recall gate.
- A whole-image byte cutoff loses small application layers above giant data layers.
- Selecting only recognized commands lets missing or unusual history hide useful payload.
- Selecting globally covered content only when it still fits the current budget wastes verified
immutable evidence.
### 4. Lease bounded multi-blob checkpoints
The current executor and ingestion format already support more than one leased descriptor, but the
binder leases one new digest and then defers the parent for 60 seconds. Adaptive execution leases a
deterministic bounded batch from the frozen selected set under configurable maximum blob count and
compressed bytes. A first eligible blob larger than the checkpoint-byte target but within its hard
per-layer bound may be leased alone so it cannot starve indefinitely.
Every digest still has its own advisory lock, lease token, attempt count, execution record, digest
verification, and final state. The executor processes the batch sequentially inside the same owned
slot and bundle. Ingestion may cover successful earlier blobs while returning a later retryable blob
to pending. A crash before durable handoff covers none of the un-ingested batch and normal exact
lease expiry/recovery applies.
Checkpoint count, byte target, retry count, lease duration, and continuation delay are scheduling
controls. They do not enter execution or selector hashes because they cannot turn an incomplete blob
into successful coverage or change the frozen selected set.
### 5. Keep image coverage explicit and policy-specific
An image is complete only when its configuration and every manifest layer position are successfully
covered under the requested execution policy. Reused and duplicate digests count as covered only
after exact policy-compatible evidence exists. Any unsupported, oversized, budget-excluded, failed,
or otherwise unselected descriptor makes the image bounded partial coverage.
Selected retryable work keeps the parent deferred. Shared active work does not consume another blob
attempt. Exhausted selected work produces terminal incomplete disposition. Findings from completed
selected blobs retain image, digest, kind, class, and position provenance and use the existing
authoritative ingestion, projection, and keycheck paths.
Config history classification affects only selection order. It never changes detector output,
finding authority, digest identity, or completion criteria.
### 6. Add adaptive modes without changing legacy rollout semantics
Existing `full`, timeout-only `canary`, and legacy `layer` meanings remain available for durable
version-one work and rollback. Two explicit version-two modes are added:
- `adaptive-canary` assigns a configured basis-point cohort across all immutable Docker manifests by
a stable versioned hash; cohort members use adaptive plans and non-members use full-image scanning.
- `adaptive` uses adaptive plans for every eligible immutable Docker claim.
Neither mode depends on a previous full-image timeout. Retry and checkpoint attempts for the same
manifest and selector retain the same assignment. Invalid mode, policy, migration, or gate state
fails closed to full-image execution before any adaptive plan is bound. Returning configuration to
`full` changes only new claims and leaves adaptive plans and audit rows intact.
The existing timeout-only canary remains independent and may continue while adaptive shadow evidence
is gathered. The broad legacy `layer` mode remains operationally disabled because its selector did
not pass recall gates.
### 7. Gate rollout with non-authoritative aggregate shadow evidence
An operator-invoked bounded shadow evaluator selects 50-100 exact immutable images whose authoritative
full-image scans completed successfully under one scan fingerprint. It executes the candidate
adaptive policy using the same downloader, validators, process containment, deadlines, and scanner
fingerprint, but under a shadow authority that cannot call normal result ingestion or mutate target
status, result reservations, global blob coverage, findings, keycheck candidates, projections, or
source counters.
For each paired control, full routed identities are read from the protected database as
`(service, provider_key_hash)` and adaptive routed identities are derived in private memory through
the same candidate normalization. Detector identities use `detector_secret_hash`. Identity sets,
raw findings, commands, provider material, image names, and bearer tokens are discarded after
intersection counts are computed.
The durable report contains only policy hashes, cohort and completion counts, aggregate full,
adaptive, and intersection counts, aggregate slot milliseconds, bounded failure counts, threshold
results, and timestamps. It records no per-image row or identity. Slot timing uses the same outer
monotonic boundary from admitted work through durable shadow sink completion for both paths; scan
subprocess duration remains a diagnostic, not the gate denominator.
A report passes only when:
- 50-100 controls completed both paths without integrity, containment, or fence failure;
- aggregate routed-identity recall, `intersection / full`, is at least 85%;
- aggregate adaptive/full slot-time ratio is at most 40%;
- every adaptive omission is represented in coverage counts; and
- no credential persistence, quarantine, projection, source-failure, or resource-bound regression
is observed.
Reports are bound to exact selector, execution, and scan policy hashes. Stale or incomplete reports
cannot authorize another policy. Adaptive canary remains fail-closed until a matching report passes.
Broad `adaptive` enablement additionally requires a stable low-percentage production canary over at
least one repository-refresh interval. Operators change rollout configuration explicitly; shadow
evidence never changes execution mode by itself.
Alternatives rejected:
- Routing shadow candidates through keycheck would make the experiment authoritative and consume
external capacity.
- Persisting per-image shadow identities creates unnecessary sensitive correlation data.
- Comparing only detector counts does not measure the routed identities the scanner is intended to
produce.
- Automatically enabling adaptive mode from a report removes the operational rollback checkpoint.
### 8. Preserve incomplete warning semantics without redundant retries
The first completed 50-control production shadow report failed closed. Routed recall was 12 of 18
identities (66.7%), adaptive/full slot time was 46.6%, and the report recorded 43 aggregate
failures. Its selection evidence showed 230 descriptors omitted by the eight-layer limit and 14
oversized descriptors, so selector recall remains the primary rollout blocker.
The same evidence exposed a separate execution defect. TruffleHog diagnostics such as
`chunk_processing` and `detector_timeout` are explicitly deterministic, non-retryable warnings.
Their findings must be retained, but the affected blob cannot establish complete coverage. The
diagnostic adapter previously discarded the non-retryable bit, causing the layer executor to
download and scan the same incomplete blob up to three times before reaching the same terminal
state. The adapter now preserves aggregate warning retryability and the layer executor terminates
that blob after the first deterministic warning. It does not mark the blob covered or remove the
report failure.
The historical report schema retained only a total failure count, so its 43 failures cannot be
decomposed exactly after the fact. Future shadow runs keep a fixed allow-list of aggregate-only
failure categories in protected memory and print their totals in the final operator summary without
changing report authority or persisting target-level evidence.
Shadow execution also suppresses target labels in finding-filter logs. Production scans retain their
existing target logging, while both private full and private layer paths emit only aggregate filter
counts. This closes a privacy gap found in the first report log without weakening normal operational
diagnostics.
## Risks / Trade-offs
- [History is malformed, misleading, or secret-bearing] -> Bound and validate it, persist only an
enum, fall back to `unknown`, and keep identity/coverage independent of classification.
- [Application secrets exist in a low-priority or giant layer] -> Scan all-fit images, keep unknown
ahead of generic/bulk classes, record partial scope, retain full controls, and enforce the 85%
routed-recall gate.
- [Unsafe policy reuse suppresses a required rescan] -> Separate hashes and permit legacy aliasing
only from exact successful compatible evidence in one fenced transaction.
- [Selector changes expand coverage during retry] -> Freeze the earliest full position map under the
selector hash before leases are issued.
- [Multi-blob checkpoints increase work lost on crash] -> Bound count/bytes and retain independent
per-blob leases and ingestion records; no pre-handoff result becomes covered.
- [Config prefetch adds Registry traffic] -> Reuse the already bounded authenticated downloader and
avoid a second fetch when the config is leased in the same plan.
- [Individual layer scans expose whiteouted historical files] -> Preserve position provenance and
describe evidence as image-content coverage, not merged-root state.
- [Shadow evaluation consumes slots] -> Keep it operator-invoked, bounded, deterministic, and
subject to the existing slot/resource controls.
- [Aggregate reports hide individual anomalies] -> Fail the whole report on incomplete paired work
and retain bounded failure counts without persisting target identity.
- [Deterministic scanner warnings consume repeated transfer and slot time] -> Preserve their
non-retryable policy, keep findings and incomplete coverage, and terminate the blob on its first
bounded attempt.
## Migration Plan
1. Ship version-one and version-two validators, new modes, and migration code while production stays
on its existing timeout-only canary.
2. Stop authoritative runtime and verify no active result, queue, or blob leases remain.
3. Add the selector-policy coverage column, aggregate shadow-report state, required timing/fence
columns, indexes, and a new migration marker.
4. Backfill every historical image-coverage row from its exact valid reservation plan. Abort and
roll back the migration transaction on an orphan, malformed plan, or invalid hash; then make the
column non-null and run exact schema validation.
5. Restart with unchanged mode and verify legacy claims, projection, keycheck, resolver, quarantine,
and source health before creating version-two work.
6. Run the private shadow evaluator for 50-100 completed controls. Keep adaptive modes fail-closed if
the matching report misses recall, timing, safety, or completion gates.
7. Enable a low deterministic `adaptive-canary`, monitor at least one repository-refresh interval,
and compare source failures, coverage reasons, routed yield, slot time, and quarantine.
8. Increase canary basis points and finally enable `adaptive` only after every gate remains satisfied.
9. Roll back immediately by setting mode to `full` or the existing timeout-only `canary`. Keep all
version-two plans, reports, and coverage rows for audit and exact future resume.
## Open Questions
- Which initial bounded checkpoint count and byte target provide the best reduction in continuation
delay without increasing crash rework materially?
- Which classifier allow-list revisions improve routed recall in the first 50-100 controls? Every
revision will receive a new selector-policy hash rather than changing an existing policy.
- What adaptive-canary basis-point sequence should operators use after the shadow gate passes?