Files
2026-09-30 20:30:56 +03:00

12 KiB

Context

DockerHub discovery currently inspects a bounded tag page but stops after the first eligible digest. Different tags often alias the same manifest or share the same ordered layers, so increasing the tag count without content-aware selection would mostly multiply duplicate work. The Hub tag response identifies platform digests but does not expose their layers; those come from the Docker Registry v2 manifest API.

GitHub and GitLab updated-target admission currently uses provider timestamps. A claimed scan then points TruffleHog at a mutable repository URL with a rolling age boundary and a depth cap. The timestamp is useful for deciding that something may have changed, but it neither identifies the exact ref nor proves which commit was scanned. Existing PostgreSQL reservations provide the fence under which an immutable plan can be attached.

Goals / Non-Goals

Goals:

  • Spend each repository's Docker budget on up to three distinct ordered layer graphs.
  • Keep Docker tag, manifest, request, cache, and emitted-target counts explicitly bounded.
  • Bind GitHub and GitLab scans to a provider-resolved ref and commit SHA after a fenced claim.
  • Use the last successfully covered SHA as the incremental boundary for the same ref.
  • Preserve exact plan and coverage evidence through retries, worker failure, and mid-scan updates.
  • Keep first-scan history bounded while making every later delta independent of rolling age and depth limits.

Non-Goals:

  • Persist a global historical inventory of Docker layers.
  • Guarantee discovery of a branch that appears and disappears between metadata polling cycles.
  • Replace repository metadata search with a global push-event feed.
  • Remove source API rate limits, target timeouts, result bounds, or queue admission bounds.
  • Claim that a bounded first scan covered history older than its configured baseline depth.

Decisions

Resolve platform manifests before selecting Docker targets

For each bounded Hub tag candidate, the resolver will choose the requested platform child digest, obtain a short-lived pull token for the public repository, and fetch that child manifest from the Docker Registry v2 API. A candidate identity contains its tag, update time, immutable platform manifest digest, and ordered layer digest tuple.

The top-level multi-platform index digest will not be emitted when a matching child digest exists. Registry responses must be JSON manifests of bounded size with valid sha256 layer digests. Token requests use a fixed Docker authentication origin and repository pull scope; credentials are never placed in cache records or logs.

Alternatives rejected:

  • Comparing tag names or manifest digests alone, because different manifests can still contain the same layer chain.
  • Pulling every candidate image before selection, because manifest metadata is sufficient and much cheaper.
  • Persisting every layer immediately, because within-resolution novelty provides the requested bounded diversity without a new authoritative subsystem.

Select up to three graphs deterministically

The source setting docker_images_per_repository is clamped to one through three. After exact ordered-tuple deduplication, selection uses stable tie breaking:

  1. The newest resolved graph.
  2. The remaining graph that contributes the most layers not present in the selected union, with recency as the tie breaker.
  3. The oldest remaining distinct graph.

If fewer distinct graphs exist, fewer targets are emitted. Reordered layer tuples remain distinct because layer order changes the image filesystem. Selected targets retain the existing immutable repository@sha256:... identity, so queue deduplication and Docker scan execution do not change.

Only complete graph-selection results are stored in the disposable tag cache. A partial manifest-resolution failure can return successfully resolved targets for the current cycle, but it is not cached as complete and is reported as partial coverage.

Resolve and bind a Git plan after claim

Repository timestamps remain coarse admission signals. Once PostgreSQL has fenced a queue row and result reservation, the worker resolves the source-provided ref hint or the repository's current default branch through the GitHub or GitLab API. The resolver returns a normalized ref and exact head SHA.

The reservation is then bound transactionally to an immutable plan containing ref, head SHA, optional covered base SHA, plan mode, and baseline bounds. Binding requires the active reservation and claim lease token. A missing or malformed revision is a retryable source failure; the worker does not silently scan a mutable URL and does not advance exact coverage.

Metadata search can identify only repository-level activity, so its exact scope is the provider-resolved default branch. Event-backed discovery may supply a more specific ref. Polling cannot guarantee capture of transient or unadvertised refs; that limitation remains explicit.

Alternatives rejected:

  • Resolving one SHA for every search result before admission, because that would spend scarce API quota on records that are never claimed.
  • Encoding SHA in queue target identity, because it would create unbounded rows and bypass existing changed-target coalescing.
  • Copying a timestamp into a covered field at claim, because failed work would look complete.

Execute pinned baselines and deltas

An initial or ref-changed plan scans the exact head with the configured first-scan depth bound. A same-ref plan with a different successfully covered head scans the pinned head with --since-commit <covered SHA> and omits rolling age and maximum-depth limits. A plan whose resolved head already equals the covered head produces an exact no-op result.

The installed TruffleHog binary's ability to accept a commit SHA as --branch is a deployment contract and will be covered by a local repository contract test. If the covered base is unavailable after a force push or ref recreation, execution falls back to a pinned bounded baseline and labels the result as such; it never reports an incremental range as covered when the base was not usable.

This delta means all commits reachable from the pinned head after the covered boundary, not merely the final filesystem diff. An add-then-delete sequence in separate new commits therefore remains visible.

Advance covered SHA only during fenced successful ingestion

target_queue stores the last successfully covered ref and head SHA. result_reservations stores the immutable claimed plan, and normalized scan metadata stores the executed plan and whether execution remained pinned. On fenced ingestion with queue disposition done, the plan in metadata must match the reservation. Only a successful pinned baseline, successful incremental scan, or exact no-op advances or confirms the covered head. Failed, deferred, unbound, or mutable fallback work leaves coverage unchanged.

Because the claimed head is immutable, a newer provider update during execution remains discoverable after completion and can create a later plan from the just-covered head.

Risks / Trade-offs

  • [Registry manifest calls increase Docker API traffic] -> Keep tag candidates and emitted graphs bounded, reuse one scoped token per repository resolution, and do not cache partial results as complete.
  • [Many tags alias one graph] -> Deduplicate exact ordered layer tuples before queue insertion.
  • [A source API is unavailable after claim] -> Produce a retryable source failure and refund through the existing bounded lifecycle without changing coverage.
  • [TruffleHog SHA branch behavior differs by version] -> Add a local contract test against the configured binary and fail closed when immutable pinning is unsupported.
  • [Force-pushed base is no longer reachable] -> Retry as a pinned bounded baseline and label the loss of incremental continuity.
  • [Default-branch resolution misses non-default branch activity] -> Record the resolved scope honestly and allow event-backed ref hints; a complete ref-event feed remains future work.
  • [First-scan depth remains bounded] -> Treat the first head as the future delta baseline without claiming unbounded historical coverage.
  • [Three Docker graphs can triple downstream work] -> Clamp the per-repository setting to three and retain all existing source, queue, timeout, and output bounds.

Migration Plan

  1. Add nullable Git coverage and reservation-plan columns through idempotent PostgreSQL schema initialization.
  2. Deploy graph resolution, plan binding, and tests with explicit configuration gates.
  3. Enable docker_images_per_repository: 3 for DockerHub while retaining the existing twenty-tag candidate bound.
  4. Enable exact Git planning for GitHub and GitLab; legacy rows start with no covered SHA and receive a pinned bounded baseline on their next admitted scan.
  5. Observe partial manifest resolution, distinct graph counts, exact no-ops, baseline resets, delta scans, failures, and strict usable yield.

Rollback is configuration-first. Set Docker images per repository back to one and disable exact Git planning; nullable schema additions remain inert and require no destructive migration.

Open Questions

  • Whether production evidence supports increasing the Docker candidate tag page beyond twenty without exhausting Hub rate limits.
  • Whether a later change should snapshot all advertised refs or consume a dedicated push-event feed for complete non-default-branch coverage.

Implementation Evidence

Implementation completed and locally validated on 2026-08-28:

  • Docker resolution uses the bounded Registry v2 bearer flow, immutable platform-child digests, ordered-layer graph deduplication, and deterministic newest/novel/oldest selection. Partial graph resolution emits only proven targets, retains the repository for retry, and is not cached as complete.
  • PostgreSQL reservations bind one canonical exact Git plan under the active queue/reservation lease fence. Covered ref/head state advances in the same fenced transaction as successful result ingestion after exact plan and execution-evidence comparison.
  • GitHub and GitLab resolve either an explicit branch ref or the provider default branch to an exact commit. Baseline, delta, no-op, and continuity-reset execution remain pinned to the bound head.
  • The checked-in core configuration enables three Docker graphs and exact Git planning with a baseline depth of 100, two ref-resolution attempts, a 10-second shared timeout, and a 1 MiB response limit.
  • python -B -m pytest -p no:cacheprovider -q tests/test_exact_git_scan_planning.py tests/test_validated_high_scanner_fixes.py::DockerTagIdentityTests tests/test_runtime_safety_layer.py tests/test_pipeline_cutover_invariants.py tests/test_migration_runtime_safety.py tests/test_pipeline_postgres_integration.py::PipelinePostgresIntegrationTests::test_exact_git_plan_binding_and_coverage_are_fenced completed with 141 passed.
  • The PostgreSQL integration scenario covers idempotent/conflicting binding, baseline to delta to no-op progression, durable exact metadata, a provider update observed during an active scan, successful fenced coverage advancement, and a post-bind worker refund that cannot advance coverage.
  • The configured local TruffleHog binary passed the exact-SHA --branch contract test included in the focused suite.
  • python -B -m py_compile app/scanner.py app/scanner_db.py app/console_runner.py tests/test_exact_git_scan_planning.py tests/test_pipeline_postgres_integration.py completed successfully.
  • openspec validate improve-core-scan-coverage --strict reported the change as valid.

Retained rollout limits are twenty Docker tag candidates, at most three emitted distinct graphs, 8 MiB per Registry manifest, 1,000 index descriptors, 2,048 layers, bounded source/queue/result limits, and PostgreSQL-only exact Git plan binding. No production migration, process restart, or rollout was performed as part of implementation. Deployment still requires the offline idempotent runtime-safety migration before restarting sources. Configuration-first rollback remains docker_images_per_repository: 1 plus disabling exact_git_planning_enabled; nullable schema additions may remain in place.