Files
2026-09-30 20:30:56 +03:00

135 lines
19 KiB
Markdown

## Context
DockerHub discovery now durably admits pages and rotates across 61 queries, but repository anchors are physically deduplicated by `(source, normalized_target)` and retain only the first inserting query. The resolver is FIFO, runs in batches of 100 only after scanable image work drains, and silently clamps image selection to three graphs. The current backlog therefore cannot provide fair keyword evidence or answer whether versions four through ten add useful findings.
The runtime is PostgreSQL-final-cutover and safety-critical. Queue leases, resolver tokens, content reservations, quarantine holds, and existing cold-policy events must remain authoritative. Normal Docker image scans already traverse every layer; this change varies immutable image/version depth and records manifest positions without enabling broad layer-fallback mode.
## Goals / Non-Goals
**Goals:**
- Run one durable experiment bounded to 1,200 unique immutable image targets and approximately 24-48 hours of observed throughput.
- Give every configured Docker query equal breadth before spending capacity on deeper versions.
- Preserve many-to-many discovery provenance while retaining one physical queue row and scan per normalized image target.
- Hold non-cohort repository anchors reversibly without bypassing active fences or unrelated hold policies.
- Keep ordinary and experiment image depths independently authoritative and validate them before runtime side effects.
- Persist enough immutable manifest and selection evidence to compare ranks 1-3 with ranks 4-10 and attribute findings to exact layers where evidence exists.
**Non-Goals:**
- Do not change Docker target normalization or merge repository aliases.
- Do not enable broad layer-only scanning or change the eight-layer recovery budget.
- Do not infer a layer for findings that lack an exact Docker layer digest.
- Do not resurrect failed, quarantined, already-completed, or independently cold targets automatically.
- Do not reconstruct historical many-to-many provenance; only fresh observations are experiment-eligible.
- Do not make experiment mutation APIs available on SQLite/file-queue execution.
## Decisions
### Fresh relational provenance
Page admission will upsert `docker_repository_query_provenance` in the same transaction that admits repository anchors. Its key is `(source, query, repository_queue_id)` and it records first/last observation, search rank, cycle/policy evidence, and observation count. Existing queue attribution may be seeded as `legacy_queue`, but only a fresh complete observation under the experiment policy is eligible.
This preserves physical deduplication while allowing one repository to belong to several keywords. A JSON column on `target_queue` was rejected because concurrent page admissions, ranking, indexing, and foreign-key validation require relational updates.
### Durable experiment state machine
The experiment uses PostgreSQL tables and compare-and-swap transitions:
`collecting -> planned -> holding -> resolving -> active -> draining -> completed -> released`
Any phase can enter `held` on configuration drift, incomplete query coverage, capacity conflict, or stale fencing. Rows carry a config hash, exact ordered-query hash, selector version/hash, counters, and timestamps. Every transaction that can change experiment provenance, repository membership, resolver state, target/binding/reservation state, queue state, or experiment-owned policy events first locks the experiment row. Cohort/target/queue locks follow that aggregate lock and global capacity is acquired last. Authority validation holds the same experiment lock while reading its multi-table snapshot, which provides a stable READ COMMITTED protocol without changing isolation for ordinary workers.
### Fair bounded cohort
Planning requires one complete fresh deep observation pass for all 61 configured queries. After that pass, each query contributes up to ten eligible previously unseen repositories; a query with fewer or zero new repositories remains in the authority with its exact desired and selected counts, and historical targets are never substituted. Repository selection is deterministic by discovery rank and stable queue identity. Scheduling proceeds round-robin across the available memberships by repository rank and query ordinal, never by global FIFO.
Every selected breadth membership must terminate with either an eligible physical rank-one immutable target or exact hashed image-unavailability evidence before activation or completion. When fresh resolver evidence excludes an existing terminal, quarantined, independently cold, or fenced immutable candidate, the resolver records only bounded hashed skip evidence and compacts the next eligible graph/alias to the experiment rank. If no eligible graph remains, it deterministically advances to the next fresh ranked repository observation under the pinned generation/policy without reactivating its queue row. Once that fresh replacement pool is exhausted, only that membership becomes terminal `skipped` with `no_eligible_physical_target`; it consumes no target/capacity slot and remains separately visible as image-level scarcity. Missing, malformed, or mutated skip evidence moves the experiment to `held`. A query that had no eligible repository at planning time has no breadth membership to satisfy and remains explicitly reported as unavailable rather than missing work.
For each query:
- up to ten available fresh repositories contribute their newest distinct immutable image;
- when the query has at least one repository, one of them, chosen deterministically by the largest valid distinct layer-graph count with stable tie breaks, contributes image ranks 2-10.
The theoretical maximum remains `61 * (10 + 9) = 1,159`; honest query scarcity lowers the actual planned maximum. A separate transactional ceiling of 1,200 unique immutable queue targets protects against planner defects and concurrency. Shared repositories/images consume physical capacity once while retaining all query selection rows.
The dispatch order is breadth rank 1 for every query, then deep ranks 2-3 across every query, then deep ranks 4-10 across every query. These are completion barriers: no wave-two target can be reserved until every physical wave-one target has a latest terminal experiment binding, and wave three similarly waits for wave two. Refunding a terminal attempt returns its target to pending and reopens the earlier barrier. A target shared across selections retains its earliest wave and is scanned once. This makes partial experiment results interpretable even if runtime is stopped early.
### Reversible holds
The planner creates an explicit reviewed experiment hold manifest. Existing fenced cold APIs apply it only to unresolved Docker repository anchors with no active lease, resolver token, content reservation, quarantine, or unrelated hold. Every event stores the prior queue state and experiment identity. Release reverses only unreversed events belonging to this experiment.
Newly observed non-cohort anchors are admitted directly into the same authorized experiment hold after planning. Far-future retry timestamps and destructive deletion were rejected because they obscure ownership and cannot be safely reversed.
The reviewed hold/reactivation set is bounded at 250,000 rows. This is above the strict maximum configured search surface of `61 * 30 * 100 = 183,000` rows. Selection uses a `LIMIT 250001` fail-closed overflow check and application locks deterministic 500-row parts in one transaction, so PostgreSQL parameter limits cannot silently truncate the reviewed set. Experiment authority continuously validates every owned event's queue state, config/policy/manifest hashes, audit hash, and reversal state.
Release is never automatic and a reviewed reactivation manifest can be generated or applied only from `completed`. Holding, resolving, active, draining, and held experiments cannot transition directly to `released`.
### Dedicated resolver lane and atomic completion
Experiment repositories use a dedicated round-robin resolver claim lane that does not wait for the ordinary image queue to empty and runs before ordinary backlog suppression. It bypasses only discovery and scan-queue-empty gating; PostgreSQL final-cutover, enabled state, API authentication, and shared rate-limit stops remain mandatory. Resolver completion atomically validates its owner/generation/token, persists immutable image targets, manifest graphs/layers, selection rank/reason, experiment target membership, and the unique-target counter.
Resolver leases are renewed under that exact owner/generation/token before and between bounded tag/manifest network stages and verified again after remote work. A stale or expired token performs no write. Authority validation atomically returns interrupted resolver memberships to pending and records `held(stale_resolver_fence)`; only that reason may automatically return to `resolving` on a later claim, and only after a complete clean authority snapshot proves every resolver fence absent. Consumed non-conclusive remote attempts use a code-pinned 300-second exponential backoff capped at 3,600 seconds. On the third consumed attempt, only that membership terminates: breadth work becomes exact hashed `remote_unavailable_after_attempt_limit` scarcity, while a deep probe keeps its already selected images and closes at the achieved depth. Existing clean `held(resolver_attempt_limit)` rows from the earlier policy are terminalized by the same rule on the next fully validated claim. Evidence, capacity, configuration, and authority conflicts remain fail-closed experiment holds; a shared cooldown that made no remote request refunds the claim attempt.
A resolver-attempt refund is never automatic. A one-time offline reviewed manifest may refund exactly two attempts only when the private immutable log snapshot contains exactly two target-bound instances of the retired local zero-graph limit defect. The manifest contains only queue/member IDs and hashes of target identity, prior error, entry evidence, log, and authority; it never contains raw targets. Application requires exact SHA approval, stopped sources, the unchanged attempt-limit hold, unfenced membership state, and an additive append-only audit row unique to that membership and recovery kind. Real partial/remote failures and every membership without exact old-bug evidence remain untouched.
The offline reviewed disposition protocol remains available for a legacy persisted `resolver_attempt_limit` hold before runtime resumes. It snapshots the deterministic fresh replacement search as IDs, ranks, conflict codes, and target-identity hashes. Exact-SHA application either replaces the held membership with the first unchanged eligible fresh repository and resets only that membership's attempts, or terminally records scarcity. Normal runtime no longer serializes the whole pilot on target-local remote exhaustion.
Only previously unseen pending immutable targets are experiment-eligible. Conflicts with `done`, `failed`, `quarantined`, or another cold policy cause deterministic fresh replacement and, after that pool is exhausted, an explicitly evidenced terminal membership skip. Historical work is never automatically reactivated.
The reviewed cohort hash covers all 61 query authorities, each query's desired count of ten, its actual selected count from zero through ten, and every selected repository identity. Repository exclusions and replacements do not rewrite it. Once every selected breadth membership has either rank-one evidence or exact terminal skip evidence, a separate immutable runtime-selection hash freezes effective repository identities, terminal scarcity evidence, and one graph-informed deep-probe choice for each query with at least one image-bearing membership. An all-skipped query has no deep probe. Claims continuously validate both hashes, so terminal evidence and `is_deep_probe` cannot mutate after selection.
### Config is the source of truth
`docker_images_per_repository` remains the ordinary FIFO resolver depth and is fixed at the reviewed production value of three while the experiment is configured. `docker_depth_experiment.deep_images_per_repository` independently fixes the dedicated experiment lane at ten. The shared selector accepts an explicit validated limit through ten without a hidden clamp, but disabling experiment activation never raises ordinary FIFO work above three. Enabled experiment configuration rejects ordinary periodic repository refresh (`docker_repository_refresh_max_per_cycle` must be zero), preventing the ordinary resolver from mutating experiment repository authority.
The experiment mapping enables provenance collection even while `enabled` is false. Both collection and activation therefore require managed PostgreSQL final-cutover, search mode, immutable digests, exact ordered queries, one representable effective pages/per-page pass policy, strict query overrides, and valid Docker platform/candidate settings. Validation runs before secrets, state, database, network, or worker initialization. The canonical semantic config hash includes ordered effective query policies, the code-pinned collection generation, and all platform selector inputs, but excludes operational `enabled`; false-to-true activation therefore preserves the reviewed frozen hash while enabled state remains a separately validated claim gate. Configuration/order/selector/generation hash drift moves an active experiment to `held` rather than silently changing its cohort, and disabling in `holding`, `resolving`, `active`, or `draining` moves it to `held`.
The pinned collection generation is persisted on each discovery pass. While the experiment mapping remains operationally disabled for reviewed collection, Docker discovery continues but ordinary Docker resolver and scan admission remain paused so an unbounded historical backlog cannot pin query rotation or consume fresh cohort candidates. After reviewed activation, that collection-only barrier is removed and the frozen experiment lanes own resolution and scanning. Until PostgreSQL contains one complete deep pass for the pinned generation, policy, and ordered-query authority, every query is forced through deep provenance collection regardless of a pre-migration 72-hour runner-state marker. Terminal page evidence includes total count and must be coherent with the page's absolute result bound; undercounts cannot complete a pass.
### Deterministic image selection beyond three
The existing first-three semantics remain: newest distinct graph, maximum marginal layer novelty, then oldest distinct graph. Ranks four through ten repeatedly choose maximum marginal layer novelty against already selected graphs, with temporal distance, recency, normalized target, and graph hash as stable tie breakers. Invalid/duplicate graphs do not consume a rank.
### Manifest and finding attribution
Every selected image persists its immutable manifest identity and ordered descriptors. Layer position 1 is from the base; `layer_count - position + 1` is from the top. Duplicate layer digests at different positions remain separate rows.
During result ingestion, an exact finding `Data.Docker.layer` digest is joined to all matching positions in that selected image and written to a finding-layer relation. Findings without exact evidence remain explicitly unattributed. Existing finding identity metadata is not rewritten.
Exact attribution is fail-closed at a code-pinned limit of 100,000 projected finding-position rows per experiment result bundle. Ingestion counts every matching duplicate position before inserting scan evidence; overflow quarantines the whole bundle without truncation, inference, or a partial database commit. Reports aggregate position fan-out in SQL and read all report relations from one read-only repeatable snapshot.
### Reservation-bound reporting
Experiment scan bindings are created atomically with queue reservations, so historical scans of the same target cannot leak into the pilot. Reports join experiment query, repository, selection, unique target, reservation, target scan, findings, layer positions, and frozen/current keycheck outcomes.
Capacity or quarantine saturation changes the experiment to `held` in the same experiment-row-locked transaction that aborts admission. There is no committed interval in which an aborted experiment admission remains active. Result readiness, renewal, refund, quarantine, and ingestion acquire the same experiment aggregate lock before reservation, queue, binding, or capacity transitions.
Reports expose aggregates only. Per-query attribution credits shared physical scans to every observing query; global totals deduplicate experiment target and scan IDs. Marginal identity yield is assigned to the minimum image rank, and rank buckets are 1-3 versus 4-10. Unattributed layer findings and incomplete keychecks remain visible rather than being dropped.
## Risks / Trade-offs
- [A query has fewer than ten fresh repositories after the complete pinned deep pass] -> Freeze its exact available count, substitute no historical targets, and report the shortage against the desired count of ten.
- [The resolver cannot find ten distinct graphs for a deep probe] -> Record actual deep depth and continue with fewer depth selections; each breadth membership still requires rank one or exact terminal image-unavailability evidence.
- [Shared layers/findings multiply through many-to-many joins] -> Preaggregate physical scan/finding identities before query attribution and report both physical and attributed totals.
- [Native scanner output lacks a layer digest] -> Count it as unattributed and do not infer a position.
- [Concurrent resolver completions exceed 1,200] -> Serialize counter updates under the locked experiment row and recheck before insert.
- [Existing queue work competes with the pilot] -> Apply reviewed reversible holds and use experiment-only claim ordering while leaving unrelated active fences untouched.
- [Runtime/config changes invalidate evidence] -> Hash exact ordered queries, selector code version, and experiment config; transition fail-closed to `held` on drift.
- [Migration fails or rollout must be reverted] -> Use additive schema only, keep activation disabled during migration, and roll back configuration without deleting provenance or experiment evidence.
## Migration Plan
1. Canonically stop the supervised runtime and apply the additive migration with experiment activation disabled.
2. Start the runtime to collect fresh many-to-many provenance through one complete deep observation pass for all 61 queries.
3. Canonically stop, verify quiescence, generate and review the deterministic cohort/hold manifest, and activate the frozen experiment definition.
4. Start runtime, resolve and dispatch the cohort in round-robin waves, and monitor the 1,200 ceiling, leases, restart counters, and quarantine.
5. Transition to draining when all selections are terminal, then generate the immutable aggregate report.
6. Keep non-cohort rows cold until a reviewed post-experiment decision; release by exact experiment event identity when approved.
Rollback disables new experiment claims and moves the experiment to `held`. Existing reservations finish through normal fencing. Additive provenance, manifest, and audit evidence remains intact for diagnosis and later resumption.
## Open Questions
None. The experiment limits, fairness rule, fresh-only eligibility, hold behavior, attribution model, and fail-closed rollout are fixed by the reviewed plan.