Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-09-09
@@ -0,0 +1,134 @@
## Context
DockerHub discovery now durably admits pages and rotates across 61 queries, but repository anchors are physically deduplicated by `(source, normalized_target)` and retain only the first inserting query. The resolver is FIFO, runs in batches of 100 only after scanable image work drains, and silently clamps image selection to three graphs. The current backlog therefore cannot provide fair keyword evidence or answer whether versions four through ten add useful findings.
The runtime is PostgreSQL-final-cutover and safety-critical. Queue leases, resolver tokens, content reservations, quarantine holds, and existing cold-policy events must remain authoritative. Normal Docker image scans already traverse every layer; this change varies immutable image/version depth and records manifest positions without enabling broad layer-fallback mode.
## Goals / Non-Goals
**Goals:**
- Run one durable experiment bounded to 1,200 unique immutable image targets and approximately 24-48 hours of observed throughput.
- Give every configured Docker query equal breadth before spending capacity on deeper versions.
- Preserve many-to-many discovery provenance while retaining one physical queue row and scan per normalized image target.
- Hold non-cohort repository anchors reversibly without bypassing active fences or unrelated hold policies.
- Keep ordinary and experiment image depths independently authoritative and validate them before runtime side effects.
- Persist enough immutable manifest and selection evidence to compare ranks 1-3 with ranks 4-10 and attribute findings to exact layers where evidence exists.
**Non-Goals:**
- Do not change Docker target normalization or merge repository aliases.
- Do not enable broad layer-only scanning or change the eight-layer recovery budget.
- Do not infer a layer for findings that lack an exact Docker layer digest.
- Do not resurrect failed, quarantined, already-completed, or independently cold targets automatically.
- Do not reconstruct historical many-to-many provenance; only fresh observations are experiment-eligible.
- Do not make experiment mutation APIs available on SQLite/file-queue execution.
## Decisions
### Fresh relational provenance
Page admission will upsert `docker_repository_query_provenance` in the same transaction that admits repository anchors. Its key is `(source, query, repository_queue_id)` and it records first/last observation, search rank, cycle/policy evidence, and observation count. Existing queue attribution may be seeded as `legacy_queue`, but only a fresh complete observation under the experiment policy is eligible.
This preserves physical deduplication while allowing one repository to belong to several keywords. A JSON column on `target_queue` was rejected because concurrent page admissions, ranking, indexing, and foreign-key validation require relational updates.
### Durable experiment state machine
The experiment uses PostgreSQL tables and compare-and-swap transitions:
`collecting -> planned -> holding -> resolving -> active -> draining -> completed -> released`
Any phase can enter `held` on configuration drift, incomplete query coverage, capacity conflict, or stale fencing. Rows carry a config hash, exact ordered-query hash, selector version/hash, counters, and timestamps. Every transaction that can change experiment provenance, repository membership, resolver state, target/binding/reservation state, queue state, or experiment-owned policy events first locks the experiment row. Cohort/target/queue locks follow that aggregate lock and global capacity is acquired last. Authority validation holds the same experiment lock while reading its multi-table snapshot, which provides a stable READ COMMITTED protocol without changing isolation for ordinary workers.
### Fair bounded cohort
Planning requires one complete fresh deep observation pass for all 61 configured queries. After that pass, each query contributes up to ten eligible previously unseen repositories; a query with fewer or zero new repositories remains in the authority with its exact desired and selected counts, and historical targets are never substituted. Repository selection is deterministic by discovery rank and stable queue identity. Scheduling proceeds round-robin across the available memberships by repository rank and query ordinal, never by global FIFO.
Every selected breadth membership must terminate with either an eligible physical rank-one immutable target or exact hashed image-unavailability evidence before activation or completion. When fresh resolver evidence excludes an existing terminal, quarantined, independently cold, or fenced immutable candidate, the resolver records only bounded hashed skip evidence and compacts the next eligible graph/alias to the experiment rank. If no eligible graph remains, it deterministically advances to the next fresh ranked repository observation under the pinned generation/policy without reactivating its queue row. Once that fresh replacement pool is exhausted, only that membership becomes terminal `skipped` with `no_eligible_physical_target`; it consumes no target/capacity slot and remains separately visible as image-level scarcity. Missing, malformed, or mutated skip evidence moves the experiment to `held`. A query that had no eligible repository at planning time has no breadth membership to satisfy and remains explicitly reported as unavailable rather than missing work.
For each query:
- up to ten available fresh repositories contribute their newest distinct immutable image;
- when the query has at least one repository, one of them, chosen deterministically by the largest valid distinct layer-graph count with stable tie breaks, contributes image ranks 2-10.
The theoretical maximum remains `61 * (10 + 9) = 1,159`; honest query scarcity lowers the actual planned maximum. A separate transactional ceiling of 1,200 unique immutable queue targets protects against planner defects and concurrency. Shared repositories/images consume physical capacity once while retaining all query selection rows.
The dispatch order is breadth rank 1 for every query, then deep ranks 2-3 across every query, then deep ranks 4-10 across every query. These are completion barriers: no wave-two target can be reserved until every physical wave-one target has a latest terminal experiment binding, and wave three similarly waits for wave two. Refunding a terminal attempt returns its target to pending and reopens the earlier barrier. A target shared across selections retains its earliest wave and is scanned once. This makes partial experiment results interpretable even if runtime is stopped early.
### Reversible holds
The planner creates an explicit reviewed experiment hold manifest. Existing fenced cold APIs apply it only to unresolved Docker repository anchors with no active lease, resolver token, content reservation, quarantine, or unrelated hold. Every event stores the prior queue state and experiment identity. Release reverses only unreversed events belonging to this experiment.
Newly observed non-cohort anchors are admitted directly into the same authorized experiment hold after planning. Far-future retry timestamps and destructive deletion were rejected because they obscure ownership and cannot be safely reversed.
The reviewed hold/reactivation set is bounded at 250,000 rows. This is above the strict maximum configured search surface of `61 * 30 * 100 = 183,000` rows. Selection uses a `LIMIT 250001` fail-closed overflow check and application locks deterministic 500-row parts in one transaction, so PostgreSQL parameter limits cannot silently truncate the reviewed set. Experiment authority continuously validates every owned event's queue state, config/policy/manifest hashes, audit hash, and reversal state.
Release is never automatic and a reviewed reactivation manifest can be generated or applied only from `completed`. Holding, resolving, active, draining, and held experiments cannot transition directly to `released`.
### Dedicated resolver lane and atomic completion
Experiment repositories use a dedicated round-robin resolver claim lane that does not wait for the ordinary image queue to empty and runs before ordinary backlog suppression. It bypasses only discovery and scan-queue-empty gating; PostgreSQL final-cutover, enabled state, API authentication, and shared rate-limit stops remain mandatory. Resolver completion atomically validates its owner/generation/token, persists immutable image targets, manifest graphs/layers, selection rank/reason, experiment target membership, and the unique-target counter.
Resolver leases are renewed under that exact owner/generation/token before and between bounded tag/manifest network stages and verified again after remote work. A stale or expired token performs no write. Authority validation atomically returns interrupted resolver memberships to pending and records `held(stale_resolver_fence)`; only that reason may automatically return to `resolving` on a later claim, and only after a complete clean authority snapshot proves every resolver fence absent. Consumed non-conclusive remote attempts use a code-pinned 300-second exponential backoff capped at 3,600 seconds. On the third consumed attempt, only that membership terminates: breadth work becomes exact hashed `remote_unavailable_after_attempt_limit` scarcity, while a deep probe keeps its already selected images and closes at the achieved depth. Existing clean `held(resolver_attempt_limit)` rows from the earlier policy are terminalized by the same rule on the next fully validated claim. Evidence, capacity, configuration, and authority conflicts remain fail-closed experiment holds; a shared cooldown that made no remote request refunds the claim attempt.
A resolver-attempt refund is never automatic. A one-time offline reviewed manifest may refund exactly two attempts only when the private immutable log snapshot contains exactly two target-bound instances of the retired local zero-graph limit defect. The manifest contains only queue/member IDs and hashes of target identity, prior error, entry evidence, log, and authority; it never contains raw targets. Application requires exact SHA approval, stopped sources, the unchanged attempt-limit hold, unfenced membership state, and an additive append-only audit row unique to that membership and recovery kind. Real partial/remote failures and every membership without exact old-bug evidence remain untouched.
The offline reviewed disposition protocol remains available for a legacy persisted `resolver_attempt_limit` hold before runtime resumes. It snapshots the deterministic fresh replacement search as IDs, ranks, conflict codes, and target-identity hashes. Exact-SHA application either replaces the held membership with the first unchanged eligible fresh repository and resets only that membership's attempts, or terminally records scarcity. Normal runtime no longer serializes the whole pilot on target-local remote exhaustion.
Only previously unseen pending immutable targets are experiment-eligible. Conflicts with `done`, `failed`, `quarantined`, or another cold policy cause deterministic fresh replacement and, after that pool is exhausted, an explicitly evidenced terminal membership skip. Historical work is never automatically reactivated.
The reviewed cohort hash covers all 61 query authorities, each query's desired count of ten, its actual selected count from zero through ten, and every selected repository identity. Repository exclusions and replacements do not rewrite it. Once every selected breadth membership has either rank-one evidence or exact terminal skip evidence, a separate immutable runtime-selection hash freezes effective repository identities, terminal scarcity evidence, and one graph-informed deep-probe choice for each query with at least one image-bearing membership. An all-skipped query has no deep probe. Claims continuously validate both hashes, so terminal evidence and `is_deep_probe` cannot mutate after selection.
### Config is the source of truth
`docker_images_per_repository` remains the ordinary FIFO resolver depth and is fixed at the reviewed production value of three while the experiment is configured. `docker_depth_experiment.deep_images_per_repository` independently fixes the dedicated experiment lane at ten. The shared selector accepts an explicit validated limit through ten without a hidden clamp, but disabling experiment activation never raises ordinary FIFO work above three. Enabled experiment configuration rejects ordinary periodic repository refresh (`docker_repository_refresh_max_per_cycle` must be zero), preventing the ordinary resolver from mutating experiment repository authority.
The experiment mapping enables provenance collection even while `enabled` is false. Both collection and activation therefore require managed PostgreSQL final-cutover, search mode, immutable digests, exact ordered queries, one representable effective pages/per-page pass policy, strict query overrides, and valid Docker platform/candidate settings. Validation runs before secrets, state, database, network, or worker initialization. The canonical semantic config hash includes ordered effective query policies, the code-pinned collection generation, and all platform selector inputs, but excludes operational `enabled`; false-to-true activation therefore preserves the reviewed frozen hash while enabled state remains a separately validated claim gate. Configuration/order/selector/generation hash drift moves an active experiment to `held` rather than silently changing its cohort, and disabling in `holding`, `resolving`, `active`, or `draining` moves it to `held`.
The pinned collection generation is persisted on each discovery pass. While the experiment mapping remains operationally disabled for reviewed collection, Docker discovery continues but ordinary Docker resolver and scan admission remain paused so an unbounded historical backlog cannot pin query rotation or consume fresh cohort candidates. After reviewed activation, that collection-only barrier is removed and the frozen experiment lanes own resolution and scanning. Until PostgreSQL contains one complete deep pass for the pinned generation, policy, and ordered-query authority, every query is forced through deep provenance collection regardless of a pre-migration 72-hour runner-state marker. Terminal page evidence includes total count and must be coherent with the page's absolute result bound; undercounts cannot complete a pass.
### Deterministic image selection beyond three
The existing first-three semantics remain: newest distinct graph, maximum marginal layer novelty, then oldest distinct graph. Ranks four through ten repeatedly choose maximum marginal layer novelty against already selected graphs, with temporal distance, recency, normalized target, and graph hash as stable tie breakers. Invalid/duplicate graphs do not consume a rank.
### Manifest and finding attribution
Every selected image persists its immutable manifest identity and ordered descriptors. Layer position 1 is from the base; `layer_count - position + 1` is from the top. Duplicate layer digests at different positions remain separate rows.
During result ingestion, an exact finding `Data.Docker.layer` digest is joined to all matching positions in that selected image and written to a finding-layer relation. Findings without exact evidence remain explicitly unattributed. Existing finding identity metadata is not rewritten.
Exact attribution is fail-closed at a code-pinned limit of 100,000 projected finding-position rows per experiment result bundle. Ingestion counts every matching duplicate position before inserting scan evidence; overflow quarantines the whole bundle without truncation, inference, or a partial database commit. Reports aggregate position fan-out in SQL and read all report relations from one read-only repeatable snapshot.
### Reservation-bound reporting
Experiment scan bindings are created atomically with queue reservations, so historical scans of the same target cannot leak into the pilot. Reports join experiment query, repository, selection, unique target, reservation, target scan, findings, layer positions, and frozen/current keycheck outcomes.
Capacity or quarantine saturation changes the experiment to `held` in the same experiment-row-locked transaction that aborts admission. There is no committed interval in which an aborted experiment admission remains active. Result readiness, renewal, refund, quarantine, and ingestion acquire the same experiment aggregate lock before reservation, queue, binding, or capacity transitions.
Reports expose aggregates only. Per-query attribution credits shared physical scans to every observing query; global totals deduplicate experiment target and scan IDs. Marginal identity yield is assigned to the minimum image rank, and rank buckets are 1-3 versus 4-10. Unattributed layer findings and incomplete keychecks remain visible rather than being dropped.
## Risks / Trade-offs
- [A query has fewer than ten fresh repositories after the complete pinned deep pass] -> Freeze its exact available count, substitute no historical targets, and report the shortage against the desired count of ten.
- [The resolver cannot find ten distinct graphs for a deep probe] -> Record actual deep depth and continue with fewer depth selections; each breadth membership still requires rank one or exact terminal image-unavailability evidence.
- [Shared layers/findings multiply through many-to-many joins] -> Preaggregate physical scan/finding identities before query attribution and report both physical and attributed totals.
- [Native scanner output lacks a layer digest] -> Count it as unattributed and do not infer a position.
- [Concurrent resolver completions exceed 1,200] -> Serialize counter updates under the locked experiment row and recheck before insert.
- [Existing queue work competes with the pilot] -> Apply reviewed reversible holds and use experiment-only claim ordering while leaving unrelated active fences untouched.
- [Runtime/config changes invalidate evidence] -> Hash exact ordered queries, selector code version, and experiment config; transition fail-closed to `held` on drift.
- [Migration fails or rollout must be reverted] -> Use additive schema only, keep activation disabled during migration, and roll back configuration without deleting provenance or experiment evidence.
## Migration Plan
1. Canonically stop the supervised runtime and apply the additive migration with experiment activation disabled.
2. Start the runtime to collect fresh many-to-many provenance through one complete deep observation pass for all 61 queries.
3. Canonically stop, verify quiescence, generate and review the deterministic cohort/hold manifest, and activate the frozen experiment definition.
4. Start runtime, resolve and dispatch the cohort in round-robin waves, and monitor the 1,200 ceiling, leases, restart counters, and quarantine.
5. Transition to draining when all selections are terminal, then generate the immutable aggregate report.
6. Keep non-cohort rows cold until a reviewed post-experiment decision; release by exact experiment event identity when approved.
Rollback disables new experiment claims and moves the experiment to `held`. Existing reservations finish through normal fencing. Additive provenance, manifest, and audit evidence remains intact for diagnosis and later resumption.
## Open Questions
None. The experiment limits, fairness rule, fresh-only eligibility, hold behavior, attribution model, and fail-closed rollout are fixed by the reviewed plan.
@@ -0,0 +1,27 @@
## Why
DockerHub discovery admitted a large deduplicated repository backlog, but FIFO resolution and the hidden three-image selector cap cannot produce a fair, bounded comparison across all configured keywords. A controlled 24-48 hour experiment is needed to measure keyword yield and the marginal value of image versions beyond the current first three without allowing the backlog to monopolize runtime capacity.
## What Changes
- Add durable many-to-many DockerHub keyword-to-repository provenance so deduplicated targets retain every discovery attribution.
- Add a reversible experiment hold and a round-robin cohort planner that takes up to ten genuinely fresh repositories per configured keyword after one complete pinned deep pass, preserving honest shortfalls without substituting historical targets.
- Scan one current image from each cohort repository and up to ten distinct image graphs from one version-rich deep probe per keyword.
- Enforce a global experiment ceiling of 1,200 unique immutable images and keep non-cohort/new repositories cold until reviewed reactivation.
- Replace the hidden three-image clamp with explicit fail-closed configuration validation and deterministic selection beyond the third image.
- Persist image rank, selection reason, manifest layer graph, layer digest, and base/top layer positions for per-keyword experiment reporting.
- Report first-three versus later-version yield using deduplicated findings, credential identities, verification outcomes, scan time, and coverage.
## Capabilities
### New Capabilities
- `docker-depth-experiment`: Durable provenance, bounded fair cohort scheduling, configurable image depth, reversible holds, layer attribution, and experiment reporting.
### Modified Capabilities
## Impact
- PostgreSQL schema and managed migrations for Docker discovery provenance, experiment cohorts, image selection metadata, and durable reporting state.
- DockerHub page admission, repository resolver scheduling, immutable target creation, and finding attribution in `app/scanner_db.py`, `app/scanner.py`, and `app/console_runner.py`.
- DockerHub configuration and validation in `app/config.yaml` and runner argument construction.
- Focused unit/PostgreSQL integration tests plus canonical offline migration and supervised runtime rollout.
@@ -0,0 +1,252 @@
## ADDED Requirements
### Requirement: Authoritative Docker image depth configuration
The system SHALL obtain separate ordinary and experiment Docker images-per-repository limits from validated configuration. While this experiment is configured, the ordinary FIFO resolver limit SHALL remain three and the dedicated experiment deep limit SHALL remain ten. The system SHALL reject invalid or incompatible collection configuration before secrets, state, database, network, or worker initialization.
#### Scenario: Disabled activation collection rollout
- **WHEN** experiment activation is disabled while provenance collection remains configured
- **THEN** Docker discovery SHALL persist fresh provenance while ordinary Docker resolver and scan admission remain paused, and the configured ordinary resolver depth SHALL remain three for later non-collection operation
#### Scenario: Reviewed false-to-true activation
- **WHEN** a disabled collection is reviewed and operational `enabled` changes from false to true without another configuration change
- **THEN** the frozen semantic configuration hash SHALL remain unchanged and enabled state SHALL be enforced separately at each activation or claim boundary
#### Scenario: Experiment depth ten
- **WHEN** the dedicated experiment resolver receives the validated deep limit of 10
- **THEN** it SHALL be allowed to select up to ten deterministic distinct image graphs without a hidden lower clamp
#### Scenario: Invalid depth
- **WHEN** Docker image depth is boolean, non-integer, below 1, above 10, or incompatible with the configured candidate-tag depth
- **THEN** startup SHALL fail before any runtime side effect
#### Scenario: Mixed discovery policies
- **WHEN** query overrides produce different effective pages or per-page policy values for experiment queries
- **THEN** startup SHALL fail because the current experiment pass authority represents one discovery policy
#### Scenario: Incompatible ordinary refresh
- **WHEN** experiment activation is enabled with periodic ordinary repository refresh greater than zero
- **THEN** startup SHALL fail before runtime side effects
### Requirement: Durable many-to-many discovery provenance
The system SHALL record every fresh Docker query-to-repository observation in the same transaction as page admission while preserving one physical repository queue identity.
#### Scenario: Repository observed by two queries
- **WHEN** two Docker queries observe the same normalized repository anchor
- **THEN** the system SHALL retain two provenance relations and one repository queue row
#### Scenario: Page admission rolls back
- **WHEN** provenance persistence or repository admission fails
- **THEN** neither the page admission nor its provenance and retry progress SHALL commit partially
#### Scenario: Page admission races authority validation
- **WHEN** page ingestion and experiment validation execute concurrently at READ COMMITTED
- **THEN** both SHALL serialize on the experiment authority row and validation SHALL NOT observe a half-committed page, hold event, queue, binding, or reservation transition
#### Scenario: Pre-migration deep marker
- **WHEN** runner state contains a deep-dispatch marker but PostgreSQL has no complete deep pass for the pinned collection generation, policy, and ordered queries
- **THEN** each configured query SHALL be forced through one generation-current deep provenance pass before the normal 72-hour policy resumes
#### Scenario: Underreported result count
- **WHEN** a Docker discovery page reports a total count below its current absolute result bound or incoherent with an empty continuation
- **THEN** that page SHALL fail or be delegated and SHALL NOT provide terminal pass evidence
### Requirement: Fresh complete cohort eligibility
The experiment SHALL use only fresh observations made under its pinned policy and SHALL remain in collecting state until one complete deep observation pass covers every configured query. It SHALL then preserve every query authority row and select up to ten eligible previously unscanned repositories per query, including an explicit actual count of zero when the completed pass produced none.
#### Scenario: Query has insufficient candidates
- **WHEN** the complete pinned deep pass leaves a configured query with fewer than ten fresh eligible repositories
- **THEN** the cohort SHALL retain exactly the available fresh repositories, SHALL NOT substitute historical or previously scanned targets, and SHALL report the unavailable count against the desired quota of ten
#### Scenario: Collection pass is incomplete
- **WHEN** no complete pinned deep pass covers all 61 configured queries
- **THEN** planning SHALL remain unavailable even if partial observations exist
#### Scenario: Legacy attribution exists
- **WHEN** a repository has only historical first-inserter queue attribution
- **THEN** that evidence SHALL NOT satisfy fresh experiment eligibility
### Requirement: Bounded fair experiment cohort
The system SHALL select up to ten fresh repositories per configured query after the complete pass, resolve one newest distinct image from each selected repository when an eligible unseen image exists, and select image ranks 2 through 10 from one deterministic version-rich image-bearing repository for each applicable query, subject to a transactional ceiling of 1,200 unique immutable image targets. A selected repository that exhausts fresh replacements without an eligible image SHALL remain explicit terminal image-level scarcity rather than causing historical substitution.
#### Scenario: Complete 61-query authority
- **WHEN** all 61 query rows are planned from a complete pinned deep pass
- **THEN** the planned selections SHALL be no more than 1,159 before cross-query deduplication, honest scarcity SHALL reduce rather than inflate that count, and physical experiment targets SHALL never exceed 1,200
#### Scenario: Conclusive breadth zero
- **WHEN** a breadth repository resolves conclusively with no eligible immutable image
- **THEN** the resolver SHALL deterministically select the next fresh ranked repository and, when that pool is exhausted, SHALL terminally mark only that membership `skipped` with exact hashed `no_eligible_physical_target` evidence without consuming a physical target slot
#### Scenario: Complete breadth authority
- **WHEN** the experiment activates or completes
- **THEN** every selected breadth membership SHALL have either a valid rank-one selection bound to an eligible physical experiment target or exact terminal image-unavailability evidence, while zero-member queries SHALL remain explicit nonmissing repository scarcity evidence
#### Scenario: Query has no image-bearing membership
- **WHEN** every selected repository for a query terminates with valid image-unavailability evidence
- **THEN** that query SHALL have no deep probe and SHALL remain separately reportable without blocking activation
#### Scenario: Concurrent shared image selection
- **WHEN** concurrent query selections resolve to the same normalized immutable image
- **THEN** the image SHALL consume one physical target slot and SHALL retain every query selection relation
### Requirement: Round-robin experiment scheduling
The system SHALL schedule breadth and depth work by pinned query ordinal rather than repository FIFO.
#### Scenario: Breadth precedes depth
- **WHEN** experiment targets become claimable
- **THEN** ranks 2-3 SHALL remain unclaimable until every physical rank-one target has a terminal latest experiment binding, and ranks 4-10 SHALL similarly wait for ranks 2-3
#### Scenario: Earlier wave is refunded
- **WHEN** a terminal attempt in an earlier wave is refunded and its physical target returns to pending
- **THEN** the earlier completion barrier SHALL reopen and later waves SHALL stop until its replacement attempt completes terminally
#### Scenario: Partial execution
- **WHEN** runtime stops before the experiment completes
- **THEN** completed work SHALL remain evenly attributable to the earliest unfinished round-robin wave
### Requirement: Deterministic distinct graph selection
The resolver SHALL preserve the existing first-three selection semantics and SHALL select ranks 4 through 10 deterministically by marginal layer novelty and stable temporal/identity tie breaks.
#### Scenario: Duplicate tags share a graph
- **WHEN** multiple tags resolve to an identical ordered layer graph
- **THEN** the graph SHALL consume at most one image rank
#### Scenario: Fewer than ten valid graphs
- **WHEN** a deep repository has fewer than ten valid distinct image graphs
- **THEN** the resolver SHALL record the actual depth without inserting invalid or duplicate replacements
### Requirement: Reversible fenced backlog hold
The system SHALL place non-cohort Docker repository anchors into a durable experiment-scoped cold state only when no active lease, resolver token, reservation, quarantine, or unrelated hold prevents the transition, and SHALL preserve the exact prior state for reviewed reactivation.
#### Scenario: Unfenced non-cohort anchor
- **WHEN** an eligible non-cohort unresolved repository is covered by the reviewed experiment hold policy
- **THEN** it SHALL become cold with a durable policy event and SHALL be ignored by ordinary resolver claims
#### Scenario: Independently fenced anchor
- **WHEN** a repository has an active or unrelated safety fence
- **THEN** the experiment SHALL leave it unchanged and record the hold conflict
#### Scenario: Reviewed release
- **WHEN** the experiment hold is released
- **THEN** the experiment SHALL already be completed and only unreversed cold events owned by that experiment SHALL restore their exact prior queue states
#### Scenario: Unsafe early release
- **WHEN** a reviewed release is requested from holding, resolving, active, draining, or held
- **THEN** release SHALL be rejected without changing owned cold events or experiment state
#### Scenario: Full reviewed search surface
- **WHEN** a reviewed hold or reactivation covers more than 100,000 rows up to the strict `61 * 30 * 100` search maximum
- **THEN** deterministic bounded parts SHALL cover every row, aggregate counts/hashes SHALL remain authoritative, and overflow beyond the reviewed 250,000-row ceiling SHALL fail rather than omit rows
### Requirement: Fresh target safety
The experiment SHALL scan only previously unseen immutable image targets and SHALL NOT automatically reactivate completed, failed, quarantined, or independently cold targets.
#### Scenario: Selected image already completed
- **WHEN** a resolver selection conflicts with an immutable target already in done state
- **THEN** the system SHALL record hashed skip evidence and deterministically continue to the next eligible fresh graph/alias without requeueing the completed target or retrying the identical selected set
#### Scenario: Selected image is quarantined
- **WHEN** a resolver selection conflicts with quarantined work
- **THEN** the system SHALL skip it and continue to the next eligible fresh candidate or replacement repository and SHALL NOT bypass quarantine
#### Scenario: No safe replacement remains
- **WHEN** every fresh candidate is terminal, independently cold, quarantined, fenced, or otherwise ineligible
- **THEN** only that repository membership SHALL become terminal `skipped` with exact hashed image-unavailability evidence, no experiment target or capacity slot SHALL be consumed, and no candidate SHALL be reactivated
#### Scenario: Terminal scarcity evidence drifts
- **WHEN** terminal repository skip state lacks its exact reason, repository identity, ordinal, or canonical evidence hash
- **THEN** experiment authority validation SHALL move the experiment to held state before activation or further mutation
### Requirement: Durable manifest and layer attribution
The system SHALL persist each selected immutable manifest and its ordered layer descriptors, including positions from base and top, and SHALL associate findings only when an exact layer digest is present.
#### Scenario: Exact layer digest finding
- **WHEN** a finding reports a Docker layer digest present at one or more manifest positions
- **THEN** the system SHALL persist every exact matching base/top position for that image
#### Scenario: Finding has no layer digest
- **WHEN** scanner evidence does not identify an exact layer digest
- **THEN** the report SHALL count the finding as layer-unattributed and SHALL NOT infer a position
### Requirement: Reservation-bound experiment evidence
The system SHALL bind experiment targets to scan reservations atomically and SHALL use those bindings to exclude historical or unrelated scans from experiment results.
#### Scenario: Experiment target is reserved
- **WHEN** an experiment target receives scan capacity and a queue lease
- **THEN** its experiment binding, reservation, and queue transition SHALL commit atomically
#### Scenario: Reservation is retried
- **WHEN** the same experiment target requires a fenced retry
- **THEN** reporting SHALL preserve each bound attempt while deduplicating final physical target totals
#### Scenario: Admission capacity is saturated
- **WHEN** experiment admission cannot reserve pipeline or quarantine capacity
- **THEN** the aborted admission intent and experiment `held` transition SHALL commit atomically under the same experiment authority lock
### Requirement: Safe experiment reporting
The system SHALL produce aggregate physical and per-query reports comparing image ranks 1-3 with ranks 4-10, including deduplicated findings, credential identities, keycheck outcomes, scan duration/errors, layer positions, overlap, coverage, and marginal minimum-rank yield without exposing secret material.
#### Scenario: Image belongs to multiple queries
- **WHEN** one physical image selection is attributed to multiple queries
- **THEN** global totals SHALL count it once while each relevant query report SHALL receive attribution and overlap SHALL be explicit
#### Scenario: Identity repeats at later rank
- **WHEN** a detector-secret or credential identity first appears at rank 2 and appears again at rank 7
- **THEN** marginal yield SHALL assign that identity to rank 2 and SHALL NOT recount it as new in ranks 4-10
#### Scenario: Report contains sensitive evidence
- **WHEN** aggregate reporting reads findings or keycheck records
- **THEN** output SHALL contain only approved IDs, hashes, enums, counts, durations, and positions and SHALL NOT contain raw credentials or secret-bearing excerpts
### Requirement: Fail-closed experiment authority
Experiment collection and activation SHALL run only on managed PostgreSQL final-cutover with search-mode immutable-digest discovery. The canonical semantic configuration hash SHALL exclude operational `enabled` and include the pinned collection generation, exact ordered effective query policies, and Docker platform filter, OS, architecture, and candidate-count values. Enabled state SHALL be validated independently. The original cohort plan hash SHALL remain immutable across deterministic exclusions/replacements, and a second frozen runtime-selection hash SHALL bind effective repositories, terminal image-scarcity evidence, and the executed deep-probe choice before activation. Experiment authority validation and every experiment-owned mutation SHALL use an experiment-row-first locked protocol. An active experiment SHALL transition to held state when ordered queries, selector version, configuration hash, runtime-selection hash, terminal skip evidence, owned hold-event/queue history, capacity, or fencing invariants drift.
#### Scenario: Query list changes during execution
- **WHEN** the configured ordered query hash differs from the pinned experiment hash
- **THEN** new experiment claims SHALL stop and the experiment SHALL enter held state
#### Scenario: File-queue fallback is active
- **WHEN** PostgreSQL final-cutover authority is unavailable
- **THEN** experiment activation and mutation SHALL be rejected
#### Scenario: Executed deep probe drifts
- **WHEN** an executed `is_deep_probe` choice differs from the frozen runtime-selection hash
- **THEN** new claims SHALL stop and the experiment SHALL enter held state
#### Scenario: Disabled during holding
- **WHEN** operational `enabled` becomes false while the experiment is in holding
- **THEN** the experiment SHALL transition fail-closed to held
#### Scenario: Long remote graph resolution
- **WHEN** candidate manifest resolution spans the original 300-second lease
- **THEN** the worker SHALL renew before and between bounded remote stages under its exact owner/generation/token, verify the token after remote work, and perform no completion write after lease loss
#### Scenario: Runtime stops during resolver work
- **WHEN** authority validation finds an expired resolver fence left by an interrupted runtime
- **THEN** it SHALL atomically return the affected membership to pending and enter `held(stale_resolver_fence)`, and a later claim MAY resume `resolving` only after full authority validation succeeds and no resolver fence remains; no other held reason SHALL automatically resume
#### Scenario: Repeated non-conclusive remote failures
- **WHEN** a resolver membership consumes three non-conclusive remote attempts
- **THEN** only that membership SHALL terminate instead of retrying forever or pinning the breadth barrier
- **AND** breadth work SHALL record exact hashed `remote_unavailable_after_attempt_limit` scarcity without a target slot, while deep work SHALL preserve already selected images and close at its achieved depth
#### Scenario: Systemic resolver conflict reaches its ceiling
- **WHEN** repeated evidence, capacity, configuration, or authority conflicts reach their safety ceiling
- **THEN** the experiment SHALL remain fail-closed in held state and SHALL NOT misclassify the conflict as target-local remote scarcity
#### Scenario: Legacy attempt-limit hold resumes under the simplified policy
- **WHEN** a fully validated claim encounters exactly one unfenced membership persisted as `held(resolver_attempt_limit)` by the earlier policy
- **THEN** it SHALL terminalize only that membership with exact hashed remote-unavailable evidence, clear the legacy experiment hold, and continue resolving the remaining cohort
#### Scenario: Reviewed disposition of a genuine attempt-limit hold
- **WHEN** an operator reviews an exact hash-only manifest for the single membership held by genuine remote failures and approves its SHA while sources are stopped
- **THEN** the system SHALL atomically use the first unchanged eligible fresh replacement repository and reset only that membership's attempts, or SHALL terminally record exact `remote_unavailable_after_attempt_limit` scarcity when the reviewed fresh replacement pool is exhausted
- **AND** it SHALL append a one-time immutable audit row, resume the experiment, preserve the global three-attempt policy, consume no target slot for a skip, and never expose or reactivate historical target identity
#### Scenario: Reviewed refund of attempts consumed by a retired local defect
- **WHEN** stopped-source offline review proves exactly two target-bound occurrences of the retired zero-graph limit defect for a membership and an operator approves the exact private manifest SHA
- **THEN** the system SHALL refund exactly two attempts, append a one-time immutable audit row, clear only that membership's obsolete local error, and resume the attempt-limit-held experiment without exposing the target identity
- **AND** memberships containing only partial or other remote failures SHALL remain unchanged, and the same recovery kind SHALL never refund a membership twice
#### Scenario: Owned hold history drifts
- **WHEN** an experiment-owned hold event, queue state, config/policy/manifest hash, audit hash, or reversal is changed outside the reviewed protocol
- **THEN** continuous authority validation SHALL stop mutation and hold the experiment before new rows or events commit
@@ -0,0 +1,52 @@
## 1. Schema And Migration
- [x] 1.1 Add migration marker and SQLite/PostgreSQL schema for Docker query provenance, immutable manifests/layers, experiment authority, query/repository/selection/target/binding relations, and finding-layer attribution
- [x] 1.2 Register required columns, keys, foreign keys, indexes, integer-width conversions, constraints, migration quiescence checks, and idempotent validation
- [x] 1.3 Add legacy first-inserter provenance seeding that is explicitly ineligible for fresh experiment coverage
## 2. Configuration And Selection
- [x] 2.1 Add strict side-effect-free Docker image-depth and experiment configuration validation with exact ordered-query/config/selector hashes
- [x] 2.2 Remove the hidden three-image clamp and implement deterministic distinct graph selection through rank ten while preserving first-three semantics
- [x] 2.3 Add production configuration for the disabled collection rollout and the pinned 61-query, up-to-10-fresh-repository, 10-image, 1,200-target experiment
## 3. Discovery Provenance
- [x] 3.1 Persist main-pass query-to-repository provenance and stable search rank atomically with each Docker discovery page
- [x] 3.2 Persist retry-pass provenance with the original query/policy evidence under retry lease fencing
- [x] 3.3 Require one complete pinned deep pass, then freeze each of all 61 queries with its exact zero-to-ten eligible unseen repository count without historical substitution
## 4. Cohort And Holds
- [x] 4.1 Implement deterministic sparse round-robin cohort planning, one deep-probe choice per nonempty query, immutable desired/actual count hashes, and serialized unique-target capacity accounting
- [x] 4.2 Generate and validate an explicit experiment hold manifest for non-cohort unresolved Docker anchors
- [x] 4.3 Apply/release experiment-owned cold events through existing fence-aware APIs without changing independently fenced targets
- [x] 4.4 Hold newly observed non-cohort anchors under the activated reviewed experiment policy
## 5. Resolver And Scheduling
- [x] 5.1 Add a dedicated fenced experiment resolver lane ordered by repository round and query ordinal
- [x] 5.2 Persist selected immutable targets, ranks/reasons, manifest identity, and ordered layer descriptors atomically with resolver completion
- [x] 5.3 Reject or replace conflicts with done, failed, quarantined, or independently cold immutable targets without implicit reactivation
- [x] 5.4 Bind experiment targets atomically to scan reservations and dispatch breadth, ranks 2-3, then ranks 4-10 round-robin
- [x] 5.5 Fail closed to held state on query/config/selector drift, capacity conflict, stale token, or non-PostgreSQL authority
## 6. Evidence And Reporting
- [x] 6.1 Associate exact Docker finding layer digests with all matching base/top manifest positions during atomic result ingestion
- [x] 6.2 Add a secret-safe aggregate experiment report separating physical totals from per-query attribution and ranks 1-3 from ranks 4-10
- [x] 6.3 Report deduplicated findings/credentials, frozen and current verification outcomes, marginal minimum-rank yield, scan duration/errors, overlap, coverage, and unattributed findings
## 7. Verification
- [x] 7.1 Add focused unit tests for strict configuration, side-effect ordering, rank 4-10 selection, deduplication, and deterministic tie breaks
- [x] 7.2 Add schema/migration and PostgreSQL integration tests for provenance atomicity, fair planning, global cap concurrency, hold/reactivation fencing, resolver completion, reservation binding, and ingestion
- [x] 7.3 Add reporting and secret-sentinel tests for shared attribution, rank boundaries, marginal identity yield, layer positions, retries, and unattributed evidence
- [x] 7.4 Run focused tests, PostgreSQL integration, broad relevant suites, strict OpenSpec validation, and independent review
## 8. Controlled Rollout
- [x] 8.1 Canonically stop runtime, apply the additive migration with experiment disabled, and restart to collect fresh provenance
- [x] 8.2 Verify one complete fresh observation pass for every query, then canonically stop and generate/review the cohort and hold manifest
- [x] 8.3 Activate the frozen experiment, restart canonically, and verify PostgreSQL/pipeline/auth/worker health, fair cohort dispatch, capacity ceiling, leases, retries, restart counters, and bytecode absence
- [x] 8.4 Leave the experiment running toward drain/report while keeping non-cohort rows reversibly cold for a later reviewed decision