Files
2026-09-30 20:30:56 +03:00

92 lines
11 KiB
Markdown

## Context
Managed DockerHub repository search is authenticated and bounded to two GET attempts per page, but standard discovery still builds one all-pages result before database admission. An unavailable expected page therefore discards successful pages and leaves the numeric query cursor on the same keyword. Database deduplication happens only after pagination, so known repositories cannot currently stop deeper requests.
DockerHub search results are ordered but not snapshot-stable. A delayed retry of page N cannot reconstruct the exact earlier result window, so the design must combine idempotent page admission with periodic deep coverage rather than claim snapshot completeness. Repository anchors and immutable digest scan targets already have durable PostgreSQL identities, but none of the existing scan, resolver, projection, or source-cycle tables is a valid leaseable queue for failed discovery-page work.
The active source uses a single DockerHub writer, a numeric main query cursor, and additive JSON state. The new query list is appended to preserve that cursor. Runtime configuration and application files may only be changed while the canonical supervisor is stopped.
## Goals / Non-Goals
**Goals:**
- Persist each valid DockerHub page before requesting or acting on later pages.
- Stop ordinary pagination after two consecutive nonempty pages containing only identities known before the current pass.
- Use a 30-page, 100-result ceiling for every normal and deep query pass.
- Give each exact query a deep pass that bypasses seen-page stopping when 72 hours have elapsed since its last durable dispatch.
- Delegate unavailable page/query work to a durable, bounded, fenced retry queue without blocking main keyword rotation.
- Add the confirmed 12 product/framework queries and disable only periodic re-resolution of completed repository anchors.
- Preserve authenticated fail-closed requests, the existing two-GET page budget, cold/failed target exclusion, and immutable-digest error retries.
**Non-Goals:**
- Treating DockerHub pagination as a stable snapshot or guaranteeing successful external acquisition every 72 hours.
- Re-enabling cold/failed targets, rescanning successful immutable digests, or changing tag/layer selection, keychecks, scan workers, or scan budgets.
- Reusing scan/resolver queues for HTTP discovery work.
- Changing GitHub, GitLab, HuggingFace, or other source behavior.
## Decisions
### Admit repository pages incrementally
The managed PostgreSQL DockerHub path will fetch page one first, validate `count`, and process the expected range sequentially. A narrow database operation will normalize and deduplicate the page, determine preexisting identities, insert bare repository anchors using the existing unresolved-anchor semantics, and return safe counts plus internal normalized identities needed for the current-pass knownness decision. It will not invoke resolver or scan claiming. The existing resolver gate runs once after pagination, not once per page.
Sequential acquisition is selected over the current parallel remainder because requests for page three and beyond must be avoidable after pages one and two prove fully known. It also bounds in-memory results and makes page-level durability explicit. The authenticated page helper and its two-total-GET budget remain unchanged.
Knownness is evaluated before page admission. A page increments the streak only when it is nonempty and every normalized repository existed before the pass. Identities inserted on an earlier page in the same pass do not count as preexisting if they appear again after result movement. Empty pages end the available range. Lookup uncertainty fails open by resetting the streak and continuing deeper.
Before each managed pass, source state records an incomplete marker keyed by exact query and effective policy hash. A matching marker forces deep behavior on the next attempt and is cleared only after a durable `completed`, `completed_with_retries`, or `query_invalid` outcome. This prevents a failed pass from treating pages it admitted before the failure as old-enough evidence for an early stop after restart.
### Delegate gaps to a PostgreSQL discovery retry queue
Add a dedicated `discovery_retry_queue`; existing target, resolver, projection, outbox, and source-cycle tables have incompatible lifecycle and foreign-key semantics. Work is coalesced by a deterministic key over source, exact query, effective policy hash, pass kind, and page/range. Rows move through `pending -> leased -> deleted`, with retryable failure returning to `pending`, expired leases reclaimable, and removed/mismatched policy work moved to `held`.
Claims use bounded `FOR UPDATE SKIP LOCKED` selection and random lease tokens. Completion and retry transitions require the exact row, owner, and token. A stale worker cannot acknowledge or replace newer work. The worker renews the same fenced lease before each page in a multi-page retry, so a bounded per-page lease cannot expire merely because an entire range takes longer than one lease interval. Retryable failures use exponential backoff with a cap and are never silently dropped at a maximum attempt count. Provider cooldown refunds the dispatch attempt only when the claim has made no remote request and uses the trusted retry time. Diagnostics store fixed safe categories, never authorization material or arbitrary response bodies.
If page one is unavailable, enqueue query-level work because the expected range is unknown. If a later page fails, enqueue page work and continue when account availability permits. If the account pool is exhausted before a remaining tail can be attempted, coalesce the tail into range work instead of creating one row per unattempted page. Successful pages remain admitted. The source cycle may advance its main query only after every observed gap is durably delegated; retry persistence failure keeps the old cursor and marks the cycle failed.
At most one due retry item is processed per normal source-loop iteration, independently of the main cursor. Main query state is saved before retry work can affect the next iteration. Retry success admits repositories before fenced acknowledgement; retry failure changes only retry state. This prevents a persistent page outage from head-of-line blocking the 61-keyword rotation.
### Represent partial success explicitly
Add `completed_with_retries` as a source-cycle outcome for a pass whose successful pages were admitted and whose gaps were durably delegated. The numeric query cursor advances for `completed`, `completed_with_retries`, and `query_invalid`, but not for `failed`, `source_failed`, or `backlog_only`. Invalid payloads, database admission failure, and retry-enqueue failure remain hard cycle failures rather than endlessly retryable transport work.
Alternative considered: keep the keyword pinned while retaining successful pages. Rejected because it preserves head-of-line blocking. Alternative considered: advance after logging a failed page without durable work. Rejected because a moving search window can make the gap unrecoverable.
### Schedule deep passes by exact query
DockerHub source state gains a versioned map keyed by exact query, not numeric list position. Each record contains the effective search policy hash and UTC deep-dispatch timestamp. A query is deep-due when the record is absent, malformed, policy-mismatched, or at least 72 hours old. Its next main rotation pass ignores only the two-known-page stop; page-one count, 30-page cap, 100-result size, authentication, retries, and incremental admission remain mandatory.
The dispatch timestamp is written only after successful page admission and durable delegation of any gaps. Scheduling from dispatch time avoids deep-pass storms during prolonged provider failure. Appended queries are immediately due without invalidating existing query timestamps. Removed queries are pruned from state and their retry work is held. This provides one deep dispatch per exact query on its first normal selection after the 72-hour boundary; it is not an external-success SLA.
State-file loss causes conservative early deep passes, not missed passes. PostgreSQL remains authoritative for retry work and repository identity.
### Apply the confirmed discovery policy
Set DockerHub defaults to 30 pages and 100 results per page and align existing per-query page overrides so every query has the same acquisition ceiling. Append exactly the 12 confirmed terms, preserving existing order and numeric cursor safety.
Set `docker_repository_refresh_max_per_cycle` to zero. This disables periodic claims of completed/resolved repository anchors only. `refresh_registry` remains enabled, so keyword search still runs; pending initial anchors and partial/error resolver rows remain eligible; new immutable digests discovered through those paths are scanned; immutable scan errors retain their existing retries.
## Risks / Trade-offs
- [The first pass of 12 new terms can inspect up to 36,000 search rows and create a large resolver backlog] -> Keep existing resolver, scan-worker, admission, and pipeline bounds; append terms rather than force a manual pass; observe backlog after natural rotation.
- [Two known pages do not prove all deeper pages are known] -> Treat stopping as an optimization and bypass it per exact query every 72 hours.
- [A delayed page retry sees a moving result window] -> Admit retries idempotently and rely on recurring deep passes for eventual coverage rather than snapshot claims.
- [A retry table can grow during a prolonged outage] -> Coalesce deterministic work keys, use range rows for unattempted tails, bound claims per loop, expose pending/oldest-age metrics, and never advance when durable delegation itself fails.
- [Sequential pages increase latency for a truly deep pass] -> Normal passes usually stop early; deep passes deliberately trade latency for bounded coverage and remain capped at 30 pages.
- [Disabling completed-anchor refresh can miss a new digest in a repository that no longer appears in search] -> This is the user's temporary policy choice; initial and failed resolution continue, and the switch can be restored independently.
- [JSON deep-state corruption can trigger extra load] -> Validate schema strictly, key by exact query/policy, and fail toward an early deep pass rather than suppressing coverage.
## Migration Plan
1. Create and validate the OpenSpec artifacts and deterministic tests before touching the live runtime.
2. Canonically stop the supervisor and verify all managed workers and PostgreSQL are down.
3. Apply the additive PostgreSQL schema migration, scanner/runner orchestration, source-state validation, tests, and configuration changes.
4. Run focused unit/SQL/migration/integration tests, then the broader relevant scanner and runtime-safety suites.
5. Canonically start the supervisor; verify PostgreSQL migration authority, pipeline readiness, DockerHub authentication, worker restart counters, retry/deep-state health, and absence of bytecode artifacts.
6. Roll back under another canonical stop by restoring code/config and leaving the additive retry table dormant; repository and immutable-target inserts are idempotent and require no destructive rollback.
## Open Questions
None. Keyword scope, 30x100 policy, 72-hour deep cadence, disabled completed-anchor refresh, preserved scan retries, and independent retry backlog were explicitly confirmed.