Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-09-09
@@ -0,0 +1,91 @@
## Context
Managed DockerHub repository search is authenticated and bounded to two GET attempts per page, but standard discovery still builds one all-pages result before database admission. An unavailable expected page therefore discards successful pages and leaves the numeric query cursor on the same keyword. Database deduplication happens only after pagination, so known repositories cannot currently stop deeper requests.
DockerHub search results are ordered but not snapshot-stable. A delayed retry of page N cannot reconstruct the exact earlier result window, so the design must combine idempotent page admission with periodic deep coverage rather than claim snapshot completeness. Repository anchors and immutable digest scan targets already have durable PostgreSQL identities, but none of the existing scan, resolver, projection, or source-cycle tables is a valid leaseable queue for failed discovery-page work.
The active source uses a single DockerHub writer, a numeric main query cursor, and additive JSON state. The new query list is appended to preserve that cursor. Runtime configuration and application files may only be changed while the canonical supervisor is stopped.
## Goals / Non-Goals
**Goals:**
- Persist each valid DockerHub page before requesting or acting on later pages.
- Stop ordinary pagination after two consecutive nonempty pages containing only identities known before the current pass.
- Use a 30-page, 100-result ceiling for every normal and deep query pass.
- Give each exact query a deep pass that bypasses seen-page stopping when 72 hours have elapsed since its last durable dispatch.
- Delegate unavailable page/query work to a durable, bounded, fenced retry queue without blocking main keyword rotation.
- Add the confirmed 12 product/framework queries and disable only periodic re-resolution of completed repository anchors.
- Preserve authenticated fail-closed requests, the existing two-GET page budget, cold/failed target exclusion, and immutable-digest error retries.
**Non-Goals:**
- Treating DockerHub pagination as a stable snapshot or guaranteeing successful external acquisition every 72 hours.
- Re-enabling cold/failed targets, rescanning successful immutable digests, or changing tag/layer selection, keychecks, scan workers, or scan budgets.
- Reusing scan/resolver queues for HTTP discovery work.
- Changing GitHub, GitLab, HuggingFace, or other source behavior.
## Decisions
### Admit repository pages incrementally
The managed PostgreSQL DockerHub path will fetch page one first, validate `count`, and process the expected range sequentially. A narrow database operation will normalize and deduplicate the page, determine preexisting identities, insert bare repository anchors using the existing unresolved-anchor semantics, and return safe counts plus internal normalized identities needed for the current-pass knownness decision. It will not invoke resolver or scan claiming. The existing resolver gate runs once after pagination, not once per page.
Sequential acquisition is selected over the current parallel remainder because requests for page three and beyond must be avoidable after pages one and two prove fully known. It also bounds in-memory results and makes page-level durability explicit. The authenticated page helper and its two-total-GET budget remain unchanged.
Knownness is evaluated before page admission. A page increments the streak only when it is nonempty and every normalized repository existed before the pass. Identities inserted on an earlier page in the same pass do not count as preexisting if they appear again after result movement. Empty pages end the available range. Lookup uncertainty fails open by resetting the streak and continuing deeper.
Before each managed pass, source state records an incomplete marker keyed by exact query and effective policy hash. A matching marker forces deep behavior on the next attempt and is cleared only after a durable `completed`, `completed_with_retries`, or `query_invalid` outcome. This prevents a failed pass from treating pages it admitted before the failure as old-enough evidence for an early stop after restart.
### Delegate gaps to a PostgreSQL discovery retry queue
Add a dedicated `discovery_retry_queue`; existing target, resolver, projection, outbox, and source-cycle tables have incompatible lifecycle and foreign-key semantics. Work is coalesced by a deterministic key over source, exact query, effective policy hash, pass kind, and page/range. Rows move through `pending -> leased -> deleted`, with retryable failure returning to `pending`, expired leases reclaimable, and removed/mismatched policy work moved to `held`.
Claims use bounded `FOR UPDATE SKIP LOCKED` selection and random lease tokens. Completion and retry transitions require the exact row, owner, and token. A stale worker cannot acknowledge or replace newer work. The worker renews the same fenced lease before each page in a multi-page retry, so a bounded per-page lease cannot expire merely because an entire range takes longer than one lease interval. Retryable failures use exponential backoff with a cap and are never silently dropped at a maximum attempt count. Provider cooldown refunds the dispatch attempt only when the claim has made no remote request and uses the trusted retry time. Diagnostics store fixed safe categories, never authorization material or arbitrary response bodies.
If page one is unavailable, enqueue query-level work because the expected range is unknown. If a later page fails, enqueue page work and continue when account availability permits. If the account pool is exhausted before a remaining tail can be attempted, coalesce the tail into range work instead of creating one row per unattempted page. Successful pages remain admitted. The source cycle may advance its main query only after every observed gap is durably delegated; retry persistence failure keeps the old cursor and marks the cycle failed.
At most one due retry item is processed per normal source-loop iteration, independently of the main cursor. Main query state is saved before retry work can affect the next iteration. Retry success admits repositories before fenced acknowledgement; retry failure changes only retry state. This prevents a persistent page outage from head-of-line blocking the 61-keyword rotation.
### Represent partial success explicitly
Add `completed_with_retries` as a source-cycle outcome for a pass whose successful pages were admitted and whose gaps were durably delegated. The numeric query cursor advances for `completed`, `completed_with_retries`, and `query_invalid`, but not for `failed`, `source_failed`, or `backlog_only`. Invalid payloads, database admission failure, and retry-enqueue failure remain hard cycle failures rather than endlessly retryable transport work.
Alternative considered: keep the keyword pinned while retaining successful pages. Rejected because it preserves head-of-line blocking. Alternative considered: advance after logging a failed page without durable work. Rejected because a moving search window can make the gap unrecoverable.
### Schedule deep passes by exact query
DockerHub source state gains a versioned map keyed by exact query, not numeric list position. Each record contains the effective search policy hash and UTC deep-dispatch timestamp. A query is deep-due when the record is absent, malformed, policy-mismatched, or at least 72 hours old. Its next main rotation pass ignores only the two-known-page stop; page-one count, 30-page cap, 100-result size, authentication, retries, and incremental admission remain mandatory.
The dispatch timestamp is written only after successful page admission and durable delegation of any gaps. Scheduling from dispatch time avoids deep-pass storms during prolonged provider failure. Appended queries are immediately due without invalidating existing query timestamps. Removed queries are pruned from state and their retry work is held. This provides one deep dispatch per exact query on its first normal selection after the 72-hour boundary; it is not an external-success SLA.
State-file loss causes conservative early deep passes, not missed passes. PostgreSQL remains authoritative for retry work and repository identity.
### Apply the confirmed discovery policy
Set DockerHub defaults to 30 pages and 100 results per page and align existing per-query page overrides so every query has the same acquisition ceiling. Append exactly the 12 confirmed terms, preserving existing order and numeric cursor safety.
Set `docker_repository_refresh_max_per_cycle` to zero. This disables periodic claims of completed/resolved repository anchors only. `refresh_registry` remains enabled, so keyword search still runs; pending initial anchors and partial/error resolver rows remain eligible; new immutable digests discovered through those paths are scanned; immutable scan errors retain their existing retries.
## Risks / Trade-offs
- [The first pass of 12 new terms can inspect up to 36,000 search rows and create a large resolver backlog] -> Keep existing resolver, scan-worker, admission, and pipeline bounds; append terms rather than force a manual pass; observe backlog after natural rotation.
- [Two known pages do not prove all deeper pages are known] -> Treat stopping as an optimization and bypass it per exact query every 72 hours.
- [A delayed page retry sees a moving result window] -> Admit retries idempotently and rely on recurring deep passes for eventual coverage rather than snapshot claims.
- [A retry table can grow during a prolonged outage] -> Coalesce deterministic work keys, use range rows for unattempted tails, bound claims per loop, expose pending/oldest-age metrics, and never advance when durable delegation itself fails.
- [Sequential pages increase latency for a truly deep pass] -> Normal passes usually stop early; deep passes deliberately trade latency for bounded coverage and remain capped at 30 pages.
- [Disabling completed-anchor refresh can miss a new digest in a repository that no longer appears in search] -> This is the user's temporary policy choice; initial and failed resolution continue, and the switch can be restored independently.
- [JSON deep-state corruption can trigger extra load] -> Validate schema strictly, key by exact query/policy, and fail toward an early deep pass rather than suppressing coverage.
## Migration Plan
1. Create and validate the OpenSpec artifacts and deterministic tests before touching the live runtime.
2. Canonically stop the supervisor and verify all managed workers and PostgreSQL are down.
3. Apply the additive PostgreSQL schema migration, scanner/runner orchestration, source-state validation, tests, and configuration changes.
4. Run focused unit/SQL/migration/integration tests, then the broader relevant scanner and runtime-safety suites.
5. Canonically start the supervisor; verify PostgreSQL migration authority, pipeline readiness, DockerHub authentication, worker restart counters, retry/deep-state health, and absence of bytecode artifacts.
6. Roll back under another canonical stop by restoring code/config and leaving the additive retry table dormant; repository and immutable-target inserts are idempotent and require no destructive rollback.
## Open Questions
None. Keyword scope, 30x100 policy, 72-hour deep cadence, disabled completed-anchor refresh, preserved scan retries, and independent retry backlog were explicitly confirmed.
@@ -0,0 +1,31 @@
## Why
DockerHub discovery currently downloads its full configured page range before database deduplication, and one exhausted page discards every successful page while blocking keyword rotation. The authenticated 30-page window makes deeper discovery possible, but it needs incremental persistence, seen-page stopping, and durable retry delegation to remain efficient and avoid silent gaps or head-of-line blocking.
## What Changes
- **BREAKING**: replace complete-or-fail DockerHub pagination with page-level durable persistence; successful pages remain admitted when another page exhausts its two request attempts.
- Stop an ordinary query pass after two consecutive nonempty pages whose repository identities were already known before the pass.
- Run every DockerHub query with an effective ceiling of 30 pages and 100 results per page.
- Bypass seen-page stopping for a full deep pass of each exact query at least once per 72-hour scheduling interval.
- Delegate failed page/query acquisition to a fenced PostgreSQL discovery retry backlog before allowing the main keyword rotation to advance.
- Add 12 DockerHub product/framework queries: `open-webui`, `ragflow`, `dify`, `flowise`, `crewai`, `n8n`, `langflow`, `autogen`, `browser-use`, `openhands`, `anythingllm`, and `agent-zero`.
- Temporarily disable periodic re-resolution of completed DockerHub repository anchors while preserving initial resolution, partial/error resolver retries, immutable-digest scan retries, and normal keyword discovery.
- Preserve authenticated fail-closed search, the two-GET page budget, target uniqueness, cold/failed target policy, and secret-safe diagnostics.
## Capabilities
### New Capabilities
- `dockerhub-incremental-discovery`: Incremental DockerHub page admission, seen-page stopping, 72-hour deep passes, and durable failed-page retry work.
### Modified Capabilities
None.
## Impact
- Affects DockerHub pagination and authentication integration in `app/scanner.py`, source-cycle orchestration/state in `app/console_runner.py`, PostgreSQL schema and retry claims in `app/scanner_db.py`, and DockerHub settings in `app/config.yaml`.
- Adds a PostgreSQL discovery retry queue with bounded leases, fencing, backoff, and configured-query/policy validation.
- Changes DockerHub query rotation from failure-blocking to durable retry delegation and adds an initial bounded backlog from 12 new deep searches.
- Requires focused unit, SQL-shape, migration, PostgreSQL integration, source-state, configuration, and runtime health verification.
- Does not add dependencies or credential formats and does not change Docker tag selection, Registry authentication, layer scanning, scan workers, keychecks, or immutable-digest retry policy.
@@ -0,0 +1,164 @@
## ADDED Requirements
### Requirement: DockerHub pages are admitted incrementally
The system SHALL validate and durably admit every successful DockerHub repository-search page before relying on later-page acquisition, while preserving target identity and cold/failed-target exclusion.
#### Scenario: Successful page precedes a later failure
- **WHEN** an expected DockerHub page is valid and a later expected page exhausts its request budget
- **THEN** repositories from the successful page remain idempotently admitted and the later failure cannot roll them back
#### Scenario: Page admission fails
- **WHEN** repository normalization or durable page admission fails
- **THEN** the source cycle fails, does not treat the page as complete, and does not advance the main query cursor
#### Scenario: Resolver processing follows pagination
- **WHEN** one or more pages admit repository anchors
- **THEN** the existing Docker resolver gate runs at most once after the pass rather than once per page
### Requirement: Ordinary discovery stops after consecutive known pages
The system SHALL stop an ordinary DockerHub query before requesting another page after two consecutive nonempty pages contain only repository identities that existed before the current pass.
#### Scenario: First two pages were previously known
- **WHEN** pages one and two are nonempty and every normalized repository identity existed before the pass
- **THEN** the pass completes without requesting page three
#### Scenario: Page contains a new repository
- **WHEN** either of the last two pages contains a repository identity not known before the pass
- **THEN** the consecutive-known counter resets and discovery continues within the effective range
#### Scenario: Current-pass duplicate appears later
- **WHEN** a repository first admitted earlier in the same pass appears on a later page
- **THEN** that identity is not treated as preexisting evidence for the later page's known-page stop
#### Scenario: Knownness cannot be determined
- **WHEN** the database cannot establish complete pre-pass knownness for a valid page
- **THEN** discovery fails open by continuing deeper rather than stopping early
#### Scenario: Previous pass ended incompletely
- **WHEN** an exact query and policy have a retained incomplete-pass marker
- **THEN** the next selected pass bypasses known-page stopping until a durable pass outcome clears the marker
### Requirement: Deep discovery bypasses seen-page stopping every 72 hours
The system SHALL make each exact configured DockerHub query deep-due no later than 72 hours after its last durable deep dispatch and SHALL bypass the consecutive-known-page stop on that query's next selected pass.
#### Scenario: Exact query reaches its deep interval
- **WHEN** at least 72 hours have elapsed since the query's last durable deep dispatch
- **THEN** its next selected pass processes the available expected range without stopping on known pages
#### Scenario: Query is new or policy changes
- **WHEN** an exact configured query has no valid state or its effective search-policy hash changes
- **THEN** that query is immediately deep-due without resetting unrelated query schedules
#### Scenario: Deep pass has durable gaps
- **WHEN** valid pages are admitted and unavailable pages are durably delegated to retry work
- **THEN** the deep dispatch timestamp advances while completion remains represented by the outstanding retry work
#### Scenario: Provider prevents acquisition
- **WHEN** neither successful pages nor durable retry delegation can establish a deep dispatch
- **THEN** the system does not advance that query's deep timestamp
### Requirement: DockerHub page gaps use a durable retry backlog
The system SHALL persist unavailable DockerHub query/page work in a dedicated PostgreSQL retry backlog before allowing the main keyword rotation to advance.
#### Scenario: Page one is unavailable
- **WHEN** page one exhausts its bounded request/account handling and the expected range is unknown
- **THEN** the system durably enqueues query-level retry work
#### Scenario: Later page is unavailable
- **WHEN** a later expected page exhausts its bounded request/account handling
- **THEN** the system durably enqueues page-level work and retains every successfully admitted page
#### Scenario: Remaining tail cannot be attempted
- **WHEN** provider/account exhaustion prevents remote attempts for a known remaining page range
- **THEN** the system coalesces the unattempted range into bounded retry work instead of creating unbounded individual rows
#### Scenario: Durable delegation succeeds
- **WHEN** every observed acquisition gap has durable retry work
- **THEN** the cycle records `completed_with_retries` and advances the main keyword independently of retry processing
#### Scenario: Durable delegation fails
- **WHEN** retry work cannot be persisted authoritatively
- **THEN** the cycle fails and retains the current keyword cursor
### Requirement: Discovery retry claims are bounded and fenced
The system SHALL coalesce retry work by source, exact query, effective policy, pass kind, and page/range; SHALL process bounded due work with expiring leases; and MUST reject stale acknowledgements.
#### Scenario: Concurrent workers claim due work
- **WHEN** multiple workers attempt to claim the same due retry row
- **THEN** at most one receives the active lease token
#### Scenario: Lease expires
- **WHEN** a worker fails to complete work before its lease expires
- **THEN** the row becomes reclaimable without deleting its attempt history
#### Scenario: Multi-page retry remains active
- **WHEN** a leased query or range retry is about to request another page
- **THEN** the worker renews the same owner-and-token fence before acquisition and stops if renewal is rejected
#### Scenario: Stale worker finishes
- **WHEN** a worker presents an obsolete owner or lease token
- **THEN** it cannot acknowledge, delete, defer, or hold the newer work
#### Scenario: Retryable acquisition fails again
- **WHEN** leased work encounters another retryable failure
- **THEN** it returns to pending with bounded exponential backoff and is not silently dropped at an attempt limit
#### Scenario: Provider cooldown follows partial remote progress
- **WHEN** an earlier page in the same claim made a remote request before a later local provider cooldown
- **THEN** the dispatch attempt is not refunded
#### Scenario: Query is removed or policy is obsolete
- **WHEN** retry work no longer matches an exact configured query and effective policy
- **THEN** it is held and cannot execute against stale discovery policy
### Requirement: Retry work does not block main keyword rotation
The system SHALL process at most a bounded amount of due discovery retry work per source-loop iteration independently of the saved main query cursor.
#### Scenario: Persistent page failure exists
- **WHEN** a retry row remains unavailable across multiple attempts
- **THEN** ordinary configured keywords continue rotating while the row follows its own backoff
#### Scenario: Retry succeeds
- **WHEN** leased page or range work returns valid repositories
- **THEN** repositories are durably admitted before the fenced retry acknowledgement
### Requirement: DockerHub search policy uses the confirmed breadth
The system SHALL use an effective ceiling of 30 pages and 100 results per page for every configured DockerHub query and SHALL include exactly the 12 confirmed new product/framework terms in addition to the existing ordered query set.
#### Scenario: Ordinary pass reaches known content
- **WHEN** a normal 30-by-100 query pass reaches two consecutive preexisting-known pages
- **THEN** it stops early despite the larger configured ceiling
#### Scenario: Deep pass remains novel
- **WHEN** a due deep pass continues to return pages containing new repositories
- **THEN** it processes up to the page-one result boundary or the 30-page safety cap
#### Scenario: Query list is expanded
- **WHEN** the configuration is loaded after deployment
- **THEN** `open-webui`, `ragflow`, `dify`, `flowise`, `crewai`, `n8n`, `langflow`, `autogen`, `browser-use`, `openhands`, `anythingllm`, and `agent-zero` each appear once at the ordered tail
### Requirement: Completed repository refresh is disabled without changing retries
The system SHALL disable periodic re-resolution of successfully completed DockerHub repository anchors while retaining normal search, initial/partial resolver work, and immutable-digest error retries.
#### Scenario: Completed anchor becomes periodically due
- **WHEN** a resolved repository anchor reaches its prior refresh interval
- **THEN** periodic policy does not claim it solely for refresh
#### Scenario: Initial or partial anchor is due
- **WHEN** an unresolved or retryable partial repository anchor is due
- **THEN** the existing resolver remains eligible to process it
#### Scenario: Immutable digest scan fails retryably
- **WHEN** an immutable Docker image scan meets the existing retry conditions
- **THEN** its current bounded retry and backoff behavior remains unchanged
### Requirement: Incremental discovery remains authenticated and secret-safe
The system MUST preserve explicit-pool fail-closed Hub bearer authentication and MUST NOT persist or emit credentials, bearer values, authorization headers, raw response bodies, arbitrary exception text, or repository targets through retry/deep-state diagnostics.
#### Scenario: Search or retry fails
- **WHEN** DockerHub page acquisition or retry processing reports an error
- **THEN** diagnostics contain only bounded status/category/count metadata required for operation
#### Scenario: Explicit account pool is unavailable
- **WHEN** no configured account can perform repository search
- **THEN** normal and retry acquisition fail closed without anonymous fallback
@@ -0,0 +1,30 @@
## 1. Runtime Safety
- [x] 1.1 Canonically stop the live supervisor and verify all managed workers and PostgreSQL are down before editing application or configuration files
## 2. Durable Discovery Storage
- [x] 2.1 Add and validate the additive PostgreSQL discovery retry queue schema, required indexes, migration metadata, and bounded lifecycle fields
- [x] 2.2 Implement strict page admission plus retry enqueue, claim, reclaim, backoff, hold, and fenced completion database operations
- [x] 2.3 Add SQL-shape and PostgreSQL migration/integration coverage for idempotency, concurrent claims, expired leases, and stale-token rejection
## 3. Incremental DockerHub Discovery
- [x] 3.1 Refactor managed DockerHub pagination to validate, deduplicate, and durably admit successful pages sequentially without invoking the resolver per page
- [x] 3.2 Implement two-consecutive-preexisting-page stopping with current-pass duplicate protection and fail-open knownness handling
- [x] 3.3 Implement per-query policy-hashed 72-hour deep scheduling that bypasses only seen-page stopping
- [x] 3.4 Delegate page-one, later-page, and unavailable-tail failures to durable retry work and advance the main cursor only after authoritative delegation
- [x] 3.5 Process bounded retry work independently of main rotation with lease fencing, backoff, policy/query validation, and safe diagnostics
## 4. Discovery Policy
- [x] 4.1 Configure every DockerHub query for a 30-page by 100-result ceiling and append the exact 12 confirmed product/framework queries once
- [x] 4.2 Disable periodic completed-anchor digest refresh while preserving initial/partial resolver and immutable-digest scan retries
- [x] 4.3 Add permanent configuration regressions for query order/count, effective breadth, disabled periodic refresh, and unchanged retry/resource policy
## 5. Verification And Deployment
- [x] 5.1 Run focused pagination, retry-queue, source-state, migration, and configuration tests without creating application bytecode
- [x] 5.2 Run broader relevant scanner, queue, runtime-safety, and PostgreSQL integration suites and complete an independent read-only review
- [x] 5.3 Strictly validate the OpenSpec change and verify no secret-bearing diagnostics or unrelated behavior changes
- [x] 5.4 Canonically start the runtime and verify PostgreSQL readiness, migrations, pipeline workers, DockerHub authentication/worker health, retry/deep state, restart counters, and absence of application bytecode