Initial server source import
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-09-09
|
||||
@@ -0,0 +1,70 @@
|
||||
## Context
|
||||
|
||||
Managed DockerHub source cycles configure an explicit account pool, but repository search bypasses it and calls the Hub search endpoint anonymously. Docker Hub rejects anonymous result windows beyond 200 entries; an isolated authenticated probe using the existing Hub bearer flow returned 100 results for pages 3, 20, 21, and 30 at a page size of 100. The standard search path also schedules every page before learning the reported result count and currently accepts successful pages when another expected page fails, so the source cycle completes and advances its numeric query cursor with incomplete discovery.
|
||||
|
||||
The account manager already provides secret-safe accounts, cached Hub bearer tokens, account rotation, endpoint-keyed cooldowns, and persisted auth events. The implementation must reuse those controls without changing tag resolution, Registry access, or immutable scan behavior.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Authenticate standard and recent DockerHub repository searches through the configured account pool.
|
||||
- Fail closed when an explicit pool has no usable account.
|
||||
- Support at most 30 authenticated search pages and avoid requests beyond the result count reported by page one.
|
||||
- Give each page exactly one additional attempt for transient transport or server failures.
|
||||
- Return a complete expected page set or fail the source cycle without enqueueing partial results or advancing its query cursor.
|
||||
- Keep logs, exceptions, and auth events free of credentials and bearer values.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Adding discovery queries, changing active page sizes, changing repository refresh cadence, or altering scan budgets.
|
||||
- Changing Docker tag selection, Registry bearer authentication, layer scanning, immutable target identity, or scan retry policy.
|
||||
- Guaranteeing that Docker Hub will indefinitely support 30 pages; the local cap remains a safety bound, not an external SLA.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Reuse Hub bearer authentication with a search-specific endpoint identity
|
||||
|
||||
Add a repository-search response helper that follows the existing tag-response pattern but uses the `hub_search` endpoint key. It obtains cached or fresh Hub bearer tokens, sends `Authorization: Bearer ...`, refreshes once after a 401, and rotates across the configured accounts for 401, 403, or 429 responses. Search cooldowns remain separate from `hub_tags` and `registry` cooldowns. The shared token cache remains per account because the same Hub token is valid for both Hub endpoints.
|
||||
|
||||
When the manager represents an explicit pool and no account is usable, the helper fails before making an anonymous request. A non-explicit legacy invocation may retain anonymous behavior for direct CLI compatibility.
|
||||
|
||||
Alternative considered: add a second login mechanism or cookie session. Rejected because the existing `/v2/auth/token` bearer flow was verified against authenticated search through page 30 and avoids another credential path.
|
||||
|
||||
### Fetch page one before parallel remainder
|
||||
|
||||
Raise the code-level page cap from 20 to 30. Standard discovery fetches page one first, validates its payload, derives the expected page count from its reported `count`, and submits only pages 2 through `min(requested, expected, 30)` concurrently. This removes ambiguous post-hoc suppression of failed pages and avoids requesting pages objectively outside the reported result set.
|
||||
|
||||
Alternative considered: keep launching all pages concurrently and ignore failures above the largest successful count. Rejected because a failed early page can make the inferred boundary unreliable, and unnecessary out-of-range requests consume account budget.
|
||||
|
||||
### Bound page retries inside the request primitive
|
||||
|
||||
The authenticated search helper gives each page at most two search GETs across transient retry, bearer refresh, and account rotation. Each GET calls `api_request` with one total network attempt; the outer page budget supplies the single bounded retry for `408`, `500`, `502`, `503`, `504`, or account-specific failures. A legacy anonymous invocation has no account handling and calls `api_request` with two total attempts directly. This keeps the page-wide budget independent of global proxy retry settings and prevents it from resetting during account rotation.
|
||||
|
||||
### Treat incomplete pagination as a source-cycle transport failure
|
||||
|
||||
Introduce a DockerHub discovery transport error analogous to the existing GitLab error. Any expected page that remains unavailable after its bounded request/account handling aborts the complete search result before tag resolution or enqueue. The configured-source runner records a failed cycle and returns without advancing the query cursor; it does not crash the long-running source process. A later cycle retries the same query, and queue uniqueness keeps successful rediscovery idempotent.
|
||||
|
||||
### Keep discovery authentication isolated from scan behavior
|
||||
|
||||
The change only replaces repository-search HTTP calls and their error propagation. Tag/manifest resolution continues to use its existing endpoint identities, retries, caches, resolver states, and immutable digest constraints.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Thirty authenticated pages increase search request volume] -> Preserve the hard cap, configured worker bounds, and account rotation; this change does not raise active source page settings.
|
||||
- [Page-one count can change while later pages are fetched] -> Treat page one as the cycle snapshot boundary; all pages within that boundary must still succeed, and DB deduplication handles overlap caused by result movement.
|
||||
- [A single page outage now rejects otherwise usable pages] -> This is intentional complete-or-fail behavior; one retry limits transient loss, and the unchanged cursor retries the query later.
|
||||
- [Fetching page one serially adds one request latency before parallel work] -> It prevents unnecessary pages and provides an authoritative expected set, which is more valuable than the small latency saving.
|
||||
- [All accounts can be temporarily unavailable] -> Fail closed and persist endpoint-specific auth state rather than silently reverting to the anonymous 200-result window.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Canonically stop the supervisor and verify all managed children are down.
|
||||
2. Deploy the scanner, runner, and focused regression tests without changing source query configuration.
|
||||
3. Run focused DockerHub authentication/pagination/cursor tests and the broader relevant scanner suites.
|
||||
4. Canonically restart the supervisor and verify PostgreSQL, pipeline workers, DockerHub source worker, and restart counters.
|
||||
5. Roll back by restoring the previous code under a canonical stop/start if authenticated search causes an operational regression; no data migration is required.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None. Authenticated access through page 30 and the explicit-pool behavior have been verified or are covered by deterministic tests.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
DockerHub repository discovery is currently anonymous even when a managed account pool is configured, so searches are limited to the anonymous 200-result window and partial page failures can silently advance the query rotation. Authenticated probing confirms the configured Hub bearer flow can retrieve at least 30 pages of 100 results, making reliable deeper pagination available without new credentials or dependencies.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Authenticate every managed DockerHub repository-search mode through the configured account pool and existing Hub bearer-token flow.
|
||||
- **BREAKING**: when an explicit DockerHub account pool is configured, fail closed if no account can authenticate instead of falling back to anonymous search.
|
||||
- Support an authenticated search window of up to 30 pages while retaining a bounded code-level limit.
|
||||
- Retry transient page failures once, rotate accounts for account-specific failures, and reject an incomplete expected page set rather than enqueueing partial discovery results.
|
||||
- Preserve the current query cursor when pagination fails, while continuing to advance it after complete or objectively exhausted pagination.
|
||||
- Keep tag resolution, Registry authentication, immutable-digest deduplication, scan retries, and queue disposition unchanged.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `dockerhub-search-pagination`: Authenticated, bounded, complete-or-fail DockerHub repository-search pagination using the managed account pool.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects DockerHub search/authentication in `app/scanner.py`, source-cycle failure propagation in `app/console_runner.py`, and focused scanner/runner tests.
|
||||
- Reuses existing configured DockerHub accounts, Hub bearer tokens, cooldowns, and auth-event persistence; no new external dependency or credential format is introduced.
|
||||
- Active query lists, refresh cadence, scan budgets, cold/failed target policy, and layer-aware scanning are out of scope.
|
||||
+68
@@ -0,0 +1,68 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Managed repository search is authenticated
|
||||
The system SHALL authenticate every standard and recent DockerHub repository-search request through the configured DockerHub account pool using a Hub bearer token.
|
||||
|
||||
#### Scenario: Configured account performs search
|
||||
- **WHEN** a managed DockerHub source cycle searches for repositories with an available account
|
||||
- **THEN** the system sends the search request with that account's bearer authorization and records success under the `hub_search` endpoint identity
|
||||
|
||||
#### Scenario: Explicit pool has no usable account
|
||||
- **WHEN** an explicit DockerHub account pool is configured but no account can authenticate or leave cooldown
|
||||
- **THEN** the system fails the repository search without making an anonymous fallback request
|
||||
|
||||
#### Scenario: Search authorization is rejected
|
||||
- **WHEN** a search request receives an account-specific 401, 403, or 429 response
|
||||
- **THEN** the system refreshes a rejected bearer once where applicable and rotates to another usable account within the configured pool
|
||||
|
||||
### Requirement: Authenticated pagination is bounded
|
||||
The system SHALL support up to 30 DockerHub search pages per query and SHALL cap larger requested page counts at 30.
|
||||
|
||||
#### Scenario: Thirty-page authenticated search
|
||||
- **WHEN** a query requests 30 pages and the first page reports at least 30 pages of results
|
||||
- **THEN** the system requests the complete page range from 1 through 30
|
||||
|
||||
#### Scenario: Request exceeds safety cap
|
||||
- **WHEN** a query requests more than 30 pages
|
||||
- **THEN** the system limits the search to pages 1 through 30 and records that the requested range was capped
|
||||
|
||||
#### Scenario: Reported result set is shorter
|
||||
- **WHEN** page one reports fewer results than the requested page range would contain
|
||||
- **THEN** the system requests only the pages required by that reported count
|
||||
|
||||
### Requirement: Page acquisition is bounded and complete
|
||||
The system SHALL give each expected search page at most two transient transport/server attempts and SHALL not return partial repository results when any expected page remains unavailable.
|
||||
|
||||
#### Scenario: Transient failure recovers
|
||||
- **WHEN** an expected page receives a retryable transport error or transient HTTP status on its first attempt and succeeds on its second attempt
|
||||
- **THEN** the system includes that page and completes discovery without another transient attempt
|
||||
|
||||
#### Scenario: Expected page remains unavailable
|
||||
- **WHEN** an expected page still fails after bounded retry and account handling
|
||||
- **THEN** the system raises a DockerHub discovery transport failure before tag resolution or repository enqueue
|
||||
|
||||
#### Scenario: Every expected page succeeds
|
||||
- **WHEN** all expected pages return valid payloads
|
||||
- **THEN** the system combines their repositories in page order and proceeds with existing deduplication and resolution behavior
|
||||
|
||||
### Requirement: Failed pagination preserves query rotation
|
||||
The system SHALL record incomplete DockerHub pagination as a failed source cycle and SHALL keep the current query cursor unchanged.
|
||||
|
||||
#### Scenario: Source cycle receives pagination failure
|
||||
- **WHEN** repository discovery raises a DockerHub discovery transport failure
|
||||
- **THEN** the source cycle finishes with failed status, enqueues no partial search result, and selects the same query for the next cycle
|
||||
|
||||
#### Scenario: Complete source cycle succeeds
|
||||
- **WHEN** repository discovery and the remaining source cycle complete normally
|
||||
- **THEN** the existing query-advance policy remains unchanged
|
||||
|
||||
### Requirement: Search authentication is secret-safe and isolated
|
||||
The system MUST NOT expose account credentials or bearer tokens through search logs, errors, or auth events, and SHALL preserve existing tag, Registry, immutable-digest, and scan-retry behavior.
|
||||
|
||||
#### Scenario: Search request fails
|
||||
- **WHEN** an authenticated search request fails or exhausts the account pool
|
||||
- **THEN** emitted diagnostics identify only the safe endpoint/status category without including usernames, credentials, bearer values, or request authorization headers
|
||||
|
||||
#### Scenario: Repository search implementation changes
|
||||
- **WHEN** authenticated search pagination is deployed
|
||||
- **THEN** existing DockerHub tag resolution, Registry authentication, target deduplication, and scan retry contracts remain unchanged
|
||||
@@ -0,0 +1,21 @@
|
||||
## 1. Runtime Safety
|
||||
|
||||
- [x] 1.1 Canonically stop the live supervisor and verify all managed child processes are down before editing `app`
|
||||
|
||||
## 2. Authenticated Search Implementation
|
||||
|
||||
- [x] 2.1 Make Hub bearer acquisition endpoint-aware and add a `hub_search` response path with explicit-pool fail-closed behavior, account rotation, and two-attempt transient request bounds
|
||||
- [x] 2.2 Raise the search safety cap to 30 pages and make standard pagination fetch page one first, derive the expected range, and reject any incomplete expected page set
|
||||
- [x] 2.3 Route recent-mode repository search through the same authenticated bounded response path
|
||||
- [x] 2.4 Add a DockerHub discovery transport failure path that records a failed source cycle without advancing the query cursor
|
||||
|
||||
## 3. Regression Coverage
|
||||
|
||||
- [x] 3.1 Cover bearer authorization, token refresh, account rotation/cooldown, and explicit-pool no-fallback behavior without exposing secrets
|
||||
- [x] 3.2 Cover the 30-page cap, page-one count boundary, ordered complete results, one transient retry, and rejection of unresolved partial pagination
|
||||
- [x] 3.3 Cover failed-cycle cursor retention and successful-cycle compatibility
|
||||
- [x] 3.4 Run focused and broader relevant scanner/runner test suites
|
||||
|
||||
## 4. Runtime Verification
|
||||
|
||||
- [x] 4.1 Canonically restart the supervisor and verify PostgreSQL, pipeline workers, DockerHub source health, restart counters, and absence of app bytecode artifacts
|
||||
Reference in New Issue
Block a user