Files
2026-09-30 20:30:56 +03:00

6.7 KiB

Context

Managed DockerHub source cycles configure an explicit account pool, but repository search bypasses it and calls the Hub search endpoint anonymously. Docker Hub rejects anonymous result windows beyond 200 entries; an isolated authenticated probe using the existing Hub bearer flow returned 100 results for pages 3, 20, 21, and 30 at a page size of 100. The standard search path also schedules every page before learning the reported result count and currently accepts successful pages when another expected page fails, so the source cycle completes and advances its numeric query cursor with incomplete discovery.

The account manager already provides secret-safe accounts, cached Hub bearer tokens, account rotation, endpoint-keyed cooldowns, and persisted auth events. The implementation must reuse those controls without changing tag resolution, Registry access, or immutable scan behavior.

Goals / Non-Goals

Goals:

  • Authenticate standard and recent DockerHub repository searches through the configured account pool.
  • Fail closed when an explicit pool has no usable account.
  • Support at most 30 authenticated search pages and avoid requests beyond the result count reported by page one.
  • Give each page exactly one additional attempt for transient transport or server failures.
  • Return a complete expected page set or fail the source cycle without enqueueing partial results or advancing its query cursor.
  • Keep logs, exceptions, and auth events free of credentials and bearer values.

Non-Goals:

  • Adding discovery queries, changing active page sizes, changing repository refresh cadence, or altering scan budgets.
  • Changing Docker tag selection, Registry bearer authentication, layer scanning, immutable target identity, or scan retry policy.
  • Guaranteeing that Docker Hub will indefinitely support 30 pages; the local cap remains a safety bound, not an external SLA.

Decisions

Reuse Hub bearer authentication with a search-specific endpoint identity

Add a repository-search response helper that follows the existing tag-response pattern but uses the hub_search endpoint key. It obtains cached or fresh Hub bearer tokens, sends Authorization: Bearer ..., refreshes once after a 401, and rotates across the configured accounts for 401, 403, or 429 responses. Search cooldowns remain separate from hub_tags and registry cooldowns. The shared token cache remains per account because the same Hub token is valid for both Hub endpoints.

When the manager represents an explicit pool and no account is usable, the helper fails before making an anonymous request. A non-explicit legacy invocation may retain anonymous behavior for direct CLI compatibility.

Alternative considered: add a second login mechanism or cookie session. Rejected because the existing /v2/auth/token bearer flow was verified against authenticated search through page 30 and avoids another credential path.

Fetch page one before parallel remainder

Raise the code-level page cap from 20 to 30. Standard discovery fetches page one first, validates its payload, derives the expected page count from its reported count, and submits only pages 2 through min(requested, expected, 30) concurrently. This removes ambiguous post-hoc suppression of failed pages and avoids requesting pages objectively outside the reported result set.

Alternative considered: keep launching all pages concurrently and ignore failures above the largest successful count. Rejected because a failed early page can make the inferred boundary unreliable, and unnecessary out-of-range requests consume account budget.

Bound page retries inside the request primitive

The authenticated search helper gives each page at most two search GETs across transient retry, bearer refresh, and account rotation. Each GET calls api_request with one total network attempt; the outer page budget supplies the single bounded retry for 408, 500, 502, 503, 504, or account-specific failures. A legacy anonymous invocation has no account handling and calls api_request with two total attempts directly. This keeps the page-wide budget independent of global proxy retry settings and prevents it from resetting during account rotation.

Treat incomplete pagination as a source-cycle transport failure

Introduce a DockerHub discovery transport error analogous to the existing GitLab error. Any expected page that remains unavailable after its bounded request/account handling aborts the complete search result before tag resolution or enqueue. The configured-source runner records a failed cycle and returns without advancing the query cursor; it does not crash the long-running source process. A later cycle retries the same query, and queue uniqueness keeps successful rediscovery idempotent.

Keep discovery authentication isolated from scan behavior

The change only replaces repository-search HTTP calls and their error propagation. Tag/manifest resolution continues to use its existing endpoint identities, retries, caches, resolver states, and immutable digest constraints.

Risks / Trade-offs

  • [Thirty authenticated pages increase search request volume] -> Preserve the hard cap, configured worker bounds, and account rotation; this change does not raise active source page settings.
  • [Page-one count can change while later pages are fetched] -> Treat page one as the cycle snapshot boundary; all pages within that boundary must still succeed, and DB deduplication handles overlap caused by result movement.
  • [A single page outage now rejects otherwise usable pages] -> This is intentional complete-or-fail behavior; one retry limits transient loss, and the unchanged cursor retries the query later.
  • [Fetching page one serially adds one request latency before parallel work] -> It prevents unnecessary pages and provides an authoritative expected set, which is more valuable than the small latency saving.
  • [All accounts can be temporarily unavailable] -> Fail closed and persist endpoint-specific auth state rather than silently reverting to the anonymous 200-result window.

Migration Plan

  1. Canonically stop the supervisor and verify all managed children are down.
  2. Deploy the scanner, runner, and focused regression tests without changing source query configuration.
  3. Run focused DockerHub authentication/pagination/cursor tests and the broader relevant scanner suites.
  4. Canonically restart the supervisor and verify PostgreSQL, pipeline workers, DockerHub source worker, and restart counters.
  5. Roll back by restoring the previous code under a canonical stop/start if authenticated search causes an operational regression; no data migration is required.

Open Questions

None. Authenticated access through page 30 and the explicit-pool behavior have been verified or are covered by deterministic tests.