Files
2026-09-30 20:30:56 +03:00

9.6 KiB

ADDED Requirements

Requirement: DockerHub pages are admitted incrementally

The system SHALL validate and durably admit every successful DockerHub repository-search page before relying on later-page acquisition, while preserving target identity and cold/failed-target exclusion.

Scenario: Successful page precedes a later failure

  • WHEN an expected DockerHub page is valid and a later expected page exhausts its request budget
  • THEN repositories from the successful page remain idempotently admitted and the later failure cannot roll them back

Scenario: Page admission fails

  • WHEN repository normalization or durable page admission fails
  • THEN the source cycle fails, does not treat the page as complete, and does not advance the main query cursor

Scenario: Resolver processing follows pagination

  • WHEN one or more pages admit repository anchors
  • THEN the existing Docker resolver gate runs at most once after the pass rather than once per page

Requirement: Ordinary discovery stops after consecutive known pages

The system SHALL stop an ordinary DockerHub query before requesting another page after two consecutive nonempty pages contain only repository identities that existed before the current pass.

Scenario: First two pages were previously known

  • WHEN pages one and two are nonempty and every normalized repository identity existed before the pass
  • THEN the pass completes without requesting page three

Scenario: Page contains a new repository

  • WHEN either of the last two pages contains a repository identity not known before the pass
  • THEN the consecutive-known counter resets and discovery continues within the effective range

Scenario: Current-pass duplicate appears later

  • WHEN a repository first admitted earlier in the same pass appears on a later page
  • THEN that identity is not treated as preexisting evidence for the later page's known-page stop

Scenario: Knownness cannot be determined

  • WHEN the database cannot establish complete pre-pass knownness for a valid page
  • THEN discovery fails open by continuing deeper rather than stopping early

Scenario: Previous pass ended incompletely

  • WHEN an exact query and policy have a retained incomplete-pass marker
  • THEN the next selected pass bypasses known-page stopping until a durable pass outcome clears the marker

Requirement: Deep discovery bypasses seen-page stopping every 72 hours

The system SHALL make each exact configured DockerHub query deep-due no later than 72 hours after its last durable deep dispatch and SHALL bypass the consecutive-known-page stop on that query's next selected pass.

Scenario: Exact query reaches its deep interval

  • WHEN at least 72 hours have elapsed since the query's last durable deep dispatch
  • THEN its next selected pass processes the available expected range without stopping on known pages

Scenario: Query is new or policy changes

  • WHEN an exact configured query has no valid state or its effective search-policy hash changes
  • THEN that query is immediately deep-due without resetting unrelated query schedules

Scenario: Deep pass has durable gaps

  • WHEN valid pages are admitted and unavailable pages are durably delegated to retry work
  • THEN the deep dispatch timestamp advances while completion remains represented by the outstanding retry work

Scenario: Provider prevents acquisition

  • WHEN neither successful pages nor durable retry delegation can establish a deep dispatch
  • THEN the system does not advance that query's deep timestamp

Requirement: DockerHub page gaps use a durable retry backlog

The system SHALL persist unavailable DockerHub query/page work in a dedicated PostgreSQL retry backlog before allowing the main keyword rotation to advance.

Scenario: Page one is unavailable

  • WHEN page one exhausts its bounded request/account handling and the expected range is unknown
  • THEN the system durably enqueues query-level retry work

Scenario: Later page is unavailable

  • WHEN a later expected page exhausts its bounded request/account handling
  • THEN the system durably enqueues page-level work and retains every successfully admitted page

Scenario: Remaining tail cannot be attempted

  • WHEN provider/account exhaustion prevents remote attempts for a known remaining page range
  • THEN the system coalesces the unattempted range into bounded retry work instead of creating unbounded individual rows

Scenario: Durable delegation succeeds

  • WHEN every observed acquisition gap has durable retry work
  • THEN the cycle records completed_with_retries and advances the main keyword independently of retry processing

Scenario: Durable delegation fails

  • WHEN retry work cannot be persisted authoritatively
  • THEN the cycle fails and retains the current keyword cursor

Requirement: Discovery retry claims are bounded and fenced

The system SHALL coalesce retry work by source, exact query, effective policy, pass kind, and page/range; SHALL process bounded due work with expiring leases; and MUST reject stale acknowledgements.

Scenario: Concurrent workers claim due work

  • WHEN multiple workers attempt to claim the same due retry row
  • THEN at most one receives the active lease token

Scenario: Lease expires

  • WHEN a worker fails to complete work before its lease expires
  • THEN the row becomes reclaimable without deleting its attempt history

Scenario: Multi-page retry remains active

  • WHEN a leased query or range retry is about to request another page
  • THEN the worker renews the same owner-and-token fence before acquisition and stops if renewal is rejected

Scenario: Stale worker finishes

  • WHEN a worker presents an obsolete owner or lease token
  • THEN it cannot acknowledge, delete, defer, or hold the newer work

Scenario: Retryable acquisition fails again

  • WHEN leased work encounters another retryable failure
  • THEN it returns to pending with bounded exponential backoff and is not silently dropped at an attempt limit

Scenario: Provider cooldown follows partial remote progress

  • WHEN an earlier page in the same claim made a remote request before a later local provider cooldown
  • THEN the dispatch attempt is not refunded

Scenario: Query is removed or policy is obsolete

  • WHEN retry work no longer matches an exact configured query and effective policy
  • THEN it is held and cannot execute against stale discovery policy

Requirement: Retry work does not block main keyword rotation

The system SHALL process at most a bounded amount of due discovery retry work per source-loop iteration independently of the saved main query cursor.

Scenario: Persistent page failure exists

  • WHEN a retry row remains unavailable across multiple attempts
  • THEN ordinary configured keywords continue rotating while the row follows its own backoff

Scenario: Retry succeeds

  • WHEN leased page or range work returns valid repositories
  • THEN repositories are durably admitted before the fenced retry acknowledgement

Requirement: DockerHub search policy uses the confirmed breadth

The system SHALL use an effective ceiling of 30 pages and 100 results per page for every configured DockerHub query and SHALL include exactly the 12 confirmed new product/framework terms in addition to the existing ordered query set.

Scenario: Ordinary pass reaches known content

  • WHEN a normal 30-by-100 query pass reaches two consecutive preexisting-known pages
  • THEN it stops early despite the larger configured ceiling

Scenario: Deep pass remains novel

  • WHEN a due deep pass continues to return pages containing new repositories
  • THEN it processes up to the page-one result boundary or the 30-page safety cap

Scenario: Query list is expanded

  • WHEN the configuration is loaded after deployment
  • THEN open-webui, ragflow, dify, flowise, crewai, n8n, langflow, autogen, browser-use, openhands, anythingllm, and agent-zero each appear once at the ordered tail

Requirement: Completed repository refresh is disabled without changing retries

The system SHALL disable periodic re-resolution of successfully completed DockerHub repository anchors while retaining normal search, initial/partial resolver work, and immutable-digest error retries.

Scenario: Completed anchor becomes periodically due

  • WHEN a resolved repository anchor reaches its prior refresh interval
  • THEN periodic policy does not claim it solely for refresh

Scenario: Initial or partial anchor is due

  • WHEN an unresolved or retryable partial repository anchor is due
  • THEN the existing resolver remains eligible to process it

Scenario: Immutable digest scan fails retryably

  • WHEN an immutable Docker image scan meets the existing retry conditions
  • THEN its current bounded retry and backoff behavior remains unchanged

Requirement: Incremental discovery remains authenticated and secret-safe

The system MUST preserve explicit-pool fail-closed Hub bearer authentication and MUST NOT persist or emit credentials, bearer values, authorization headers, raw response bodies, arbitrary exception text, or repository targets through retry/deep-state diagnostics.

Scenario: Search or retry fails

  • WHEN DockerHub page acquisition or retry processing reports an error
  • THEN diagnostics contain only bounded status/category/count metadata required for operation

Scenario: Explicit account pool is unavailable

  • WHEN no configured account can perform repository search
  • THEN normal and retry acquisition fail closed without anonymous fallback