9.6 KiB
ADDED Requirements
Requirement: DockerHub pages are admitted incrementally
The system SHALL validate and durably admit every successful DockerHub repository-search page before relying on later-page acquisition, while preserving target identity and cold/failed-target exclusion.
Scenario: Successful page precedes a later failure
- WHEN an expected DockerHub page is valid and a later expected page exhausts its request budget
- THEN repositories from the successful page remain idempotently admitted and the later failure cannot roll them back
Scenario: Page admission fails
- WHEN repository normalization or durable page admission fails
- THEN the source cycle fails, does not treat the page as complete, and does not advance the main query cursor
Scenario: Resolver processing follows pagination
- WHEN one or more pages admit repository anchors
- THEN the existing Docker resolver gate runs at most once after the pass rather than once per page
Requirement: Ordinary discovery stops after consecutive known pages
The system SHALL stop an ordinary DockerHub query before requesting another page after two consecutive nonempty pages contain only repository identities that existed before the current pass.
Scenario: First two pages were previously known
- WHEN pages one and two are nonempty and every normalized repository identity existed before the pass
- THEN the pass completes without requesting page three
Scenario: Page contains a new repository
- WHEN either of the last two pages contains a repository identity not known before the pass
- THEN the consecutive-known counter resets and discovery continues within the effective range
Scenario: Current-pass duplicate appears later
- WHEN a repository first admitted earlier in the same pass appears on a later page
- THEN that identity is not treated as preexisting evidence for the later page's known-page stop
Scenario: Knownness cannot be determined
- WHEN the database cannot establish complete pre-pass knownness for a valid page
- THEN discovery fails open by continuing deeper rather than stopping early
Scenario: Previous pass ended incompletely
- WHEN an exact query and policy have a retained incomplete-pass marker
- THEN the next selected pass bypasses known-page stopping until a durable pass outcome clears the marker
Requirement: Deep discovery bypasses seen-page stopping every 72 hours
The system SHALL make each exact configured DockerHub query deep-due no later than 72 hours after its last durable deep dispatch and SHALL bypass the consecutive-known-page stop on that query's next selected pass.
Scenario: Exact query reaches its deep interval
- WHEN at least 72 hours have elapsed since the query's last durable deep dispatch
- THEN its next selected pass processes the available expected range without stopping on known pages
Scenario: Query is new or policy changes
- WHEN an exact configured query has no valid state or its effective search-policy hash changes
- THEN that query is immediately deep-due without resetting unrelated query schedules
Scenario: Deep pass has durable gaps
- WHEN valid pages are admitted and unavailable pages are durably delegated to retry work
- THEN the deep dispatch timestamp advances while completion remains represented by the outstanding retry work
Scenario: Provider prevents acquisition
- WHEN neither successful pages nor durable retry delegation can establish a deep dispatch
- THEN the system does not advance that query's deep timestamp
Requirement: DockerHub page gaps use a durable retry backlog
The system SHALL persist unavailable DockerHub query/page work in a dedicated PostgreSQL retry backlog before allowing the main keyword rotation to advance.
Scenario: Page one is unavailable
- WHEN page one exhausts its bounded request/account handling and the expected range is unknown
- THEN the system durably enqueues query-level retry work
Scenario: Later page is unavailable
- WHEN a later expected page exhausts its bounded request/account handling
- THEN the system durably enqueues page-level work and retains every successfully admitted page
Scenario: Remaining tail cannot be attempted
- WHEN provider/account exhaustion prevents remote attempts for a known remaining page range
- THEN the system coalesces the unattempted range into bounded retry work instead of creating unbounded individual rows
Scenario: Durable delegation succeeds
- WHEN every observed acquisition gap has durable retry work
- THEN the cycle records
completed_with_retriesand advances the main keyword independently of retry processing
Scenario: Durable delegation fails
- WHEN retry work cannot be persisted authoritatively
- THEN the cycle fails and retains the current keyword cursor
Requirement: Discovery retry claims are bounded and fenced
The system SHALL coalesce retry work by source, exact query, effective policy, pass kind, and page/range; SHALL process bounded due work with expiring leases; and MUST reject stale acknowledgements.
Scenario: Concurrent workers claim due work
- WHEN multiple workers attempt to claim the same due retry row
- THEN at most one receives the active lease token
Scenario: Lease expires
- WHEN a worker fails to complete work before its lease expires
- THEN the row becomes reclaimable without deleting its attempt history
Scenario: Multi-page retry remains active
- WHEN a leased query or range retry is about to request another page
- THEN the worker renews the same owner-and-token fence before acquisition and stops if renewal is rejected
Scenario: Stale worker finishes
- WHEN a worker presents an obsolete owner or lease token
- THEN it cannot acknowledge, delete, defer, or hold the newer work
Scenario: Retryable acquisition fails again
- WHEN leased work encounters another retryable failure
- THEN it returns to pending with bounded exponential backoff and is not silently dropped at an attempt limit
Scenario: Provider cooldown follows partial remote progress
- WHEN an earlier page in the same claim made a remote request before a later local provider cooldown
- THEN the dispatch attempt is not refunded
Scenario: Query is removed or policy is obsolete
- WHEN retry work no longer matches an exact configured query and effective policy
- THEN it is held and cannot execute against stale discovery policy
Requirement: Retry work does not block main keyword rotation
The system SHALL process at most a bounded amount of due discovery retry work per source-loop iteration independently of the saved main query cursor.
Scenario: Persistent page failure exists
- WHEN a retry row remains unavailable across multiple attempts
- THEN ordinary configured keywords continue rotating while the row follows its own backoff
Scenario: Retry succeeds
- WHEN leased page or range work returns valid repositories
- THEN repositories are durably admitted before the fenced retry acknowledgement
Requirement: DockerHub search policy uses the confirmed breadth
The system SHALL use an effective ceiling of 30 pages and 100 results per page for every configured DockerHub query and SHALL include exactly the 12 confirmed new product/framework terms in addition to the existing ordered query set.
Scenario: Ordinary pass reaches known content
- WHEN a normal 30-by-100 query pass reaches two consecutive preexisting-known pages
- THEN it stops early despite the larger configured ceiling
Scenario: Deep pass remains novel
- WHEN a due deep pass continues to return pages containing new repositories
- THEN it processes up to the page-one result boundary or the 30-page safety cap
Scenario: Query list is expanded
- WHEN the configuration is loaded after deployment
- THEN
open-webui,ragflow,dify,flowise,crewai,n8n,langflow,autogen,browser-use,openhands,anythingllm, andagent-zeroeach appear once at the ordered tail
Requirement: Completed repository refresh is disabled without changing retries
The system SHALL disable periodic re-resolution of successfully completed DockerHub repository anchors while retaining normal search, initial/partial resolver work, and immutable-digest error retries.
Scenario: Completed anchor becomes periodically due
- WHEN a resolved repository anchor reaches its prior refresh interval
- THEN periodic policy does not claim it solely for refresh
Scenario: Initial or partial anchor is due
- WHEN an unresolved or retryable partial repository anchor is due
- THEN the existing resolver remains eligible to process it
Scenario: Immutable digest scan fails retryably
- WHEN an immutable Docker image scan meets the existing retry conditions
- THEN its current bounded retry and backoff behavior remains unchanged
Requirement: Incremental discovery remains authenticated and secret-safe
The system MUST preserve explicit-pool fail-closed Hub bearer authentication and MUST NOT persist or emit credentials, bearer values, authorization headers, raw response bodies, arbitrary exception text, or repository targets through retry/deep-state diagnostics.
Scenario: Search or retry fails
- WHEN DockerHub page acquisition or retry processing reports an error
- THEN diagnostics contain only bounded status/category/count metadata required for operation
Scenario: Explicit account pool is unavailable
- WHEN no configured account can perform repository search
- THEN normal and retry acquisition fail closed without anonymous fallback