Files
2026-09-30 20:30:56 +03:00

165 lines
9.6 KiB
Markdown

## ADDED Requirements
### Requirement: DockerHub pages are admitted incrementally
The system SHALL validate and durably admit every successful DockerHub repository-search page before relying on later-page acquisition, while preserving target identity and cold/failed-target exclusion.
#### Scenario: Successful page precedes a later failure
- **WHEN** an expected DockerHub page is valid and a later expected page exhausts its request budget
- **THEN** repositories from the successful page remain idempotently admitted and the later failure cannot roll them back
#### Scenario: Page admission fails
- **WHEN** repository normalization or durable page admission fails
- **THEN** the source cycle fails, does not treat the page as complete, and does not advance the main query cursor
#### Scenario: Resolver processing follows pagination
- **WHEN** one or more pages admit repository anchors
- **THEN** the existing Docker resolver gate runs at most once after the pass rather than once per page
### Requirement: Ordinary discovery stops after consecutive known pages
The system SHALL stop an ordinary DockerHub query before requesting another page after two consecutive nonempty pages contain only repository identities that existed before the current pass.
#### Scenario: First two pages were previously known
- **WHEN** pages one and two are nonempty and every normalized repository identity existed before the pass
- **THEN** the pass completes without requesting page three
#### Scenario: Page contains a new repository
- **WHEN** either of the last two pages contains a repository identity not known before the pass
- **THEN** the consecutive-known counter resets and discovery continues within the effective range
#### Scenario: Current-pass duplicate appears later
- **WHEN** a repository first admitted earlier in the same pass appears on a later page
- **THEN** that identity is not treated as preexisting evidence for the later page's known-page stop
#### Scenario: Knownness cannot be determined
- **WHEN** the database cannot establish complete pre-pass knownness for a valid page
- **THEN** discovery fails open by continuing deeper rather than stopping early
#### Scenario: Previous pass ended incompletely
- **WHEN** an exact query and policy have a retained incomplete-pass marker
- **THEN** the next selected pass bypasses known-page stopping until a durable pass outcome clears the marker
### Requirement: Deep discovery bypasses seen-page stopping every 72 hours
The system SHALL make each exact configured DockerHub query deep-due no later than 72 hours after its last durable deep dispatch and SHALL bypass the consecutive-known-page stop on that query's next selected pass.
#### Scenario: Exact query reaches its deep interval
- **WHEN** at least 72 hours have elapsed since the query's last durable deep dispatch
- **THEN** its next selected pass processes the available expected range without stopping on known pages
#### Scenario: Query is new or policy changes
- **WHEN** an exact configured query has no valid state or its effective search-policy hash changes
- **THEN** that query is immediately deep-due without resetting unrelated query schedules
#### Scenario: Deep pass has durable gaps
- **WHEN** valid pages are admitted and unavailable pages are durably delegated to retry work
- **THEN** the deep dispatch timestamp advances while completion remains represented by the outstanding retry work
#### Scenario: Provider prevents acquisition
- **WHEN** neither successful pages nor durable retry delegation can establish a deep dispatch
- **THEN** the system does not advance that query's deep timestamp
### Requirement: DockerHub page gaps use a durable retry backlog
The system SHALL persist unavailable DockerHub query/page work in a dedicated PostgreSQL retry backlog before allowing the main keyword rotation to advance.
#### Scenario: Page one is unavailable
- **WHEN** page one exhausts its bounded request/account handling and the expected range is unknown
- **THEN** the system durably enqueues query-level retry work
#### Scenario: Later page is unavailable
- **WHEN** a later expected page exhausts its bounded request/account handling
- **THEN** the system durably enqueues page-level work and retains every successfully admitted page
#### Scenario: Remaining tail cannot be attempted
- **WHEN** provider/account exhaustion prevents remote attempts for a known remaining page range
- **THEN** the system coalesces the unattempted range into bounded retry work instead of creating unbounded individual rows
#### Scenario: Durable delegation succeeds
- **WHEN** every observed acquisition gap has durable retry work
- **THEN** the cycle records `completed_with_retries` and advances the main keyword independently of retry processing
#### Scenario: Durable delegation fails
- **WHEN** retry work cannot be persisted authoritatively
- **THEN** the cycle fails and retains the current keyword cursor
### Requirement: Discovery retry claims are bounded and fenced
The system SHALL coalesce retry work by source, exact query, effective policy, pass kind, and page/range; SHALL process bounded due work with expiring leases; and MUST reject stale acknowledgements.
#### Scenario: Concurrent workers claim due work
- **WHEN** multiple workers attempt to claim the same due retry row
- **THEN** at most one receives the active lease token
#### Scenario: Lease expires
- **WHEN** a worker fails to complete work before its lease expires
- **THEN** the row becomes reclaimable without deleting its attempt history
#### Scenario: Multi-page retry remains active
- **WHEN** a leased query or range retry is about to request another page
- **THEN** the worker renews the same owner-and-token fence before acquisition and stops if renewal is rejected
#### Scenario: Stale worker finishes
- **WHEN** a worker presents an obsolete owner or lease token
- **THEN** it cannot acknowledge, delete, defer, or hold the newer work
#### Scenario: Retryable acquisition fails again
- **WHEN** leased work encounters another retryable failure
- **THEN** it returns to pending with bounded exponential backoff and is not silently dropped at an attempt limit
#### Scenario: Provider cooldown follows partial remote progress
- **WHEN** an earlier page in the same claim made a remote request before a later local provider cooldown
- **THEN** the dispatch attempt is not refunded
#### Scenario: Query is removed or policy is obsolete
- **WHEN** retry work no longer matches an exact configured query and effective policy
- **THEN** it is held and cannot execute against stale discovery policy
### Requirement: Retry work does not block main keyword rotation
The system SHALL process at most a bounded amount of due discovery retry work per source-loop iteration independently of the saved main query cursor.
#### Scenario: Persistent page failure exists
- **WHEN** a retry row remains unavailable across multiple attempts
- **THEN** ordinary configured keywords continue rotating while the row follows its own backoff
#### Scenario: Retry succeeds
- **WHEN** leased page or range work returns valid repositories
- **THEN** repositories are durably admitted before the fenced retry acknowledgement
### Requirement: DockerHub search policy uses the confirmed breadth
The system SHALL use an effective ceiling of 30 pages and 100 results per page for every configured DockerHub query and SHALL include exactly the 12 confirmed new product/framework terms in addition to the existing ordered query set.
#### Scenario: Ordinary pass reaches known content
- **WHEN** a normal 30-by-100 query pass reaches two consecutive preexisting-known pages
- **THEN** it stops early despite the larger configured ceiling
#### Scenario: Deep pass remains novel
- **WHEN** a due deep pass continues to return pages containing new repositories
- **THEN** it processes up to the page-one result boundary or the 30-page safety cap
#### Scenario: Query list is expanded
- **WHEN** the configuration is loaded after deployment
- **THEN** `open-webui`, `ragflow`, `dify`, `flowise`, `crewai`, `n8n`, `langflow`, `autogen`, `browser-use`, `openhands`, `anythingllm`, and `agent-zero` each appear once at the ordered tail
### Requirement: Completed repository refresh is disabled without changing retries
The system SHALL disable periodic re-resolution of successfully completed DockerHub repository anchors while retaining normal search, initial/partial resolver work, and immutable-digest error retries.
#### Scenario: Completed anchor becomes periodically due
- **WHEN** a resolved repository anchor reaches its prior refresh interval
- **THEN** periodic policy does not claim it solely for refresh
#### Scenario: Initial or partial anchor is due
- **WHEN** an unresolved or retryable partial repository anchor is due
- **THEN** the existing resolver remains eligible to process it
#### Scenario: Immutable digest scan fails retryably
- **WHEN** an immutable Docker image scan meets the existing retry conditions
- **THEN** its current bounded retry and backoff behavior remains unchanged
### Requirement: Incremental discovery remains authenticated and secret-safe
The system MUST preserve explicit-pool fail-closed Hub bearer authentication and MUST NOT persist or emit credentials, bearer values, authorization headers, raw response bodies, arbitrary exception text, or repository targets through retry/deep-state diagnostics.
#### Scenario: Search or retry fails
- **WHEN** DockerHub page acquisition or retry processing reports an error
- **THEN** diagnostics contain only bounded status/category/count metadata required for operation
#### Scenario: Explicit account pool is unavailable
- **WHEN** no configured account can perform repository search
- **THEN** normal and retry acquisition fail closed without anonymous fallback