3.4 KiB
DockerHub discovery retry coalescing fails for an unbounded retry time
Status
Open in the accepted live server runtime. Reproduced from the complete DockerHub producer state and PostgreSQL authority on 2026-09-25. A minimal source fix and PostgreSQL regression coverage have been added locally but have not been deployed to the live runtime during the worker artifact validation.
Impact
When DockerHub discovery cannot acquire a search page and delegates work without
an available_after timestamp, an already pending retry row cannot be coalesced.
The producer reports a generic DockerHub discovery retry delegation failed, the
source cycle fails, and the configured query does not advance. This can keep the
managed DockerHub producer in a restart loop while other sources remain healthy.
The complete producer history contained 265 failed DockerHub cycles with this masked message. A protected diagnostic cycle using the captured zero-available-auth state reproduced the original database exception exactly.
Raw evidence
- Producer log:
build/live-trace-20260925/dockerhub-discovery.log - Captured runner state:
build/live-trace-20260925/runner_state_dockerhub.json - Complete retry/pass/cycle/lock capture:
build/live-trace-20260925/dockerhub-discovery-db-raw.json - Unmasked exception:
build/live-trace-20260925/dockerhub-unmasked-cycle-failure.json
The unmasked exception is PostgreSQL IndeterminateDatatype, SQLSTATE 42P18:
could not determine data type of parameter $1. It originates in
ScannerDB.enqueue_discovery_retry() while coalescing an existing pending row.
Root cause
app/scanner_db.py used an untyped nullable placeholder in the PostgreSQL
expression:
WHEN available_after IS NULL OR ? IS NULL THEN NULL
When available_after=None, PostgreSQL has no typed expression from which it can
infer the placeholder type. SQLite accepts the same query, so existing SQLite
coalescing coverage did not expose the problem. Existing PostgreSQL integration
coverage inserted a retry but did not coalesce the same row with a null retry time.
The outer discovery helper intentionally replaced the original database exception with a generic delegation error, which obscured the SQLSTATE in normal managed logs.
Local correction
The placeholder is now explicitly typed as the schema's text timestamp form:
WHEN available_after IS NULL OR CAST(? AS TEXT) IS NULL THEN NULL
tests/test_pipeline_postgres_integration.py now coalesces the same retry with no
available_after value and checks that the existing row is reused.
The corrected statement was also executed against the live PostgreSQL schema in a transaction and rolled back. It matched one pending row and completed without an exception. Local focused results:
- DockerHub incremental discovery tests: 27 passed.
- SQLite retry lifecycle regression: 1 passed.
- PostgreSQL integration test: skipped locally because no disposable PostgreSQL DSN was configured; the corrected SQL shape was verified transactionally against the live schema without persisting a change.
Operational note
The configured DockerHub credential pool was independently exhausted at capture time: ten entries were invalid and the remaining entry was rate-limited. Correcting retry coalescing preserves the failed work and stops this database error, but it does not make an unavailable credential pool healthy.