10 KiB
10 KiB
Runtime Reliability Parking Lot
Last reviewed: 2026-08-27
This file records observed follow-up work that is intentionally outside completed changes. Evidence must be refreshed before opening a new change.
P1: GitLab discovery exits on one transient API timeout
- Status: implemented by
stabilize-gitlab-discovery-retries; all 9 tasks complete and the change is ready for archive. - Evidence: four GitLab source restarts during the current runtime window matched four
GET /api/v4/projectsread timeouts. - Current behavior: direct API calls get one attempt; an exhausted discovery request escapes the source cycle and exits the supervised child.
- Impact: avoidable source churn and discovery delay. The saved query position is retained, so no confirmed target loss was observed.
- Candidate direction: bounded direct-request retries plus a source-cycle network disposition that does not terminate the child.
- Validation: inject timeout-then-success and timeout-exhaustion cases; verify query state, auth state, backoff, and no source restart.
P1: GitLab scans have an incomplete RC=1 coverage gap
- Status: resolved by
stabilize-gitlab-trufflehog-lifecycle; the production canary and bounded historical replay completed successfully. - Evidence: 12 of 175 recent GitLab attempts exited RC=1 without a fatal diagnostic or
finished scanningmarker. - Resolution evidence: 120 process attempts passed the rollout gate with 117 RC=0 completion markers and zero recurrence of the old signature. A three-target exact-signature replay completed cleanly on the first new attempt for every target; a further 30-minute soak reached 142 attempts and 139 RC=0 completion markers with the old signature still at zero. A later three-target sample produced two clean completions and one explicit retryable
trufflehogdeferral instead of the old terminal classification. - Current behavior: all 12 were classified as non-retryable
command_exitand became terminal failures after one attempt. - Impact: these repositories have no confirmed complete scan.
- Candidate direction: run controlled Git
--local-devA/B tests, then consider extending explicit completion and bounded retry semantics beyond Docker. - Validation: canary completion-marker rate, unexplained RC=1 rate, partial-finding preservation, retry exhaustion, and non-Git source isolation.
P1: Runtime-wide external Windows termination has no autonomous recovery
- Evidence: at
2026-08-20 19:59:40 MSK, PostgreSQL and unrelated source children were terminated together with Windows exception0x40010004; PostgreSQL shut down after its startup process also failed, and the supervisor was no longer alive to recover it. - Impact: the pipeline remained offline until the next authority-checked
start_runtime.ps1launch. - Current state: PostgreSQL completed WAL recovery without corruption, all core sources and pipeline children restarted, and the post-recovery soak showed no child restarts or database failures.
- Candidate direction: run the authority-checked runtime under an external Windows Service or Task Scheduler watchdog that can restart a dead supervisor without weakening singleton or cluster-identity checks.
Resolved: Updated core targets were permanently suppressed
- Status: resolved by
rescan-updated-core-targets; the change is ready for archive. - Evidence: the first production canary admitted three GitHub and three GitLab updated targets across six separate cycles, never exceeding the configured one-per-cycle cap. All six scans completed cleanly with exact remote revision snapshots, zero errors, zero findings, and applied queue completion.
- Isolation: DockerHub recorded no revision observations or updated admissions and continues to use digest identity. HuggingFace established its newest-modified revision baseline without an updated-target surge.
- Rollout decision: retain the cap at one per source cycle and the 24-hour per-target cooldown until longer-running yield data justifies a change.
Resolved: Core discovery omitted an exact OpenAI query
- Status: resolved by
restore-openai-discovery-coverage; the change is ready for archive. - Evidence: the first bounded cycles fetched 44 GitHub and 42 GitLab results, admitted one changed GitLab target, and discovered 18 new DockerHub repository identities under the exact
openaiquery. Source bounds remained one GitHub/GitLab page with at most five claims and two DockerHub pages with at most 20 claims. - Funnel: all 18 DockerHub admissions reached terminal queue state (16 done, two explicit registry-access failures). Their scans produced 347 findings, including 53 OpenAI finding rows representing 24 distinct credentials. All 24 were genuinely new and received an authoritative API result: 20 quota/no-balance, four invalid/revoked, and zero alive.
- Durability: all 18 scan projections and 53 OpenAI keycheck projections completed with released capacity; the OpenAI candidate and publication backlogs drained to zero.
- Rollout decision: retain the bounded rotating query. It restored missing supply but did not produce a usable credential in the initial sample, so broader limits are not justified yet.
Resolved: Bounded OpenAI ecosystem keyword wave
- Status: implemented by
expand-openai-ecosystem-discovery; all nine source-specific queries completed one bounded production cycle. - Git evidence: the three GitHub queries fetched 6, 0, and 3 results without new or updated admissions. GitLab
openai-apiandopenai-agentsadmitted nothing;librechatadmitted one bounded updated target whose scan completed cleanly without findings. - Docker evidence:
librechat,lobechat, andopenai-proxyadmitted 20, 19, and 16 immutable-digest targets. All 55 reached terminal queue state: 50 done and five explicitauth_invalidfailures. Exact query scans produced 597 finding rows. - Credential funnel:
librechatproduced five genuinely new credentials (GCP two, Gemini one, GitHub one, OpenAI one);lobechatproduced two (AWS one, GitHub one);openai-proxyproduced none. All seven received authoritative API outcomes and none was alive or usable. The OpenAI credential was explicitly invalid/revoked. - Durability: every exact cohort candidate completed and all 147 cohort scan/keycheck projection jobs completed with released capacity.
- Rollout decision: retain the bounded wave provisionally without increasing any page, claim, worker, or global concurrency limit. Reassess after the next natural rotation; remove zero-yield terms if they remain empty rather than expanding this wave.
P2: Heavy Docker images can exceed bounded resources
- Evidence: in the 197-attempt Docker canary, six scans reached the 600-second timeout, three hit explicit Windows
VirtualAlloc/paging limits, and one had a mixed stream/timeout outcome. - Current behavior: failures are explicit, partial findings are retained, and retryable outcomes use the bounded queue policy. The old unexplained RC=1 signature remained at zero.
- Impact: a small set of large images does not receive confirmed full coverage.
- Candidate direction: a separate heavy-image lane with lower concurrency and a larger target budget; do not raise global limits without host-level measurements.
- Validation: completion gain, host commit usage, scan-slot fairness, source throughput, and retry amplification.
P2: Historical Docker replay remains intentionally gradual
- Evidence: the first exact-signature batch of 10 produced seven completed targets, one terminal registry-access failure, and two explicit deferred retries, with no old RC=1 recurrence. A second six-target batch produced four completed targets, one terminal registry-access failure, and one timeout deferral; all 41 emitted finding UIDs were globally unique and the timeout retained 35 partial findings.
- Current behavior: the historical terminal set has not been mass-requeued.
- Candidate direction: increase exact-signature batches conservatively while monitoring registry traffic and heavy-image failures.
P2: GitHub manifests are not mined for Docker image references
- Evidence: 1,954 retained
github_archive_filesartifacts contained 158 unique image references, but only about 48 plausible namespaced Docker Hub repositories were incremental after filtering bases, local names, and existing queue identities. - Current behavior: changed compose/workflow files are scanned for secrets only;
image:,FROM, anddocker://references do not feed Docker resolution. The boundedgithub_archive_filessource is currently disabled. - Impact: a small but potentially fresher source of Docker repositories is omitted. The measured one-time opportunity is much smaller than the existing unresolved Docker repository backlog.
- Candidate direction: after Docker resolver throughput is healthy, parse exact image references from changed compose, Kubernetes/Helm, workflow, and Dockerfile artifacts; resolve concrete tags directly to immutable platform digests and filter common base images.
- Validation: incremental repositories and digests, API calls per admitted target, source freshness, strict usable-key yield, duplicate rate, and GitHub quota cost.
P2: Background supervisor launch nonce can parse as an option
- Evidence: one authority-checked launch on 2026-08-25 generated a URL-safe nonce beginning with an option-like prefix;
argparsereportedargument --launch-nonce: expected one argument. An immediate canonical retry succeeded. - Impact: a rare transient startup refusal; no PostgreSQL or queue mutation occurred.
- Candidate direction: emit the hidden value as
--launch-nonce=<value>and add a command-construction test with a leading-hyphen nonce.
P3: GitLab removed/private repository churn
- Evidence: 10 recent clone-error scan events came from four targets; three targets exhausted all three retries.
- Current behavior: clone preparation errors are retryable even when a repository appears deleted, private, or deletion-scheduled.
- Impact: bounded but avoidable repeated work.
- Candidate direction: distinguish permanent not-found/access outcomes from transient clone transport failures before retrying.
Needs Fresh Validation: Projector recovery/index path
- Earlier investigation suggested a rollback/missing-index weakness in projector recovery.
- Current state is healthy: ingester and projector report
ready, with no active lease error. - Before creating a change, reproduce or recover the original query/error evidence and determine whether the issue still exists.