5.7 KiB
Context
The scanner runs continuously and appends findings to found_secrets.jsonl and scanner_active.db. Keycheckers classify provider credentials into current-state status files under runtime/keychecks/<service>/ and opportunistically write rows to keycheck_results for dashboard visibility.
The current behavior has several failure modes:
- Hourly keychecks can replay a multi-GB
found_secrets.jsonl, delaying or blocking later services in the batch. - Some checkers implement custom input readers, so global tail behavior does not apply consistently.
- Known keys are skipped before writing a new occurrence row, so repeated source/query hits for an already alive key are invisible in DB/dashboard history.
- Per-result DB writes compete with scanner writes and can fail under SQLite locks, causing file state and DB observations to diverge.
- Dashboard views mix file current-state, DB current-state, historical occurrences, and provider-specific usable access semantics.
Goals / Non-Goals
Goals:
- Make current-state status files remain the authoritative per-service status store.
- Record historical occurrences for known keys without forcing API rechecks.
- Make hourly keychecks bounded and incremental enough to keep up with continuous scanning.
- Make DB writes resilient and explainable, with visible lag/error indicators.
- Give operators dashboard presets for current usable keys, historical usable occurrences, no-quota/limited states, and pipeline health.
Non-Goals:
- Replace provider status files with the SQLite DB.
- Revalidate every known alive key on every hourly cycle.
- Guarantee zero SQLite lock contention while scanner writers are active.
- Redesign detector extraction or provider-specific validity semantics beyond accounting and visibility.
Decisions
-
Keep status files as current-state truth.
Rationale: checkers already compact/move keys between
Alive,NoBalance,Dead,Network, and related files. Replacing this would be riskier than making DB observations catch up.Alternative considered: make
keycheck_resultsthe source of truth. Rejected because the active DB is large, frequently locked, and dashboard reads must not block scanner writes. -
Introduce occurrence recording for skipped known keys.
When a checker sees a candidate that is already known in status files or checked files, it should write a lightweight occurrence row with the cached status, source line, and finding attribution. It must not call provider APIs unless retry flags or recheck flags require it.
Alternative considered: only record fresh API checks. Rejected because this hides repeated source/query yield for already alive keys.
-
Use a shared bounded/incremental input reader for all checkers.
Checkers should use common reader helpers rather than hand-rolled full-file loops. The minimum implementation can use a tail window; the target implementation should store per-service high-watermark offsets so hourly runs neither replay old data nor miss data outside a fixed tail window.
Alternative considered: keep reading full JSONL and rely on skip sets. Rejected because the input is already multi-GB and causes long stalls.
-
Centralize DB recording or make per-checker DB writes lock-tolerant.
The preferred direction is batching result/occurrence rows through
keycheck_runnerafter each service completes. A smaller intermediate step is to avoid schema initialization on every single keycheck write and retry lock failures with bounded backoff.Alternative considered: ignore DB write failures because files are authoritative. Rejected because dashboard and source attribution depend on DB visibility.
-
Distinguish access tiers from raw provider statuses.
Dashboard should classify provider statuses into operator-facing tiers such as
usable_llm,alive_unproven_llm,no_quota, andquota_limited, while still allowing exact status filtering for values likeBEDROCK,VERTEX,VALID_RATE_LIMITED, andALIVE.
Risks / Trade-offs
- Cached occurrence rows could be mistaken for fresh provider rechecks -> Label occurrence rows with a result source such as
cached_statusversusapi_check. - Tail windows can miss old-but-newly-unchecked lines after downtime or file rewrites -> Prefer high-watermark offsets and detect file truncation/rotation.
- DB batching can still fail if SQLite is locked for extended periods -> Keep files authoritative and surface DB write lag/errors in dashboard.
- Provider semantics differ: e.g. Gemini
VALID_RATE_LIMITED, AWSBEDROCK, GCPVERTEX-> Keep exact statuses available and use access tiers only as an additional view. - Rechecking known alive keys too often can spend quota or trigger provider limits -> Occurrence recording must not imply revalidation.
Migration Plan
- Add shared input high-watermark/tail behavior to all keycheckers, starting with hand-rolled readers.
- Add cached occurrence recording for known keys using existing status maps.
- Add batched DB write path or harden lock retry behavior.
- Update dashboard to show file current-state, DB observation freshness, and historical/current presets separately.
- Backfill/repair missing occurrence attribution from current status files and recent findings where safe.
Rollback: disable cached occurrence writes and fall back to existing status-file behavior; status files remain unchanged.
Open Questions
- Should high-watermark state live in
runtime/keychecks/<service>/state.jsonor a sharedruntime/state/keycheck_offsets.json? - Should cached occurrences be written for every repeated finding or deduped per service/key/finding/source per day?
- Which dashboard panel should be considered the primary operator view: file current-state or DB latest-current-state?