83 lines
5.7 KiB
Markdown
83 lines
5.7 KiB
Markdown
## Context
|
|
|
|
The scanner runs continuously and appends findings to `found_secrets.jsonl` and `scanner_active.db`. Keycheckers classify provider credentials into current-state status files under `runtime/keychecks/<service>/` and opportunistically write rows to `keycheck_results` for dashboard visibility.
|
|
|
|
The current behavior has several failure modes:
|
|
|
|
- Hourly keychecks can replay a multi-GB `found_secrets.jsonl`, delaying or blocking later services in the batch.
|
|
- Some checkers implement custom input readers, so global tail behavior does not apply consistently.
|
|
- Known keys are skipped before writing a new occurrence row, so repeated source/query hits for an already alive key are invisible in DB/dashboard history.
|
|
- Per-result DB writes compete with scanner writes and can fail under SQLite locks, causing file state and DB observations to diverge.
|
|
- Dashboard views mix file current-state, DB current-state, historical occurrences, and provider-specific usable access semantics.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals:**
|
|
|
|
- Make current-state status files remain the authoritative per-service status store.
|
|
- Record historical occurrences for known keys without forcing API rechecks.
|
|
- Make hourly keychecks bounded and incremental enough to keep up with continuous scanning.
|
|
- Make DB writes resilient and explainable, with visible lag/error indicators.
|
|
- Give operators dashboard presets for current usable keys, historical usable occurrences, no-quota/limited states, and pipeline health.
|
|
|
|
**Non-Goals:**
|
|
|
|
- Replace provider status files with the SQLite DB.
|
|
- Revalidate every known alive key on every hourly cycle.
|
|
- Guarantee zero SQLite lock contention while scanner writers are active.
|
|
- Redesign detector extraction or provider-specific validity semantics beyond accounting and visibility.
|
|
|
|
## Decisions
|
|
|
|
1. Keep status files as current-state truth.
|
|
|
|
Rationale: checkers already compact/move keys between `Alive`, `NoBalance`, `Dead`, `Network`, and related files. Replacing this would be riskier than making DB observations catch up.
|
|
|
|
Alternative considered: make `keycheck_results` the source of truth. Rejected because the active DB is large, frequently locked, and dashboard reads must not block scanner writes.
|
|
|
|
2. Introduce occurrence recording for skipped known keys.
|
|
|
|
When a checker sees a candidate that is already known in status files or checked files, it should write a lightweight occurrence row with the cached status, source line, and finding attribution. It must not call provider APIs unless retry flags or recheck flags require it.
|
|
|
|
Alternative considered: only record fresh API checks. Rejected because this hides repeated source/query yield for already alive keys.
|
|
|
|
3. Use a shared bounded/incremental input reader for all checkers.
|
|
|
|
Checkers should use common reader helpers rather than hand-rolled full-file loops. The minimum implementation can use a tail window; the target implementation should store per-service high-watermark offsets so hourly runs neither replay old data nor miss data outside a fixed tail window.
|
|
|
|
Alternative considered: keep reading full JSONL and rely on skip sets. Rejected because the input is already multi-GB and causes long stalls.
|
|
|
|
4. Centralize DB recording or make per-checker DB writes lock-tolerant.
|
|
|
|
The preferred direction is batching result/occurrence rows through `keycheck_runner` after each service completes. A smaller intermediate step is to avoid schema initialization on every single keycheck write and retry lock failures with bounded backoff.
|
|
|
|
Alternative considered: ignore DB write failures because files are authoritative. Rejected because dashboard and source attribution depend on DB visibility.
|
|
|
|
5. Distinguish access tiers from raw provider statuses.
|
|
|
|
Dashboard should classify provider statuses into operator-facing tiers such as `usable_llm`, `alive_unproven_llm`, `no_quota`, and `quota_limited`, while still allowing exact status filtering for values like `BEDROCK`, `VERTEX`, `VALID_RATE_LIMITED`, and `ALIVE`.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
- Cached occurrence rows could be mistaken for fresh provider rechecks -> Label occurrence rows with a result source such as `cached_status` versus `api_check`.
|
|
- Tail windows can miss old-but-newly-unchecked lines after downtime or file rewrites -> Prefer high-watermark offsets and detect file truncation/rotation.
|
|
- DB batching can still fail if SQLite is locked for extended periods -> Keep files authoritative and surface DB write lag/errors in dashboard.
|
|
- Provider semantics differ: e.g. Gemini `VALID_RATE_LIMITED`, AWS `BEDROCK`, GCP `VERTEX` -> Keep exact statuses available and use access tiers only as an additional view.
|
|
- Rechecking known alive keys too often can spend quota or trigger provider limits -> Occurrence recording must not imply revalidation.
|
|
|
|
## Migration Plan
|
|
|
|
1. Add shared input high-watermark/tail behavior to all keycheckers, starting with hand-rolled readers.
|
|
2. Add cached occurrence recording for known keys using existing status maps.
|
|
3. Add batched DB write path or harden lock retry behavior.
|
|
4. Update dashboard to show file current-state, DB observation freshness, and historical/current presets separately.
|
|
5. Backfill/repair missing occurrence attribution from current status files and recent findings where safe.
|
|
|
|
Rollback: disable cached occurrence writes and fall back to existing status-file behavior; status files remain unchanged.
|
|
|
|
## Open Questions
|
|
|
|
- Should high-watermark state live in `runtime/keychecks/<service>/state.json` or a shared `runtime/state/keycheck_offsets.json`?
|
|
- Should cached occurrences be written for every repeated finding or deduped per service/key/finding/source per day?
|
|
- Which dashboard panel should be considered the primary operator view: file current-state or DB latest-current-state?
|