Initial server source import
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-06-22
|
||||
@@ -0,0 +1,82 @@
|
||||
## Context
|
||||
|
||||
The scanner runs continuously and appends findings to `found_secrets.jsonl` and `scanner_active.db`. Keycheckers classify provider credentials into current-state status files under `runtime/keychecks/<service>/` and opportunistically write rows to `keycheck_results` for dashboard visibility.
|
||||
|
||||
The current behavior has several failure modes:
|
||||
|
||||
- Hourly keychecks can replay a multi-GB `found_secrets.jsonl`, delaying or blocking later services in the batch.
|
||||
- Some checkers implement custom input readers, so global tail behavior does not apply consistently.
|
||||
- Known keys are skipped before writing a new occurrence row, so repeated source/query hits for an already alive key are invisible in DB/dashboard history.
|
||||
- Per-result DB writes compete with scanner writes and can fail under SQLite locks, causing file state and DB observations to diverge.
|
||||
- Dashboard views mix file current-state, DB current-state, historical occurrences, and provider-specific usable access semantics.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make current-state status files remain the authoritative per-service status store.
|
||||
- Record historical occurrences for known keys without forcing API rechecks.
|
||||
- Make hourly keychecks bounded and incremental enough to keep up with continuous scanning.
|
||||
- Make DB writes resilient and explainable, with visible lag/error indicators.
|
||||
- Give operators dashboard presets for current usable keys, historical usable occurrences, no-quota/limited states, and pipeline health.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Replace provider status files with the SQLite DB.
|
||||
- Revalidate every known alive key on every hourly cycle.
|
||||
- Guarantee zero SQLite lock contention while scanner writers are active.
|
||||
- Redesign detector extraction or provider-specific validity semantics beyond accounting and visibility.
|
||||
|
||||
## Decisions
|
||||
|
||||
1. Keep status files as current-state truth.
|
||||
|
||||
Rationale: checkers already compact/move keys between `Alive`, `NoBalance`, `Dead`, `Network`, and related files. Replacing this would be riskier than making DB observations catch up.
|
||||
|
||||
Alternative considered: make `keycheck_results` the source of truth. Rejected because the active DB is large, frequently locked, and dashboard reads must not block scanner writes.
|
||||
|
||||
2. Introduce occurrence recording for skipped known keys.
|
||||
|
||||
When a checker sees a candidate that is already known in status files or checked files, it should write a lightweight occurrence row with the cached status, source line, and finding attribution. It must not call provider APIs unless retry flags or recheck flags require it.
|
||||
|
||||
Alternative considered: only record fresh API checks. Rejected because this hides repeated source/query yield for already alive keys.
|
||||
|
||||
3. Use a shared bounded/incremental input reader for all checkers.
|
||||
|
||||
Checkers should use common reader helpers rather than hand-rolled full-file loops. The minimum implementation can use a tail window; the target implementation should store per-service high-watermark offsets so hourly runs neither replay old data nor miss data outside a fixed tail window.
|
||||
|
||||
Alternative considered: keep reading full JSONL and rely on skip sets. Rejected because the input is already multi-GB and causes long stalls.
|
||||
|
||||
4. Centralize DB recording or make per-checker DB writes lock-tolerant.
|
||||
|
||||
The preferred direction is batching result/occurrence rows through `keycheck_runner` after each service completes. A smaller intermediate step is to avoid schema initialization on every single keycheck write and retry lock failures with bounded backoff.
|
||||
|
||||
Alternative considered: ignore DB write failures because files are authoritative. Rejected because dashboard and source attribution depend on DB visibility.
|
||||
|
||||
5. Distinguish access tiers from raw provider statuses.
|
||||
|
||||
Dashboard should classify provider statuses into operator-facing tiers such as `usable_llm`, `alive_unproven_llm`, `no_quota`, and `quota_limited`, while still allowing exact status filtering for values like `BEDROCK`, `VERTEX`, `VALID_RATE_LIMITED`, and `ALIVE`.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Cached occurrence rows could be mistaken for fresh provider rechecks -> Label occurrence rows with a result source such as `cached_status` versus `api_check`.
|
||||
- Tail windows can miss old-but-newly-unchecked lines after downtime or file rewrites -> Prefer high-watermark offsets and detect file truncation/rotation.
|
||||
- DB batching can still fail if SQLite is locked for extended periods -> Keep files authoritative and surface DB write lag/errors in dashboard.
|
||||
- Provider semantics differ: e.g. Gemini `VALID_RATE_LIMITED`, AWS `BEDROCK`, GCP `VERTEX` -> Keep exact statuses available and use access tiers only as an additional view.
|
||||
- Rechecking known alive keys too often can spend quota or trigger provider limits -> Occurrence recording must not imply revalidation.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add shared input high-watermark/tail behavior to all keycheckers, starting with hand-rolled readers.
|
||||
2. Add cached occurrence recording for known keys using existing status maps.
|
||||
3. Add batched DB write path or harden lock retry behavior.
|
||||
4. Update dashboard to show file current-state, DB observation freshness, and historical/current presets separately.
|
||||
5. Backfill/repair missing occurrence attribution from current status files and recent findings where safe.
|
||||
|
||||
Rollback: disable cached occurrence writes and fall back to existing status-file behavior; status files remain unchanged.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should high-watermark state live in `runtime/keychecks/<service>/state.json` or a shared `runtime/state/keycheck_offsets.json`?
|
||||
- Should cached occurrences be written for every repeated finding or deduped per service/key/finding/source per day?
|
||||
- Which dashboard panel should be considered the primary operator view: file current-state or DB latest-current-state?
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
Keycheck current-state files, DB observations, and dashboard views can diverge, making it unclear whether usable provider keys are still being found and which sources produced them. This is urgent because the scanner is running continuously, but large input files, known-key skipping, and SQLite lock behavior can hide fresh usable findings from operator-visible stats.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add reliable keycheck accounting for fresh checks and known-key occurrences.
|
||||
- Track current-state file counts separately from DB-observed validation rows.
|
||||
- Ensure hourly keychecks process recent findings efficiently without replaying multi-GB JSONL inputs from the beginning.
|
||||
- Preserve source/query/finding attribution even when a key was already classified as alive, dead, no-balance, or limited.
|
||||
- Surface keycheck pipeline health, write lag, skipped-known counts, and usable/no-quota status in dashboard views.
|
||||
- Reduce DB lock impact on keycheck result recording so file state and DB state remain explainably consistent.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `keycheck-accounting`: Defines reliable current-state, historical occurrence, and dashboard visibility behavior for keycheck results.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
## Impact
|
||||
|
||||
- `app/keycheck_runner.py` and `app/keycheckers/*`: keycheck execution, input reading, skip behavior, and result recording.
|
||||
- `app/scanner_db.py`: keycheck DB writes, lock handling, and occurrence recording.
|
||||
- `app/dashboard.py`: operator-facing keycheck and usable-key reporting.
|
||||
- Runtime files under `runtime/keychecks/`: current-state status files remain authoritative but gain clearer relationship to DB observations.
|
||||
- Runtime DB `scanner_active.db`: keycheck rows and attribution semantics become more complete and auditable.
|
||||
@@ -0,0 +1,76 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Current-state status files remain authoritative
|
||||
The system SHALL keep per-service keycheck status files as the authoritative current-state classification for keys.
|
||||
|
||||
#### Scenario: Key status changes after recheck
|
||||
- **WHEN** a checker revalidates a key and receives a new status
|
||||
- **THEN** the key MUST be removed from other status files for that service and written to the status file for the new status
|
||||
|
||||
#### Scenario: Dashboard compares files and DB
|
||||
- **WHEN** dashboard displays keycheck totals
|
||||
- **THEN** it MUST be clear whether each count comes from current-state status files or from DB observation rows
|
||||
|
||||
### Requirement: Known key occurrences are recorded
|
||||
The system SHALL record an occurrence when a checker sees a candidate that is already known in service status files or checked files.
|
||||
|
||||
#### Scenario: Known alive key appears in a new finding
|
||||
- **WHEN** a key already classified as alive appears in a new scanner finding
|
||||
- **THEN** the system MUST record the new source/query/finding occurrence without requiring a provider API recheck
|
||||
|
||||
#### Scenario: Known dead key appears in a new finding
|
||||
- **WHEN** a key already classified as dead appears in a new scanner finding
|
||||
- **THEN** the system MUST record the new source/query/finding occurrence with a cached dead status
|
||||
|
||||
#### Scenario: Cached occurrence is distinguishable from API recheck
|
||||
- **WHEN** an occurrence row is written without calling the provider API
|
||||
- **THEN** the row MUST indicate that the status came from cached current-state classification
|
||||
|
||||
### Requirement: Keycheck input processing is bounded and consistent
|
||||
The system SHALL avoid replaying the full scanner JSONL input on every hourly keycheck run.
|
||||
|
||||
#### Scenario: Hourly keychecks run on a multi-GB input file
|
||||
- **WHEN** `found_secrets.jsonl` is large
|
||||
- **THEN** each checker MUST process only a bounded recent range or an incremental range since its last processed offset
|
||||
|
||||
#### Scenario: Checker has a custom input loop
|
||||
- **WHEN** a checker reads scanner findings
|
||||
- **THEN** it MUST use shared keycheck input-reading behavior or implement equivalent high-watermark/tail semantics
|
||||
|
||||
#### Scenario: Input file rotates or shrinks
|
||||
- **WHEN** a stored high-watermark offset is larger than the current input file size
|
||||
- **THEN** the system MUST reset the offset safely and continue processing without crashing
|
||||
|
||||
### Requirement: Keycheck DB observation writes are resilient
|
||||
The system SHALL make keycheck DB observation writes resilient to active scanner DB contention.
|
||||
|
||||
#### Scenario: SQLite database is temporarily locked
|
||||
- **WHEN** a keycheck result or occurrence is ready to record and SQLite is locked
|
||||
- **THEN** the system MUST retry with bounded backoff before reporting a DB write failure
|
||||
|
||||
#### Scenario: DB write fails after retries
|
||||
- **WHEN** all DB write retries fail
|
||||
- **THEN** the status file write MUST remain intact and the failure MUST be visible in logs or dashboard health
|
||||
|
||||
#### Scenario: Schema initialization would contend with active writers
|
||||
- **WHEN** a checker records a single result row
|
||||
- **THEN** it MUST NOT run schema initialization or migration DDL as part of that per-result write path
|
||||
|
||||
### Requirement: Dashboard exposes keycheck pipeline health
|
||||
The dashboard SHALL expose keycheck pipeline health and freshness separately from provider status counts.
|
||||
|
||||
#### Scenario: DB observations lag behind status files
|
||||
- **WHEN** status files are newer than the latest DB keycheck row
|
||||
- **THEN** dashboard MUST show that DB observation data is stale relative to file current-state
|
||||
|
||||
#### Scenario: Keycheck run is stuck on a service
|
||||
- **WHEN** the keychecks process has not advanced past a service for longer than expected
|
||||
- **THEN** dashboard or supervisor-visible status MUST make the stuck service and elapsed time visible
|
||||
|
||||
#### Scenario: Operator wants current usable keys
|
||||
- **WHEN** an operator selects current usable key view
|
||||
- **THEN** dashboard MUST use provider-specific access tiers while retaining exact status filters such as `BEDROCK`, `VERTEX`, `ALIVE`, and `VALID_RATE_LIMITED`
|
||||
|
||||
#### Scenario: Operator wants historical source yield
|
||||
- **WHEN** an operator selects historical yield view
|
||||
- **THEN** dashboard MUST include cached known-key occurrences so source/query yield is not lost after rechecks or known-key skips
|
||||
@@ -0,0 +1,36 @@
|
||||
## 1. Input Processing
|
||||
|
||||
- [x] 1.1 Add shared keycheck input state storage for per-service file path, file size, inode/signature if available, and last processed byte offset.
|
||||
- [x] 1.2 Extend shared keycheck input reader to support high-watermark processing with safe reset on file truncation or rotation.
|
||||
- [x] 1.3 Migrate hand-rolled readers in OpenAI, OpenRouter, and Gemini to the shared input reader.
|
||||
- [x] 1.4 Keep a bounded tail fallback for first run or missing state, with clear logging of the active input mode.
|
||||
|
||||
## 2. Known-Key Occurrence Recording
|
||||
|
||||
- [x] 2.1 Add a shared helper that resolves cached status for a key from checked/status files without provider API calls.
|
||||
- [x] 2.2 Add occurrence recording for skipped known keys, including service, cached status, source line, finding payload, and detector.
|
||||
- [x] 2.3 Mark cached occurrence rows distinctly from API-check rows in metadata.
|
||||
- [x] 2.4 Deduplicate cached occurrences per service/key/finding/source to avoid unbounded repeated rows.
|
||||
|
||||
## 3. DB Write Reliability
|
||||
|
||||
- [x] 3.1 Ensure per-result keycheck DB writes do not run schema initialization or migration DDL.
|
||||
- [x] 3.2 Add bounded lock retry/backoff for keycheck result and occurrence writes.
|
||||
- [x] 3.3 Add keycheck runner counters for DB write success, retry, and failure counts per service.
|
||||
- [x] 3.4 Log DB write failures without preventing status-file updates.
|
||||
|
||||
## 4. Dashboard Visibility
|
||||
|
||||
- [x] 4.1 Add dashboard panel comparing status-file current-state counts with latest DB observation timestamps.
|
||||
- [x] 4.2 Add keycheck pipeline health panel showing last completed service, current/stuck service, run duration, and DB write errors.
|
||||
- [x] 4.3 Add current usable-key view based on provider-specific access tiers and exact statuses.
|
||||
- [x] 4.4 Add historical source-yield view that includes cached known-key occurrences.
|
||||
- [x] 4.5 Label cached-status rows separately from fresh API-check rows in validation tables.
|
||||
|
||||
## 5. Verification
|
||||
|
||||
- [x] 5.1 Add a smoke test or script that runs a checker over a small fixture and verifies status-file writes plus DB occurrence rows.
|
||||
- [x] 5.2 Verify OpenAI, OpenRouter, Gemini, AWS, and GCP checkers use bounded/incremental input processing.
|
||||
- [x] 5.3 Verify a known alive key found in a new source/query creates a cached occurrence without an API recheck.
|
||||
- [x] 5.4 Verify dashboard can show current-state file counts even when DB observations are stale.
|
||||
- [x] 5.5 Run `python -m py_compile` on changed Python modules.
|
||||
Reference in New Issue
Block a user