Initial server source import
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-06-15
|
||||
@@ -0,0 +1,89 @@
|
||||
## Context
|
||||
|
||||
> Supersession, 2026-07-27: the approved cumulative-lag/memory/backpressure architecture supersedes this change's earlier no-new-queue and conservative-concurrency constraints. PostgreSQL is now sole authority, sources hand off S:-backed durable bundles, JSONL/status files are asynchronous projections, normal keychecks use PostgreSQL candidates, and the fair global scan limit is three. Process recycling, GC trimming, and reduced concurrency are not correctness mechanisms.
|
||||
|
||||
The scanner runs multiple supervised source processes that share runtime state, queue files, logs, and an observability database. Recent logs show source restarts caused by Windows `PermissionError` during state-file replacement, while dashboard and supervisor read the same state files for status display. Coverage is also limited by conservative artifact size caps and by discovery inputs that are broad but not always high-signal.
|
||||
|
||||
The original incremental constraints below remain historical context only where contradicted by the supersession above.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Prevent transient Windows state-file read/write races from crashing source processes.
|
||||
- Increase scan coverage through configurable size limit bumps while preserving conservative concurrency.
|
||||
- Improve discovery signal by focusing metadata searches on provider, host, and framework terms.
|
||||
- Strengthen `package_git` as a discovery backbone through better repository URL extraction and canonicalization.
|
||||
- Make GitHub Actions and GitLab CI scanning use better parsed seeds and modestly larger per-cycle target sets.
|
||||
- Continue provider-specific detector and keychecker additions with safe validation and context routing.
|
||||
|
||||
**Non-Goals:**
|
||||
- No deferred/deep queue in this change.
|
||||
- No new checked/skipped classification model in this change.
|
||||
- No scheduler rewrite or central scoring engine.
|
||||
- No broad generic `api_key`, `secret`, or `token` terms in normal repo/package metadata discovery by default.
|
||||
- No removal of existing command-line entry points or queue file formats.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Use retrying unique-temp state writes instead of locking
|
||||
|
||||
State writes will use a unique temporary file name and retry `os.replace` on transient Windows permission failures. This keeps the existing JSON state model and avoids cross-process lock files or SQLite migration.
|
||||
|
||||
Alternatives considered:
|
||||
- Lock files: rejected for now because they add another failure mode and require all readers/writers to cooperate.
|
||||
- SQLite state: rejected as too large for this incremental reliability fix.
|
||||
- Ignoring failed state writes: rejected because auth/query state must remain observable.
|
||||
|
||||
### Increase limits through configuration first
|
||||
|
||||
Package, Postman, and CI artifact limits will be increased in config while keeping source worker counts conservative. The implementation will not introduce deferred queues or new outcome classes; oversized skips remain visible through existing logs and target scan records.
|
||||
|
||||
Alternatives considered:
|
||||
- Remove limits entirely: rejected because large archives can exhaust disk, CPU, and scan slots.
|
||||
- Add deep-lane queues: deferred to a separate design because it needs stronger classification semantics.
|
||||
|
||||
### Keep metadata discovery provider-focused
|
||||
|
||||
Repository/package metadata search should use provider names, API hosts, framework terms, and ecosystem terms. Exact secret variable terms should be reserved for code/artifact-oriented searches where content is actually searched.
|
||||
|
||||
Alternatives considered:
|
||||
- Add `api_key` globally: rejected because most sources search names/descriptions/readmes and this produces low-signal security-tool/tutorial results.
|
||||
|
||||
### Improve existing CI seed flow before increasing volume
|
||||
|
||||
CI sources should first parse more existing seed formats from DB records, package candidates, and findings. After parsing improves, `ci_seed_scan_limit` and per-cycle target counts can be raised modestly.
|
||||
|
||||
Alternatives considered:
|
||||
- Increase CI volume immediately: rejected because current logs show many unparseable/known seeds, so raw volume would mostly amplify waste.
|
||||
|
||||
### Add provider support as detector plus keychecker pairs
|
||||
|
||||
Provider expansion should follow the Qwen/DashScope pattern: contextual detector, safe keychecker, and routing safeguards for ambiguous `sk-...` formats.
|
||||
|
||||
Alternatives considered:
|
||||
- Detector-only additions: rejected for providers where validation is feasible because they increase unverified noise.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Windows state retry may hide a persistent file access problem for a few hundred milliseconds -> surface the final error after bounded retries.
|
||||
- Higher artifact size caps increase runtime and disk pressure -> keep workers conservative and rely on existing scan slot limits.
|
||||
- Provider-focused queries may miss generic projects that leak keys -> use package/artifact/code-oriented paths for exact env var searches instead of metadata search.
|
||||
- CI seed parsing improvements may still leave many stale/known targets -> postpone TTL/revisit policy until outcome semantics are revisited.
|
||||
- Context routing can misclassify ambiguous keys when context is weak -> only route away from a provider when strong provider-specific context is present.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Apply state write retry first and monitor source restarts.
|
||||
2. Increase size caps in config and monitor disk usage, scan duration, and skipped/error counts.
|
||||
3. Adjust discovery query lists and verify target volume remains healthy.
|
||||
4. Improve `package_git` metadata URL extraction/canonicalization.
|
||||
5. Improve CI seed parsing, then raise CI limits modestly.
|
||||
6. Add the next provider detector/keychecker pair using the established pattern.
|
||||
|
||||
Rollback is straightforward for each step: revert the code change for state writes, restore previous config caps/query lists, or disable individual CI/provider changes.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should oversized artifacts remain checked under current semantics, or should that become a separate future change?
|
||||
- Which provider detector/keychecker pair should be prioritized after Qwen/DashScope?
|
||||
- What disk and scan-duration thresholds should trigger reducing size caps again?
|
||||
@@ -0,0 +1,28 @@
|
||||
## Why
|
||||
|
||||
The scanner is now broad enough that small reliability issues and conservative defaults are limiting throughput: source uptime resets from Windows state-file races, large artifacts are skipped too aggressively, and CI/package discovery often spends cycles on low-value or already-known targets. This change stabilizes source execution and widens coverage incrementally without introducing new queues or a scheduler rewrite.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Make runner state persistence resilient to transient Windows file-lock races so sources do not crash when supervisor or dashboard reads state concurrently.
|
||||
- Increase configurable size limits for package, Postman, and CI artifacts in a controlled way while keeping worker counts conservative.
|
||||
- Refine discovery queries so repository/package metadata searches focus on provider, host, and framework terms instead of generic secret words.
|
||||
- Improve `package_git` discovery quality by broadening metadata extraction and canonicalization while preserving the current queue model.
|
||||
- Improve CI source target selection by parsing more seed formats and scanning more relevant GitHub Actions and GitLab CI repositories per cycle.
|
||||
- Continue the provider-specific detector plus keychecker pattern, including context-based routing for generic key formats.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `scanner-runtime-stability`: resilient source state persistence and restart behavior for supervised scanner processes.
|
||||
- `scan-coverage-sizing`: configurable artifact size coverage for packages, Postman artifacts, and CI logs/artifacts.
|
||||
- `source-discovery-targeting`: higher-signal source queries, package git discovery, and CI seed target selection.
|
||||
- `provider-key-validation`: provider-specific custom detectors, keycheckers, and context routing for ambiguous key formats.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `app/console_runner.py`, `app/supervisor.py`, `app/dashboard.py`, `app/scanner.py`, `app/scanner_db.py`, `app/config.yaml`, and selected `app/keycheckers/**` modules.
|
||||
- Affected runtime data: `runtime/state/runner_state_*.json`, `runtime/queues/*`, scanner logs, target scan records, and keycheck results.
|
||||
- No breaking changes to command-line entry points or existing queue file formats are intended.
|
||||
+48
@@ -0,0 +1,48 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: PostgreSQL Candidate And Current-State Authority
|
||||
Normal provider workers SHALL claim one fenced PostgreSQL candidate at a time and SHALL transactionally insert the result, update `keycheck_current_state`, complete that candidate, and create a projection job.
|
||||
|
||||
#### Scenario: Compatibility files lag or fail
|
||||
- **WHEN** keycheck JSONL or status projection is delayed or quarantined
|
||||
- **THEN** the committed PostgreSQL result and current state SHALL remain authoritative and unrelated candidates SHALL continue
|
||||
|
||||
#### Scenario: Explicit compatibility input
|
||||
- **WHEN** an operator selects `--input-mode jsonl`
|
||||
- **THEN** the bounded legacy JSONL reader MAY be used, but `--input` SHALL be rejected in normal PostgreSQL mode
|
||||
|
||||
### Requirement: Provider Support Uses Detector And Keychecker Pair
|
||||
New provider API key support SHALL include both detection and validation when a safe validation endpoint is available.
|
||||
|
||||
#### Scenario: Provider has safe validation endpoint
|
||||
- **WHEN** support is added for a provider with a non-generating authentication or model-list endpoint
|
||||
- **THEN** the change SHALL include a detector or detector routing rule and a keychecker for that provider
|
||||
|
||||
#### Scenario: Provider has no safe validation endpoint
|
||||
- **WHEN** a provider does not have a safe validation endpoint
|
||||
- **THEN** the detector MAY be added only with explicit documentation that validation is unavailable
|
||||
|
||||
### Requirement: Ambiguous Key Formats Use Context Routing
|
||||
The scanner SHALL use nearby provider-specific context to route ambiguous key formats to the correct keychecker when possible.
|
||||
|
||||
#### Scenario: Qwen context around generic key
|
||||
- **WHEN** an `sk-...` key is found near Qwen or DashScope context
|
||||
- **THEN** the key SHALL be routed away from unrelated generic `sk-...` checkers such as DeepSeek when the context does not also identify that provider
|
||||
|
||||
#### Scenario: Weak context around generic key
|
||||
- **WHEN** an ambiguous key has no strong provider-specific context
|
||||
- **THEN** the scanner SHALL avoid speculative rerouting that would suppress the existing detector result
|
||||
|
||||
### Requirement: Keycheckers Avoid Token-Generating Probes By Default
|
||||
Provider keycheckers SHALL prefer safe non-generating validation endpoints where available.
|
||||
|
||||
#### Scenario: Provider exposes model-list endpoint
|
||||
- **WHEN** a provider exposes an authenticated model-list or account-status endpoint
|
||||
- **THEN** the keychecker SHALL use that endpoint before considering any generation-style probe
|
||||
|
||||
### Requirement: Keycheck Results Remain Linkable To Findings
|
||||
Provider keycheckers SHALL emit detector names and result metadata that can be linked back to scanner findings.
|
||||
|
||||
#### Scenario: Custom detector result is validated
|
||||
- **WHEN** a keychecker validates a key from a custom detector finding
|
||||
- **THEN** the result SHALL include a stable detector name and source metadata sufficient for DB link repair
|
||||
+33
@@ -0,0 +1,33 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Configurable Larger Artifact Coverage
|
||||
The scanner SHALL allow package, Postman, GitHub Actions, and GitLab CI artifact size limits to be raised through configuration without code changes.
|
||||
|
||||
#### Scenario: Package artifact limit is increased
|
||||
- **WHEN** npm or PyPI `max_artifact_size_mb` is configured to a larger value
|
||||
- **THEN** package scans SHALL use the configured size limit for download and scan decisions
|
||||
|
||||
#### Scenario: CI artifact limits are increased
|
||||
- **WHEN** GitHub Actions or GitLab CI artifact archive and file limits are configured to larger values
|
||||
- **THEN** CI scans SHALL use those configured limits for artifact download and extraction decisions
|
||||
|
||||
### Requirement: Conservative Concurrency Preserved
|
||||
The scanner SHALL preserve per-source worker controls so larger artifact limits do not automatically increase concurrent heavy scans.
|
||||
|
||||
#### Scenario: Size limits increase with unchanged workers
|
||||
- **WHEN** artifact size limits are raised in configuration and worker counts are unchanged
|
||||
- **THEN** the scanner SHALL keep using the configured worker counts for the affected source
|
||||
|
||||
### Requirement: Oversized Artifact Visibility
|
||||
The scanner SHALL record existing oversized-artifact skip reasons in logs and target scan records using the current result model.
|
||||
|
||||
#### Scenario: Artifact remains over configured limit
|
||||
- **WHEN** an artifact exceeds the configured size limit
|
||||
- **THEN** the scanner SHALL record a skipped result with the size-limit reason using existing logging and target scan recording paths
|
||||
|
||||
### Requirement: No New Queue Semantics
|
||||
The scanner SHALL NOT introduce a deferred or deep artifact queue as part of this change.
|
||||
|
||||
#### Scenario: Artifact is too large for current run
|
||||
- **WHEN** an artifact is skipped because it exceeds the configured limit
|
||||
- **THEN** the scanner SHALL handle it through the current scan result and queue behavior without creating a new queue type
|
||||
+40
@@ -0,0 +1,40 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Durable Asynchronous Result Handoff Supersedes Source Publication
|
||||
PostgreSQL SHALL be the sole result authority. A fair global permit count of three SHALL end after a pre-reserved result bundle is fsynced and atomically renamed on `S:`, without covering database ingest, JSONL projection, or keychecks.
|
||||
|
||||
#### Scenario: Projector or database is blocked after bundle handoff
|
||||
- **WHEN** a source completes the durable ready rename
|
||||
- **THEN** its scan permit SHALL be released while the bundle remains recoverable and downstream backlog applies capacity-based admission pressure
|
||||
|
||||
### Requirement: Correctness Does Not Depend On Recycling
|
||||
The runtime SHALL NOT use GC trimming, source lifetime limits, private-memory restarts, periodic recycling, reduced concurrency, or time-based restarts to preserve correctness.
|
||||
|
||||
#### Scenario: Runtime memory grows after warmup
|
||||
- **WHEN** a managed process reports increasing private memory
|
||||
- **THEN** admission and durable queue capacity SHALL provide backpressure without recycling the process or reducing the configured three scan permits
|
||||
|
||||
### Requirement: Resilient Runner State Writes
|
||||
The scanner SHALL persist runner state using a unique temporary file per write attempt and SHALL retry replacement when the operating system reports a transient file access error.
|
||||
|
||||
#### Scenario: Concurrent state read during write
|
||||
- **WHEN** a supervised source writes `runner_state_<source>.json` while supervisor or dashboard reads the same state file
|
||||
- **THEN** the source process SHALL retry the replacement and continue without crashing when the file becomes available within the retry window
|
||||
|
||||
#### Scenario: Persistent state write failure
|
||||
- **WHEN** the state file cannot be replaced after the bounded retry window
|
||||
- **THEN** the source process SHALL surface the final write error instead of silently discarding state changes
|
||||
|
||||
### Requirement: State Writes Avoid Shared Temp Path Contention
|
||||
The scanner SHALL avoid using a single shared `.tmp` path for repeated state writes from supervised processes.
|
||||
|
||||
#### Scenario: Multiple state write attempts overlap
|
||||
- **WHEN** two state write attempts occur close together for the same state file
|
||||
- **THEN** each attempt SHALL use a distinct temporary file path before replacing the final state file
|
||||
|
||||
### Requirement: Restart Noise Reduction
|
||||
The supervisor SHALL no longer restart sources due only to transient state-file replacement races that resolve within the retry window.
|
||||
|
||||
#### Scenario: GitLab state replacement race resolves
|
||||
- **WHEN** the GitLab source hits a transient Windows file lock while saving state
|
||||
- **THEN** the source SHALL complete the state save after retry and its `up` timer SHALL not reset because of that transient race
|
||||
+48
@@ -0,0 +1,48 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Metadata Discovery Uses High-Signal Queries
|
||||
Repository, package, and image metadata discovery SHALL prefer provider names, API hosts, framework names, and ecosystem terms over generic secret-related words.
|
||||
|
||||
#### Scenario: Repo metadata query list is reviewed
|
||||
- **WHEN** default repository or package metadata queries are configured
|
||||
- **THEN** broad terms such as `api_key`, `secret`, and `token` SHALL NOT be added by default unless the source searches content rather than metadata
|
||||
|
||||
#### Scenario: Provider query is configured
|
||||
- **WHEN** a provider such as Qwen, DashScope, Groq, or OpenRouter is targeted
|
||||
- **THEN** provider names, API hostnames, and framework terms SHALL be eligible for metadata discovery queries
|
||||
|
||||
### Requirement: Exact Secret Terms Reserved For Content-Oriented Sources
|
||||
Exact environment variable and API host searches SHALL be used for content-oriented sources such as Postman/API artifacts, code search, and CI artifacts rather than generic metadata searches.
|
||||
|
||||
#### Scenario: Env var search term is added
|
||||
- **WHEN** an exact term such as `DASHSCOPE_API_KEY` is added to discovery
|
||||
- **THEN** it SHALL be applied to a source that can inspect file or artifact content
|
||||
|
||||
### Requirement: Package Git Repository Canonicalization
|
||||
`package_git` discovery SHALL canonicalize repository URLs from package metadata before queueing targets.
|
||||
|
||||
#### Scenario: Package metadata contains issue URL
|
||||
- **WHEN** package metadata contains a repository-like URL ending in `/issues`
|
||||
- **THEN** `package_git` discovery SHALL normalize it to the canonical repository URL when possible
|
||||
|
||||
#### Scenario: Package metadata contains repository URL variants
|
||||
- **WHEN** package metadata contains `.git`, branch, tree, or homepage variants for the same repository
|
||||
- **THEN** `package_git` discovery SHALL avoid queueing duplicate normalized repository targets
|
||||
|
||||
### Requirement: CI Seed Parsing From Existing Data
|
||||
GitHub Actions and GitLab CI target discovery SHALL parse repository/project seeds from existing scanner DB records, package git candidates, findings, and target scan metadata where possible.
|
||||
|
||||
#### Scenario: Seed record contains package git target JSON
|
||||
- **WHEN** a CI source examines a package git candidate or target scan record with repository metadata
|
||||
- **THEN** it SHALL derive a GitHub repository or GitLab project seed when the URL provider matches the CI source
|
||||
|
||||
#### Scenario: Seed record cannot identify repository
|
||||
- **WHEN** no repository or project can be derived from a seed record
|
||||
- **THEN** the CI source SHALL count it as unparseable and continue processing other seeds
|
||||
|
||||
### Requirement: CI Scan Volume Is Configurable
|
||||
CI source scan volume SHALL remain controlled by existing per-source configuration values.
|
||||
|
||||
#### Scenario: CI seed limit is raised
|
||||
- **WHEN** `ci_seed_scan_limit` or `ci_max_repos_per_cycle` is increased in configuration
|
||||
- **THEN** the CI source SHALL use the configured value without requiring code changes
|
||||
@@ -0,0 +1,48 @@
|
||||
## 1. Runtime Stability
|
||||
|
||||
- [x] 1.1 Update `console_runner.save_state()` to write through a unique temp file per attempt.
|
||||
- [x] 1.2 Add bounded retry/backoff around `os.replace` for transient `PermissionError` and related Windows access errors.
|
||||
- [x] 1.3 Ensure failed state replacement after retries still surfaces the final error.
|
||||
- [x] 1.4 Add or run a focused smoke test that simulates state writes while the state file is repeatedly read.
|
||||
|
||||
## 2. Artifact Size Coverage
|
||||
|
||||
- [x] 2.1 Increase package artifact size limits in `config.yaml` while keeping npm/PyPI worker counts unchanged.
|
||||
- [x] 2.2 Increase Postman artifact size limits in `config.yaml` while keeping Postman worker counts unchanged.
|
||||
- [x] 2.3 Increase GitHub Actions and GitLab CI artifact archive/file limits in `config.yaml` while keeping CI worker counts unchanged.
|
||||
- [x] 2.4 Verify oversized artifacts still record existing skipped reasons in logs and target scan records.
|
||||
|
||||
## 3. Metadata Discovery Targeting
|
||||
|
||||
- [x] 3.1 Review default repository/package/image metadata query lists and keep generic `api_key`, `secret`, and `token` terms out of those defaults.
|
||||
- [x] 3.2 Add or retain provider/framework metadata queries for high-signal discovery terms such as Qwen, DashScope, Groq, OpenRouter, LiteLLM, LangChain, and LlamaIndex.
|
||||
- [x] 3.3 Ensure exact env var/API host terms are used only for content-oriented discovery paths such as Postman/API artifacts, code-like artifact search, or CI artifacts.
|
||||
|
||||
## 4. Package Git Discovery
|
||||
|
||||
- [x] 4.1 Extend package metadata extraction to inspect repository, homepage, bugs, and related package metadata fields for GitHub/GitLab repository URLs.
|
||||
- [x] 4.2 Canonicalize package-derived repository URLs by removing `.git`, issue paths, branch/tree paths, and other non-repository suffixes when possible.
|
||||
- [x] 4.3 Deduplicate package git targets by normalized repository URL before queueing.
|
||||
- [x] 4.4 Gradually increase `package_git.pages` and verify target volume, duplicate rate, and source runtime remain acceptable.
|
||||
|
||||
## 5. CI Seed Selection
|
||||
|
||||
- [x] 5.1 Improve GitHub Actions seed parsing from scanner DB target scans, findings, and package git candidate records.
|
||||
- [x] 5.2 Improve GitLab CI seed parsing from scanner DB target scans, findings, and package git candidate records.
|
||||
- [x] 5.3 Verify `skipped_unparseable` counts decrease for CI source discovery.
|
||||
- [x] 5.4 Increase `ci_seed_scan_limit` and `ci_max_repos_per_cycle` modestly after seed parsing improves.
|
||||
- [x] 5.5 Verify CI sources still respect configured worker and artifact limits.
|
||||
|
||||
## 6. Provider Validation Pattern
|
||||
|
||||
- [x] 6.1 Keep Qwen/DashScope context routing from sending strong Qwen-context `sk-...` keys to unrelated generic checkers.
|
||||
- [x] 6.2 Pick the next provider candidate with a safe non-generating validation endpoint.
|
||||
- [x] 6.3 Add the next provider using the detector plus keychecker pattern.
|
||||
- [x] 6.4 Ensure new provider keycheck results include stable detector names and metadata for DB link repair.
|
||||
|
||||
## 7. Verification
|
||||
|
||||
- [x] 7.1 Run Python compilation checks for modified Python modules.
|
||||
- [x] 7.2 Validate OpenSpec specs and task status for this change.
|
||||
- [x] 7.3 Run targeted smoke commands for state persistence, package discovery, and CI seed discovery.
|
||||
- [x] 7.4 Inspect supervisor status and recent logs after deployment to confirm source restarts and skipped/unparseable counts improved.
|
||||
Reference in New Issue
Block a user