76 lines
6.3 KiB
Markdown
76 lines
6.3 KiB
Markdown
## Context
|
|
|
|
The configured sources contain 543 query occurrences covering 133 normalized terms. Canonical PostgreSQL lineage links discovery scans to credentials through both candidate and result records, and the dashboard already defines the strict `usable_llm` tier. Applying that rule across every provider identified 38 non-operational terms with zero historical strict-usable linkage despite 117,492 scans and 1,890.6 cumulative scanner-hours. The same terms consumed 18,641 scans and 317.1 scanner-hours in the latest 30-day window.
|
|
|
|
The runtime configuration is code-authority protected. Query rotation state, target queues, scan history, and credential history are separate persisted authorities and must not be rewritten to deploy this change.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals:**
|
|
- Remove only globally zero-yield terms with enough exposure to support a conservative decision.
|
|
- Apply one decision consistently anywhere the exact normalized term is configured.
|
|
- Preserve operational source sentinels and all terms with any demonstrated strict-usable linkage.
|
|
- Remove stale per-query overrides and verify deterministic post-prune query sets.
|
|
- Deploy through the coordinated authority lifecycle and verify normal source rotation.
|
|
|
|
**Non-Goals:**
|
|
- Deleting, reprioritizing, or rewriting existing target backlog or historical records.
|
|
- Optimizing for one provider, broad `alive`, raw findings, or candidate volume.
|
|
- Changing detector routing, keycheck classification, source concurrency, scan limits, or query-state files.
|
|
- Claiming that a retired term can never produce a useful credential in the future.
|
|
|
|
## Decisions
|
|
|
|
### Use all-provider strict-usable evidence
|
|
|
|
A term is eligible only when no credential linked to that term has ever reached the canonical dashboard `usable_llm` tier and no earliest-origin credential attributed to it has reached that tier. Candidate/result scan links are unioned before attribution so migrated and resolver-routed credentials are not lost.
|
|
|
|
The decision is global by case-normalized exact term. If a term produced one strict-usable credential for any provider or source, it remains configured everywhere. This is more conservative than pruning source-term pairs independently and avoids removing cross-provider terms such as `groq`, `llm`, `chat`, `rag`, or `langchain`.
|
|
|
|
Alternative: use broad `status_group=alive` or OpenAI-only yield. Rejected because broad alive contains unproven and historically misclassified statuses, while provider-only analysis can remove terms that work for another provider.
|
|
|
|
### Require meaningful exposure
|
|
|
|
A zero-yield term qualifies when either it has at least 30 linked credential observations, or it has at least 200 completed scan events and 20 cumulative scanner-hours. The credential branch tests precision; the cost branch catches terms that repeatedly consume work without reaching candidate intake. The threshold is applied to all retained history, with the latest 30-day cost recorded as corroborating evidence.
|
|
|
|
Alternative: remove every zero-yield term. Rejected because recent and low-sample terms have insufficient evidence. Those terms remain canaries.
|
|
|
|
### Exempt source-operational sentinels
|
|
|
|
`gharchive`, `gharchive-files`, and `gists` are sole query tokens used to operate dedicated sources rather than interchangeable discovery keywords. They remain even though they have no strict-usable attribution. Emptying those lists would disable or invalidate source operation rather than merely prune a search term.
|
|
|
|
### Retire the approved cohort consistently
|
|
|
|
Remove these 32 terms from GitHub, GitLab, DockerHub, npm, PyPI, and package-git: `autonomous`, `benchmarks`, `claw`, `code-assistant`, `codegen`, `dspy`, `embedding`, `embeddings`, `eval`, `evals`, `gateway`, `grok`, `haystack`, `inference`, `inference-api`, `knowledge`, `llamaindex`, `model`, `model-router`, `models`, `ollama`, `orchestration`, `prompts`, `replicate`, `rerank`, `reranker`, `retrieval`, `router`, `tokenizer`, `tool-use`, `vector`, and `vllm`.
|
|
|
|
Remove `chatgpt`, `gpt`, and `moonshot` from those six sources and Postman. Remove `openai` from GitHub, GitLab, DockerHub, and Postman. Remove `dashscope-intl.aliyuncs.com` and `generativelanguage.googleapis.com` from Postman.
|
|
|
|
This removes 219 occurrences. Resulting list sizes are GitHub 64, GitLab 46, DockerHub 46, npm 43, PyPI 43, package-git 43, and Postman 33.
|
|
|
|
### Keep deployment configuration-only
|
|
|
|
Delete the three `openai` query overrides together with the query entries. Existing rotation reads `query_index` modulo the current list length, so no persisted state edit is needed. Runtime is stopped before editing and restarted only after tests and strict OpenSpec validation.
|
|
|
|
Alternative: rewrite query indices or purge queued targets attributed to removed terms. Rejected because both mutate independent durable authority and are unnecessary for preventing future discovery.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
- [Historical zero yield may not predict future supply] -> Keep low-sample terms, retain all historical evidence, and make rollback a configuration-only restoration.
|
|
- [Earliest-origin attribution can hide useful rediscovery] -> Require zero strict-usable linkage across every scan link in addition to zero origin yield.
|
|
- [Large list reduction changes rotation cadence] -> Verify exact list sizes and allow normal modulo-based state handling; do not edit source state.
|
|
- [Completed OpenAI rollout previously required the literal term] -> Record the requirement retirement explicitly and retain higher-signal bounded ecosystem queries.
|
|
- [Authority drift during a live edit] -> Use coordinated stop, test, and canonical start rather than relying on fail-close shutdown.
|
|
|
|
## Migration Plan
|
|
|
|
1. Add configuration contract tests for the exact retired set, retained sentinels, uniqueness, post-prune sizes, and absence of orphaned overrides.
|
|
2. Stop the authenticated runtime coordinately.
|
|
3. Remove the 219 query occurrences and three matching overrides from `app/config.yaml`; do not edit state or queue data.
|
|
4. Run focused query tests, configuration/runtime safety tests as applicable, and strict OpenSpec validation.
|
|
5. Restart through `start_runtime.ps1` and verify authenticated supervisor, PostgreSQL, pipeline readiness, source processes, and query-list loading.
|
|
6. Roll back by restoring the configuration entries and overrides through the same coordinated lifecycle if source health regresses.
|
|
|
|
## Open Questions
|
|
|
|
None.
|