Initial server source import
This commit is contained in:
@@ -0,0 +1,78 @@
|
||||
## Context
|
||||
|
||||
The exact `openai` rollout proved that bounded source queries can restore unseen credential supply without changing scanner or keycheck authority: 18 DockerHub targets produced 24 genuinely new OpenAI credentials, all with explicit terminal API outcomes. None were usable, so increasing that broad query is not justified. Historical production attribution instead shows usable OpenAI outcomes behind narrower agent, chatbot, and conversation ecosystems, while source semantics differ enough that one shared keyword list is inefficient.
|
||||
|
||||
The existing exact-query override mechanism already validates and applies `pages`, `per_page`, and `max_targets`. This change can therefore remain configuration-only plus focused contract tests. Runtime configuration is immutable-authority covered, so deployment must use coordinated stop/start and must not modify persisted query state directly.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Test nine high-signal OpenAI integration and deployable-application queries across the sources where their search semantics fit.
|
||||
- Bound discovery admission as well as scan claims for every new query.
|
||||
- Run each new query once promptly after deployment without directly editing query state.
|
||||
- Attribute the canary through durable scans, candidates, provider results, and projections.
|
||||
- Keep or remove each query based on its own measured useful yield and operational cost.
|
||||
|
||||
**Non-Goals:**
|
||||
- Expanding broad generic terms such as `gpt`, `llm`, `chatgpt`, `ai`, or `model`.
|
||||
- Increasing global scan concurrency, source worker counts, or updated-target promotion limits.
|
||||
- Changing detector routing, OpenAI checker classification, known-credential caching, or projection semantics.
|
||||
- Re-enabling inactive broad sources or guaranteeing a valid funded credential.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Use source-specific first-wave queries
|
||||
|
||||
The first wave is:
|
||||
- GitHub: `OPENAI_API_KEY`, `api.openai.com`, `openai-agents`.
|
||||
- GitLab: `openai-api`, `openai-agents`, `librechat`.
|
||||
- DockerHub: `librechat`, `lobechat`, `openai-proxy`.
|
||||
|
||||
GitHub can search README content, so direct environment and endpoint signatures are appropriate. GitLab project search is metadata-oriented, so branded slug terms are used. DockerHub searches repository metadata and then resolves immutable image digests, so deployable project and proxy names are used.
|
||||
|
||||
Alternative: add the same list to all sources. Rejected because it lengthens every rotation and ignores source search semantics. Alternative: expand historically broad terms. Rejected because those cohorts produced volume without usable OpenAI outcomes.
|
||||
|
||||
### Bound every query independently
|
||||
|
||||
Bounds are:
|
||||
- GitHub `OPENAI_API_KEY`: `pages=1`, `per_page=25`, `max_targets=5`.
|
||||
- GitHub `api.openai.com`: `1/25/5`.
|
||||
- GitHub `openai-agents`: `1/50/5`.
|
||||
- GitLab `openai-api`: `1/50/5`.
|
||||
- GitLab `openai-agents`: `1/50/5`.
|
||||
- GitLab `librechat`: `1/25/5`.
|
||||
- DockerHub `librechat`, `lobechat`, and `openai-proxy`: each `2/10/10`.
|
||||
|
||||
Page and page-size limits bound fetched/admitted identities; `max_targets` separately bounds claims in the active cycle. The existing global three-slot limit, GitLab one-updated-target-per-cycle cap, 24-hour update cooldown, and Docker digest requirement remain unchanged.
|
||||
|
||||
### Place the wave at each stopped source's current rotation index
|
||||
|
||||
After a coordinated runtime stop, capture each source's persisted `query_index` and insert that source's three-query block at the same index. Restarting then exercises the block naturally. Successful discovery advances through the block; failures and backlog-only work retain the current query under existing semantics. State files are never edited.
|
||||
|
||||
Alternative: append and wait for a full rotation. Rejected because DockerHub cycles can be long and attribution would be delayed. Alternative: edit persisted state. Rejected because state is runtime authority and direct edits would weaken recovery evidence.
|
||||
|
||||
### Evaluate individual query funnels
|
||||
|
||||
The canary reports fetched, new/updated admissions, durable scan outcomes, findings, genuinely new OpenAI credentials, `api_check` versus `cached_status`, explicit provider outcomes, first-alive/usable counts, projection drain, and runtime health. Aggregate volume alone cannot justify retention.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [README signatures produce placeholders] -> Count genuinely new credentials and explicit API outcomes; remove queries with only placeholder/dead yield.
|
||||
- [A query admits more work than its claim cap] -> Keep `pages × per_page` small and inspect the exact query-attributed queue until terminal.
|
||||
- [Docker scans are expensive or inaccessible] -> Cap each query at 20 repositories and 10 claims, retain digest authority, and classify registry failures separately.
|
||||
- [Three added terms lengthen source rotations] -> Keep only terms that add distinct credential or usable yield after the canary.
|
||||
- [Config deployment triggers authority fail-close] -> Stop coordinately before editing and restart only through `start_runtime.ps1`.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add focused tests for exact source membership, uniqueness, and all nine bounds.
|
||||
2. Coordinately stop the runtime and capture persisted source indices.
|
||||
3. Insert each three-query block at its source's current index; do not edit state files.
|
||||
4. Run focused and full regression suites plus strict OpenSpec validation with bytecode writes disabled.
|
||||
5. Start through the authoritative runtime script and verify PostgreSQL, pipeline, core sources, keychecks, and recorder.
|
||||
6. Observe each exact query cohort through terminal scans and keycheck projection, then retain or remove each query based on measured evidence.
|
||||
7. Roll back any low-value query by removing that query and override during a coordinated stop/start. No schema or data rollback is required.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None. A second wave remains gated on this canary's per-query useful-yield evidence.
|
||||
Reference in New Issue
Block a user