Initial server source import

This commit is contained in:
sashatrask
2026-09-30 20:30:56 +03:00
commit 170dd941b9
498 changed files with 261563 additions and 0 deletions
@@ -0,0 +1,134 @@
## Context
The bounded Docker-depth experiment showed that all five currently alive
credentials were already present in the newest selected image. Image ranks
2-10 consumed about 7.5 scanner-hours, added twelve globally new historical
credentials, and added no currently alive credential. The next experiment
therefore spends a larger bounded budget on repository breadth while scanning
only the newest eligible immutable image from each selected repository.
The complete frozen 61-query discovery pass contains enough fresh provenance
for this cohort. Most non-cohort repository anchors are currently held by the
depth experiment and can become eligible only after that experiment completes
and its exact reviewed cold events are reversed. PostgreSQL remains the sole
authority for cohort, leases, reservations, policy events, capacity, and
reporting evidence.
## Goals / Non-Goals
**Goals:**
- Select exactly 2,000 previously unscanned physical repositories from the
existing complete frozen discovery pass without using yield.
- Balance selection across the 52 keywords that have an eligible remaining
pool: 38 repositories each, then one additional repository for the first 24
still-eligible keywords in pinned query order.
- Physically deduplicate repositories across keywords while retaining their
frozen many-to-many keyword provenance.
- Resolve and scan at most one newest eligible immutable image per selected
repository under the existing experiment authority and fencing protocol.
- Hand authority over only after the depth experiment is completed and its
reviewed release has restored every owned non-cohort row.
- Report globally deduplicated credential and currently-alive yield, scanner
cost, and secret-safe per-keyword attribution.
**Non-Goals:**
- Do not scan older image ranks or reinterpret `rank1` as a filesystem layer.
- Do not infer an adaptive-depth trigger from the zero depth-only alive sample.
- Do not reactivate scanned, failed, quarantined, independently cold, fenced,
or incompletely released repository rows.
- Do not run the depth and breadth experiments concurrently.
- Do not expose repositories, image targets, credentials, hashes, or raw
scanner evidence in operator output.
## Decisions
### Versioned experiment profiles
The existing Docker experiment tables and worker protocol are reused. The
reviewed selector version identifies one of two exact code-pinned profiles:
the existing depth profile remains unchanged, while the breadth profile fixes
61 queries, a per-query ceiling of 39, shallow/deep image depth of one, and a
global physical target limit of 2,000. Configuration validation accepts only
one complete reviewed profile; arbitrary mixtures remain invalid.
The breadth profile's theoretical capacity is the global physical limit, not
`query_count * repositories_per_query`, because cross-keyword physical
deduplication and the global stop are part of planning. Schema bounds are
widened only enough to store the reviewed profile. Runtime authority still
compares every persisted value with the exact supplied reviewed profile.
### Deterministic balanced physical selection
Planning reads the existing complete frozen deep discovery pass and orders each
query's eligible repository observations by search rank and stable queue ID.
It walks pinned queries round-robin, taking the next candidate whose physical
queue ID has not already been selected, until each query owns at most 39
repositories or the global count reaches exactly 2,000. Duplicate observations
are skipped within the selecting query rather than consuming quota.
Given the reviewed frozen pool this produces 38 repositories for each of 52
nonempty keywords and a 39th for the first 24 of those keywords in pinned order;
the nine empty keywords remain explicit zero rows. Manifest generation fails
closed unless it reaches exactly 2,000 unique physical repositories with this
distribution. All fresh query observations remain in relational provenance and
report attribution joins the selected physical repository back to that frozen
evidence rather than duplicating scan work.
### Exact prior-release eligibility
A repository with no policy-event history remains eligible under the existing
rules. A repository with history is eligible only when every event belongs to
a completed/released Docker experiment and forms an exact cold/reactivate pair:
the reverse event names the cold event, restores its recorded prior state, and
matches experiment, config, policy, manifest, and audit evidence. Unreversed,
unrelated, malformed, or mixed history is ineligible and causes no automatic
repair. The old depth cohort is excluded independently through its persisted
experiment membership.
This exception permits the newly reviewed experiment to hold and later release
rows that were safely released by the prior experiment without weakening the
append-only policy audit chain.
### Reviewed authority handoff
The current depth experiment must first become `completed`. Runtime is stopped
canonically, the terminal aggregate report and release manifest are generated,
and exact-SHA reviewed release restores only that experiment's owned cold rows.
The breadth cohort and hold manifests are then generated and applied while
sources remain stopped. Activation rejects any other unreleased, nonterminal,
or fenced Docker experiment authority.
The breadth experiment uses the existing state machine, experiment-row-first
locking, resolver tokens, target reservations, finite retries, capacity
accounting, and reversible holds. Since both shallow and deep image limits are
one, every image-bearing membership produces only selection rank one and a
single dispatch wave.
### Secret-safe decision report
The terminal report treats target-scoped finding fingerprints as location
evidence, not secret novelty. Primary outcomes are globally deduplicated
credentials, currently alive credentials, and each per scanner-hour. It also
reports physical repository/image coverage, scan states, keycheck completeness,
and per-keyword attribution from frozen provenance. No raw identity or target
material is emitted.
## Risks / Trade-offs
- [A selected repository has no eligible image] -> Use the existing deterministic
fresh replacement path; fail the reviewed cohort if exact 2,000 repository
ownership cannot be preserved before activation.
- [Cross-keyword overlap biases ownership] -> Use pinned query-order round-robin,
preserve all frozen query provenance, and distinguish physical totals from
keyword attribution credits.
- [Prior policy history is malformed] -> Leave the row ineligible and fail
closed rather than guessing or rewriting history.
- [Two experiment authorities overlap] -> Reject planning/activation until the
prior experiment is completed, reviewed, released, and fence-free.
- [The cohort is large] -> Keep the hard 2,000-target ceiling, existing shared
capacity limits, two Docker workers, and finite retry/hold behavior.
- [Migration or rollout fails] -> Keep schema changes additive where possible,
stop before authority mutation, and retain manifests and audit rows for a
deterministic retry.