Initial server source import
This commit is contained in:
@@ -0,0 +1,134 @@
|
||||
## Context
|
||||
|
||||
The bounded Docker-depth experiment showed that all five currently alive
|
||||
credentials were already present in the newest selected image. Image ranks
|
||||
2-10 consumed about 7.5 scanner-hours, added twelve globally new historical
|
||||
credentials, and added no currently alive credential. The next experiment
|
||||
therefore spends a larger bounded budget on repository breadth while scanning
|
||||
only the newest eligible immutable image from each selected repository.
|
||||
|
||||
The complete frozen 61-query discovery pass contains enough fresh provenance
|
||||
for this cohort. Most non-cohort repository anchors are currently held by the
|
||||
depth experiment and can become eligible only after that experiment completes
|
||||
and its exact reviewed cold events are reversed. PostgreSQL remains the sole
|
||||
authority for cohort, leases, reservations, policy events, capacity, and
|
||||
reporting evidence.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Select exactly 2,000 previously unscanned physical repositories from the
|
||||
existing complete frozen discovery pass without using yield.
|
||||
- Balance selection across the 52 keywords that have an eligible remaining
|
||||
pool: 38 repositories each, then one additional repository for the first 24
|
||||
still-eligible keywords in pinned query order.
|
||||
- Physically deduplicate repositories across keywords while retaining their
|
||||
frozen many-to-many keyword provenance.
|
||||
- Resolve and scan at most one newest eligible immutable image per selected
|
||||
repository under the existing experiment authority and fencing protocol.
|
||||
- Hand authority over only after the depth experiment is completed and its
|
||||
reviewed release has restored every owned non-cohort row.
|
||||
- Report globally deduplicated credential and currently-alive yield, scanner
|
||||
cost, and secret-safe per-keyword attribution.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not scan older image ranks or reinterpret `rank1` as a filesystem layer.
|
||||
- Do not infer an adaptive-depth trigger from the zero depth-only alive sample.
|
||||
- Do not reactivate scanned, failed, quarantined, independently cold, fenced,
|
||||
or incompletely released repository rows.
|
||||
- Do not run the depth and breadth experiments concurrently.
|
||||
- Do not expose repositories, image targets, credentials, hashes, or raw
|
||||
scanner evidence in operator output.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Versioned experiment profiles
|
||||
|
||||
The existing Docker experiment tables and worker protocol are reused. The
|
||||
reviewed selector version identifies one of two exact code-pinned profiles:
|
||||
the existing depth profile remains unchanged, while the breadth profile fixes
|
||||
61 queries, a per-query ceiling of 39, shallow/deep image depth of one, and a
|
||||
global physical target limit of 2,000. Configuration validation accepts only
|
||||
one complete reviewed profile; arbitrary mixtures remain invalid.
|
||||
|
||||
The breadth profile's theoretical capacity is the global physical limit, not
|
||||
`query_count * repositories_per_query`, because cross-keyword physical
|
||||
deduplication and the global stop are part of planning. Schema bounds are
|
||||
widened only enough to store the reviewed profile. Runtime authority still
|
||||
compares every persisted value with the exact supplied reviewed profile.
|
||||
|
||||
### Deterministic balanced physical selection
|
||||
|
||||
Planning reads the existing complete frozen deep discovery pass and orders each
|
||||
query's eligible repository observations by search rank and stable queue ID.
|
||||
It walks pinned queries round-robin, taking the next candidate whose physical
|
||||
queue ID has not already been selected, until each query owns at most 39
|
||||
repositories or the global count reaches exactly 2,000. Duplicate observations
|
||||
are skipped within the selecting query rather than consuming quota.
|
||||
|
||||
Given the reviewed frozen pool this produces 38 repositories for each of 52
|
||||
nonempty keywords and a 39th for the first 24 of those keywords in pinned order;
|
||||
the nine empty keywords remain explicit zero rows. Manifest generation fails
|
||||
closed unless it reaches exactly 2,000 unique physical repositories with this
|
||||
distribution. All fresh query observations remain in relational provenance and
|
||||
report attribution joins the selected physical repository back to that frozen
|
||||
evidence rather than duplicating scan work.
|
||||
|
||||
### Exact prior-release eligibility
|
||||
|
||||
A repository with no policy-event history remains eligible under the existing
|
||||
rules. A repository with history is eligible only when every event belongs to
|
||||
a completed/released Docker experiment and forms an exact cold/reactivate pair:
|
||||
the reverse event names the cold event, restores its recorded prior state, and
|
||||
matches experiment, config, policy, manifest, and audit evidence. Unreversed,
|
||||
unrelated, malformed, or mixed history is ineligible and causes no automatic
|
||||
repair. The old depth cohort is excluded independently through its persisted
|
||||
experiment membership.
|
||||
|
||||
This exception permits the newly reviewed experiment to hold and later release
|
||||
rows that were safely released by the prior experiment without weakening the
|
||||
append-only policy audit chain.
|
||||
|
||||
### Reviewed authority handoff
|
||||
|
||||
The current depth experiment must first become `completed`. Runtime is stopped
|
||||
canonically, the terminal aggregate report and release manifest are generated,
|
||||
and exact-SHA reviewed release restores only that experiment's owned cold rows.
|
||||
The breadth cohort and hold manifests are then generated and applied while
|
||||
sources remain stopped. Activation rejects any other unreleased, nonterminal,
|
||||
or fenced Docker experiment authority.
|
||||
|
||||
The breadth experiment uses the existing state machine, experiment-row-first
|
||||
locking, resolver tokens, target reservations, finite retries, capacity
|
||||
accounting, and reversible holds. Since both shallow and deep image limits are
|
||||
one, every image-bearing membership produces only selection rank one and a
|
||||
single dispatch wave.
|
||||
|
||||
### Secret-safe decision report
|
||||
|
||||
The terminal report treats target-scoped finding fingerprints as location
|
||||
evidence, not secret novelty. Primary outcomes are globally deduplicated
|
||||
credentials, currently alive credentials, and each per scanner-hour. It also
|
||||
reports physical repository/image coverage, scan states, keycheck completeness,
|
||||
and per-keyword attribution from frozen provenance. No raw identity or target
|
||||
material is emitted.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [A selected repository has no eligible image] -> Use the existing deterministic
|
||||
fresh replacement path; fail the reviewed cohort if exact 2,000 repository
|
||||
ownership cannot be preserved before activation.
|
||||
- [Cross-keyword overlap biases ownership] -> Use pinned query-order round-robin,
|
||||
preserve all frozen query provenance, and distinguish physical totals from
|
||||
keyword attribution credits.
|
||||
- [Prior policy history is malformed] -> Leave the row ineligible and fail
|
||||
closed rather than guessing or rewriting history.
|
||||
- [Two experiment authorities overlap] -> Reject planning/activation until the
|
||||
prior experiment is completed, reviewed, released, and fence-free.
|
||||
- [The cohort is large] -> Keep the hard 2,000-target ceiling, existing shared
|
||||
capacity limits, two Docker workers, and finite retry/hold behavior.
|
||||
- [Migration or rollout fails] -> Keep schema changes additive where possible,
|
||||
stop before authority mutation, and retain manifests and audit rows for a
|
||||
deterministic retry.
|
||||
Reference in New Issue
Block a user