Files
truf-server/openspec/changes/stabilize-gitlab-discovery-retries/design.md
T
2026-09-30 20:30:56 +03:00

56 lines
3.3 KiB
Markdown

## Context
GitLab project discovery calls `api_request` without an explicit retry budget. Direct requests therefore receive one attempt, and `ApiRequestError` escapes `run_configured_source`, which terminates the supervised child. The supervisor restarts the child and the persisted query position prevents confirmed target loss, but every transient 30-second read timeout creates avoidable churn and delay.
## Goals / Non-Goals
**Goals:**
- Retry idempotent GitLab project discovery requests with a small bounded budget.
- Keep the GitLab child alive when the bounded network budget is exhausted.
- Persist a failed source cycle without advancing the query or damaging auth state.
- Preserve existing rate-limit and authentication handling.
**Non-Goals:**
- Change TruffleHog scan commands or target retry policy.
- Retry non-idempotent requests.
- Hide persistent GitLab outages or loop without delay.
- Change other source APIs in this change.
## Decisions
### Configure attempts at the GitLab source boundary
GitLab source configuration will provide `discovery_request_attempts: 3` and `discovery_retry_delay: 5`. These values flow only into project discovery calls and use the existing `api_request` retry implementation.
Alternative considered: change the direct-request default globally. Rejected because it would silently alter every source and API call without source-specific evidence.
### Convert exhausted discovery transport errors into failed cycles
GitLab project discovery will wrap exhausted request transport failures in a dedicated `GitLabDiscoveryTransportError`. `run_configured_source` will catch only that error for GitLab, roll back any open database transaction, finish the source cycle as failed, retain the current query, and return to the normal configured-source cooldown. It will not mark the token invalid or terminate the process. Payload and programming errors retain fail-fast behavior.
Alternative considered: rely on supervisor restart. Rejected because process restart is expensive error handling for an ordinary transient network condition.
### Preserve rate-limit and auth paths
HTTP 429, 401, and 403 responses continue through the existing GitLab API and auth-pool policy. Only transport failures and the existing retryable HTTP status set use the discovery retry budget.
## Risks / Trade-offs
- [A persistent outage lengthens a failed cycle] -> Bound attempts to three and delay to five seconds.
- [A failed page could cause partial discovery ambiguity] -> Discard the cycle's fetched list on failure and retain the current query for the next cycle.
- [A catch could hide programming errors] -> Catch only `ApiRequestError` for the GitLab source; all other exceptions retain fail-fast behavior.
- [Auth state could be corrupted] -> Do not invoke token cooldown or invalidation for transport-only failures.
## Migration Plan
1. Add retry and failed-cycle tests using injected timeout responses.
2. Deploy the GitLab-only configuration and error handling.
3. Restart the managed runtime and observe at least one full query rotation or an injected exhaustion test.
4. Roll back by removing the GitLab source settings and narrow exception handler.
## Open Questions
- Whether the same direct-request policy should later be adopted by other sources requires separate evidence.