Files
2026-09-30 20:30:56 +03:00

6.8 KiB

Context

GitHub, GitLab, and HuggingFace discovery return stable target identities together with remote update timestamps. The runner currently discards those timestamps and the PostgreSQL queue permanently deduplicates targets by (source, normalized_target). A completed repository or Space is therefore never scanned again even when its content changes.

Production discovery is already saturated: less than one percent of fetched records are new target identities and the core queue is often empty. The change must restore changed-content coverage without turning every rediscovery into a rescan, growing one queue row per revision, or creating an unbounded backlog.

Goals / Non-Goals

Goals:

  • Preserve a bounded remote content-update signal for GitHub, GitLab, and HuggingFace targets.
  • Rescan a completed target only when discovery observes content newer than the revision covered by its latest claim.
  • Coalesce multiple remote updates into one mutable queue row and one pending follow-up.
  • Bound changed-target admission and expose it separately from new-target admission.
  • Roll out without treating every legacy row as changed.

Non-Goals:

  • Requeue terminal failures or actively leased/deferred targets.
  • Rescan every known target on a timer.
  • Change DockerHub's digest-based identity and refresh behavior.
  • Guarantee that provider activity timestamps always represent content changes.
  • Enable inactive historical sources or enlarge global scan concurrency.

Decisions

Carry one normalized discovery record

The runner will represent eligible discovery results internally as a target plus an optional UTC remote_modified_at. GitHub uses pushed_at, GitLab uses last_activity_at, and HuggingFace uses lastModified. Missing, malformed, or non-monotonic timestamps remain valid discovery results but cannot trigger an updated-target rescan.

HuggingFace discovery will use a newest-modified feed so old Spaces changed recently are observable. Identity-only known-page stopping will be disabled when updated-target rescans are enabled; hard page and result limits remain the discovery bound.

Alternatives rejected:

  • Repository updated_at on GitHub, because metadata-only edits are not content pushes.
  • A HEAD-SHA request per repository, because it multiplies API traffic and rate-limit exposure.
  • Revision-aware early stopping, because a known first page does not prove later pages contain no changed targets.

Keep one queue row and two remote timestamps

target_queue will gain nullable remote_modified_at and scan_remote_modified_at columns. Discovery monotonically advances remote_modified_at. Claiming atomically copies the currently observed value into scan_remote_modified_at, recording what that scan covers.

The separate claim snapshot is required because a remote update can arrive while a scan is running. On completion, a newer observed timestamp remains ahead of the claimed timestamp and is eligible for exactly one later scan. Encoding revisions into normalized_target was rejected because it would grow queue rows and weaken queue authority.

For a legacy completed row with no claim snapshot, completed_at is the rollout baseline. It is eligible only when the first valid remote timestamp observed is newer than that completion. This prevents a migration surge while still admitting updates that occurred after the historical scan.

Observe and requeue atomically under a hard budget

A PostgreSQL transaction will upsert discovery observations and requeue at most updated_rescan_max_per_cycle eligible rows for one source. Eligibility requires:

  • status='done';
  • a strictly newer observed timestamp than scan_remote_modified_at, or than completed_at for a legacy row;
  • elapsed updated_rescan_cooldown_seconds since completion;
  • no lease, reservation, claim, or resolver authority.

Eligible rows are locked with FOR UPDATE SKIP LOCKED. Requeue resets only retry/completion scheduling fields required for a normal pending claim. Failed, pending, deferred, in-progress, unresolved, unchanged, and invalid-timestamp rows are never promoted by this path.

The initial production setting is one updated target per source cycle. Existing backlog-first behavior remains enabled, so a source drains its admitted work before discovery can admit more; refresh_registry is not enabled by this change.

Alternatives rejected:

  • Global requeue_done=True, because it requeues unchanged rows on every cycle.
  • Comparing only remote time with local completion time forever, because provider and host clocks differ and a mid-scan update can be lost.
  • A separate maintenance queue, because it duplicates existing lease, reservation, and completion authority.

Account for updated targets separately

source_cycles will gain queued_updated_count. queued_new_count keeps its current meaning. Source logs and dashboard aggregation will report changed-target admissions separately so rollout volume and yield can be audited.

Risks / Trade-offs

  • [GitLab activity can change without a repository push] -> Use a one-target-per-cycle cap and cooldown; report updated admissions separately.
  • [Provider clock skew] -> Require strict monotonicity and use the claimed remote timestamp after the first revision-aware scan.
  • [More discovery API traffic after disabling identity-only early stop] -> Keep existing hard page/per-page limits and source intervals.
  • [Changed targets consume capacity without useful findings] -> Start at one per cycle and compare updated-target yield before increasing the cap.
  • [Schema rollout while runtime is active] -> Stop the authority-managed runtime, apply schema through the normal initialization path, run tests, then restart through start_runtime.ps1.
  • [Rollback leaves nullable columns] -> Disable the feature in source configuration; nullable columns and metrics are backward-compatible and can remain.

Migration Plan

  1. Add nullable queue columns, the source-cycle counter, and a partial eligibility index through idempotent schema initialization.
  2. Deploy code and tests with updated-target rescans disabled by default.
  3. Enable GitHub, GitLab, and HuggingFace with a cap of one and a conservative cooldown; leave DockerHub unchanged.
  4. Restart the authority-managed runtime and verify discovery, queue authority, projection, and source health.
  5. Observe changed-target admission, completion, findings, and worker occupancy before changing any cap.

Rollback is configuration-first: disable updated-target rescans and restart the managed runtime. Existing pending work completes under normal queue semantics; no destructive data migration is required.

Open Questions

  • Whether production evidence supports different cooldowns per source after the initial canary.
  • Whether a later change should add provider-specific immutable revisions when APIs can supply them without extra requests.