Files
2026-09-30 20:30:56 +03:00

6.3 KiB

Context

The scanner is already organized around independent sources managed by console_runner.py and supervisor.py. Each source discovers targets, writes them to per-source queue files, scans targets through TruffleHog, records results in JSONL and SQLite, and can stop pagination when consecutive pages contain only known targets. Existing package sources already download and extract npm/PyPI artifacts before running TruffleHog filesystem scans.

Postman artifacts fit this model as filesystem scan targets, but they need separate discovery and enrichment. Public Postman web search is comparatively fragile, while GitHub code search exposes many real *.postman_collection.json and *.postman_environment.json files. npm and PyPI can also expose Postman artifacts during the package extraction windows that already exist.

Goals / Non-Goals

Goals:

  • Add a first-class postman source with standard queue, checked, supervisor, dashboard, and database behavior.
  • Discover Postman collection/environment JSON from GitHub code search using the existing GitHub auth pool.
  • Support initial backfill over up to the GitHub Search API result cap per query and ongoing tail scans over recently indexed pages.
  • Filter discovered GitHub code artifacts by last file commit age so stale files can be skipped during backfill.
  • Rotate across multiple GitHub tokens and pause only when all usable tokens are rate-limited.
  • Cache discovered Postman artifacts durably so package-derived artifacts survive temp directory cleanup.
  • Harvest Postman artifacts from npm and PyPI extraction flows without disrupting existing package scans.
  • Enrich findings with Postman-specific context derived from auth configuration, headers, query params, request bodies, environment variables, and endpoint hosts.

Non-Goals:

  • Scraping Postman's own web application in the first implementation.
  • Replacing TruffleHog detectors with custom regex-only detection.
  • Adding new keychecker services as part of this change.
  • Scanning private Postman workspaces through the Postman API.

Decisions

  1. Use GitHub code search as the primary discovery channel.

GitHub code search has authenticated API support, predictable pagination, and high-quality results for filename:postman_collection.json <query> and filename:postman_environment.json <query>. Public Postman web pages and postman.com links are lower-yield and more likely to change without notice. Direct Postman URL discovery can be added later as an additional provider without changing the scanner contract.

  1. Treat Postman artifacts as durable filesystem scan targets.

The source will download or copy each discovered artifact into a runtime Postman cache and scan a temporary directory containing the cached JSON. This reuses the existing TruffleHog filesystem path and avoids keeping npm/PyPI extraction directories alive.

  1. Identify GitHub-discovered targets by repo:path:sha.

The same file at the same SHA must not be rescanned, while a new SHA for the same path must be queued again. This matches the existing queue/checked model and makes stop_on_seen_pages useful for tail scans.

  1. Identify package-harvested targets by content hash.

npm/PyPI packages can contain duplicate Postman artifacts across versions or package names. A SHA-256 content hash provides stable dedupe and allows different origins to point to the same cached artifact without rescanning identical content.

  1. Filter GitHub code artifacts by path commit age after discovery.

The code search response does not include reliable file modification dates. The implementation will query the latest commit for each repo:path and skip artifacts older than max_file_age_days. This costs extra core API requests but keeps the code search query simple and reliable.

  1. Use a source-local GitHub token pool for discovery.

Postman discovery may make many GitHub requests in one cycle. A per-request token pool can rotate across all configured GitHub auth entries, cool down only the token that failed, and sleep when no token remains available. This is more efficient than one token per source cycle.

  1. Keep Postman enrichment separate from detection.

TruffleHog remains responsible for finding candidate secrets. Postman enrichment will add context and confidence by correlating findings with request auth, headers, variables, endpoints, and placeholder detection. This avoids increasing false positives from regex-only scans.

Risks / Trade-offs

  • GitHub code search is rate-limited to roughly 10 requests per minute per token -> throttle to a configurable safe RPM and rotate across the auth pool.
  • Commit-age filtering adds extra core API requests -> make max_file_age_days configurable and cache commit metadata per repo:path:sha within a cycle.
  • Broad queries such as ai can hit the 1000-result search cap and include noisy results -> use a Postman-specific query list and allow per-source query tuning.
  • Package harvesting adds small overhead during npm/PyPI scans -> limit file walking to reasonable extensions, max file size, and known Postman filename patterns.
  • Postman variables often contain placeholders rather than live secrets -> classify placeholders separately and keep TruffleHog verification/keycheckers as the authority for live/dead status.
  • All tokens may become unavailable -> sleep until the earliest known reset time, or a configured fallback such as 30 minutes when no reset is known.

Migration Plan

  1. Add the postman source disabled by default in config.yaml.
  2. Add queue/dashboard/database support for postman without changing existing source behavior.
  3. Run a small verification cycle with pages: 1, per_page: 10, and max_targets set.
  4. Run the one-time backfill with pages: 10, per_page: 100, stop_on_seen_pages: false, and max_file_age_days: 365.
  5. Switch the source to tail mode with pages: 1-3 and stop_on_seen_pages: true.
  6. Enable npm/PyPI harvesting after the base Postman source is verified.

Rollback is to disable sources.postman.enabled, leave its queues/cache intact, and continue running existing sources unchanged.

Open Questions

  • Whether the default backfill query list should include broad terms like ai, or keep only higher-intent terms such as openai, anthropic, gemini, llm, rag, and agent.
  • Whether package-harvested artifacts should always be enqueued, or only when they include auth/secret-related markers.