8.8 KiB
ADDED Requirements
Requirement: Postman source registration
The system SHALL provide a first-class postman source that can be configured, selected, supervised, queued, scanned, and displayed consistently with existing scanner sources.
Scenario: Configured Postman source is selectable
- WHEN a config contains
sources.postman.enabled: trueand the runner is invoked with--source postman - THEN the runner SHALL execute only the Postman source cycle using Postman source settings
Scenario: Postman queue files are used
- WHEN the Postman source prepares targets
- THEN the system SHALL use
todo_postman.txtandchecked_postman.txtunder the configured queue directory
Requirement: GitHub code search discovery
The Postman source SHALL discover public Postman artifacts from GitHub code search using configured queries and artifact kinds.
Scenario: Collection files are discovered
- WHEN
search_kindsincludescollectionand the query isopenai - THEN discovery SHALL search GitHub code for
filename:postman_collection.json openai
Scenario: Environment files are discovered
- WHEN
search_kindsincludesenvironmentand the query isopenai - THEN discovery SHALL search GitHub code for
filename:postman_environment.json openai
Scenario: Search pagination is bounded
- WHEN
pagesis10andper_pageis100 - THEN discovery SHALL request no more than 1000 search results per query and kind
Requirement: GitHub auth pool rotation
The Postman GitHub discovery flow SHALL use all available tokens from the configured GitHub auth pool before sleeping for rate limits.
Scenario: Token rotates per GitHub request
- WHEN multiple GitHub auth entries are available
- THEN GitHub code search, commit lookup, and content download requests SHALL rotate across available tokens
Scenario: One token is rate-limited
- WHEN a GitHub request returns a primary or secondary rate limit for the current token
- THEN the system SHALL mark only that token unavailable until its reset time or configured cooldown and continue with another available token
Scenario: All tokens are unavailable
- WHEN every configured GitHub token is rate-limited or temporarily unavailable
- THEN the source SHALL sleep until the earliest known reset time, or for the configured fallback cooldown when no reset time is known
Scenario: Token is invalid
- WHEN a GitHub request returns an authentication-invalid response for a token
- THEN the system SHALL exclude that token from the current cycle and report the authentication failure without marking other tokens invalid
Requirement: Backfill freshness filtering
The Postman source SHALL support a backfill mode that can scan deep code search pages while skipping GitHub artifacts older than a configured file age.
Scenario: Recent file is queued
- WHEN a GitHub code search result has a latest path commit within
max_file_age_days - THEN the target SHALL be eligible for queueing
Scenario: Old file is skipped
- WHEN a GitHub code search result has a latest path commit older than
max_file_age_days - THEN the target SHALL not be queued and SHALL be counted as skipped by freshness filtering
Scenario: Freshness filtering is disabled
- WHEN
max_file_age_daysis0 - THEN the source SHALL not perform path commit age filtering
Requirement: Tail mode early stop
The Postman source SHALL support ongoing tail scans that stop pagination after consecutive known pages.
Scenario: Known page increments stop counter
- WHEN
stop_on_seen_pagesis enabled and every normalized target on a fetched page already exists intodo_postman.txtorchecked_postman.txt - THEN the source SHALL count that page as known
Scenario: Tail pagination stops
- WHEN the known page count reaches
seen_page_thresholdaftermin_pages_before_stop - THEN the source SHALL stop fetching additional pages for that query and artifact kind
Requirement: Postman target identity and deduplication
The system SHALL normalize Postman targets so identical artifacts are not rescanned while changed artifacts are scanned again.
Scenario: GitHub target identity includes SHA
- WHEN a Postman target is discovered from GitHub code search
- THEN its normalized target SHALL include source, repository, path, and file SHA
Scenario: GitHub file changes
- WHEN the same GitHub repository and path is discovered with a new SHA
- THEN the system SHALL treat it as a new Postman target
Scenario: Package target identity uses content hash
- WHEN a Postman artifact is harvested from npm or PyPI
- THEN its normalized target SHALL include the artifact content SHA-256 hash
Requirement: Durable Postman artifact cache
The system SHALL store discovered Postman artifact content in durable runtime cache before scanning.
Scenario: GitHub content is cached
- WHEN a GitHub code search target is queued for scanning
- THEN the system SHALL download the artifact content and store it under the configured Postman cache directory
Scenario: Package content is cached before cleanup
- WHEN npm or PyPI extraction finds a Postman artifact
- THEN the system SHALL copy the artifact into the durable Postman cache before the extraction directory is removed
Scenario: Cache size is constrained
- WHEN an artifact exceeds the configured maximum Postman artifact size
- THEN the system SHALL skip the artifact and record a bounded error or skip reason
Requirement: Postman artifact scanning
The Postman source SHALL scan cached Postman collection and environment artifacts with TruffleHog filesystem scanning.
Scenario: Cached artifact is scanned
- WHEN a Postman target points to a cached JSON artifact
- THEN the scanner SHALL run TruffleHog against a temporary filesystem directory containing that artifact
Scenario: Findings are persisted
- WHEN TruffleHog reports findings for a Postman target
- THEN the system SHALL persist findings to existing JSONL outputs and scanner database tables with source
postman
Scenario: Scan finishes
- WHEN a Postman target scan completes with findings, errors, skipped status, or clean status
- THEN the target SHALL be moved from
todo_postman.txttochecked_postman.txt
Requirement: npm and PyPI Postman harvesting
The npm and PyPI source flows SHALL harvest Postman artifacts discovered during existing package extraction and enqueue them for the Postman source.
Scenario: npm package contains collection
- WHEN an extracted npm package contains a file matching
*.postman_collection.json - THEN the system SHALL cache the file and enqueue a Postman target with npm package origin metadata
Scenario: PyPI package contains environment
- WHEN an extracted PyPI artifact contains a file matching
*.postman_environment.json - THEN the system SHALL cache the file and enqueue a Postman target with PyPI package origin metadata
Scenario: Existing package scan continues
- WHEN Postman harvesting fails for one package artifact
- THEN the original npm or PyPI scan SHALL still complete and record the harvesting failure without failing unrelated package scanning
Requirement: Postman-aware enrichment
The system SHALL enrich Postman findings with contextual classification derived from Postman structure without replacing TruffleHog detection.
Scenario: Header credential is classified
- WHEN a finding appears in a Postman request header such as
Authorizationorx-api-key - THEN enrichment SHALL record the context location and infer credential kind from header type, value shape, and endpoint host when possible
Scenario: Environment variable is classified
- WHEN a finding appears in a Postman environment variable value
- THEN enrichment SHALL record the variable name and classify provider or credential kind when supported by value shape or associated request endpoints
Scenario: Placeholder is detected
- WHEN a Postman value is a placeholder such as
{{API_KEY}},<api_key>,YOUR_API_KEY,example, orchangeme - THEN enrichment SHALL classify it as placeholder or low confidence rather than a live secret
Requirement: Observability for Postman source
The system SHALL expose Postman source activity through existing logs, queue counts, source cycle metrics, target scan records, findings, errors, and dashboard views.
Scenario: Source cycle is recorded
- WHEN a Postman source cycle runs
- THEN the scanner database SHALL record source cycle metrics including fetched, queued, scanned, found, error, skipped, and queue counts
Scenario: Dashboard shows Postman queues
- WHEN Postman queue files exist
- THEN the dashboard SHALL include Postman queue counts in the current queues view