352 lines
7.8 KiB
Markdown
352 lines
7.8 KiB
Markdown
# Detector Notes
|
|
|
|
Working notes about TruffleHog detector behavior and local post-processing ideas.
|
|
|
|
## GitHub / GitLab Noise
|
|
|
|
Current TruffleHog source contains both modern and legacy detectors.
|
|
|
|
GitHub v2 detects modern prefixed PATs:
|
|
|
|
```text
|
|
(ghp|gho|ghu|ghs|ghr|github_pat)_[a-zA-Z0-9_]{36,255}
|
|
```
|
|
|
|
GitHub v1 detects legacy 40-character hex tokens near words such as `github`, `gh`, `pat`, or `token`:
|
|
|
|
```text
|
|
(?:github|gh|pat|token).{0,40}([a-f0-9]{40})
|
|
```
|
|
|
|
GitHubOauth2 detects a 20-character client id and a 40-character client secret near `github`:
|
|
|
|
```text
|
|
client_id: [a-zA-Z0-9]{20}
|
|
client_secret: [a-f0-9]{40}
|
|
Raw = client_id
|
|
RawV2 = client_id + client_secret
|
|
```
|
|
|
|
GitLab v2 detects modern PATs:
|
|
|
|
```text
|
|
glpat-[a-zA-Z0-9\-=_]{20,22}
|
|
```
|
|
|
|
GitLab v1 detects any 20-22 character token-like value near `gitlab` and skips `glpat-` so v2 can handle it:
|
|
|
|
```text
|
|
gitlab ... ([a-zA-Z0-9\-=_]{20,22})
|
|
```
|
|
|
|
Observed local results show high false-positive volume for unverified GitHub v1, GitHubOauth2, and GitLab v1 detections. The current scanner post-filter drops unverified GitHub/GitLab findings that do not match known modern token prefixes. This reduces noise but can hide real legacy/OAuth credentials if verification cannot run.
|
|
|
|
TODO: Prefer a confidence model over hard dropping:
|
|
|
|
```text
|
|
verified -> high confidence
|
|
modern prefix shape -> high/medium confidence
|
|
legacy GitHub v1 / GitHubOauth2 / GitLab v1 -> low confidence unless verified
|
|
```
|
|
|
|
Dashboard should hide low-confidence findings by default, but allow explicit review.
|
|
|
|
## GCP
|
|
|
|
### GCP service account JSON
|
|
|
|
Detector: `GCP`
|
|
|
|
TruffleHog detects JSON blobs containing `auth_provider_x509_cert_url` and parses service-account style credentials.
|
|
|
|
Useful fields already present in the JSON:
|
|
|
|
```text
|
|
type
|
|
project_id
|
|
private_key_id
|
|
private_key
|
|
client_email
|
|
client_id
|
|
auth_uri
|
|
token_uri
|
|
auth_provider_x509_cert_url
|
|
client_x509_cert_url
|
|
```
|
|
|
|
TruffleHog output behavior:
|
|
|
|
```text
|
|
Raw = client_email, or full key JSON if client_email is missing
|
|
RawV2 = full cleaned credential JSON
|
|
Redacted = client_email
|
|
ExtraData.project = project_id
|
|
AnalysisInfo.principal = client_email
|
|
AnalysisInfo.type = type
|
|
```
|
|
|
|
Practical enrichment fields:
|
|
|
|
```text
|
|
gcp_project_id
|
|
gcp_client_email
|
|
gcp_client_id
|
|
gcp_private_key_id
|
|
gcp_credential_type
|
|
```
|
|
|
|
This detector has enough context to verify/function without extra source-code lookup if `RawV2` is preserved.
|
|
|
|
### GCP Application Default Credentials
|
|
|
|
Detector: `GCPApplicationDefaultCredentials`
|
|
|
|
TruffleHog detects ADC JSON containing `client_secret` and `.apps.googleusercontent.com` client IDs.
|
|
|
|
Useful fields:
|
|
|
|
```text
|
|
client_id
|
|
client_secret
|
|
refresh_token
|
|
type
|
|
```
|
|
|
|
TruffleHog output behavior:
|
|
|
|
```text
|
|
Raw = client_id without .apps.googleusercontent.com suffix
|
|
RawV2 = client_id_without_suffix + refresh_token
|
|
Redacted = shortened refresh_token
|
|
ExtraData may contain verification details when verified
|
|
```
|
|
|
|
Risk: `RawV2` is concatenated and does not retain `client_secret` cleanly. The raw finding JSON may not be enough to reconstruct the original ADC JSON unless the source line/file is available.
|
|
|
|
Practical enrichment fields:
|
|
|
|
```text
|
|
gcp_client_id
|
|
gcp_refresh_token_redacted
|
|
gcp_credential_type
|
|
```
|
|
|
|
TODO: For ADC findings, use source context around the finding to parse the whole JSON and preserve `client_secret`/`refresh_token` as structured fields.
|
|
|
|
### Google AQ authentication keys
|
|
|
|
`AQ.` credentials are currently classified through the Gemini Developer API at
|
|
`generativelanguage.googleapis.com`. That result does not establish Vertex AI access.
|
|
|
|
TODO: Add a separate Vertex AI Express probe for `AQ.` credentials against the supported
|
|
`aiplatform.googleapis.com` key-authenticated methods. Keep Gemini Developer API and Vertex
|
|
results independent, and do not infer access to the full project/location-scoped Vertex API
|
|
from the key prefix or from a successful Gemini Developer API check.
|
|
|
|
## Azure
|
|
|
|
### Azure Container Registry
|
|
|
|
Detector: `AzureContainerRegistry`
|
|
|
|
Detector finds registry hosts and ACR password-like values.
|
|
|
|
Patterns:
|
|
|
|
```text
|
|
registry: <name>.azurecr.io
|
|
password: [a-zA-Z0-9+/]{42}+ACR[a-zA-Z0-9]{6}
|
|
```
|
|
|
|
TruffleHog output behavior:
|
|
|
|
```text
|
|
Raw = password
|
|
RawV2 = {"username":"<registry>","password":"<password>"}
|
|
Redacted = registry name
|
|
```
|
|
|
|
Verification uses:
|
|
|
|
```text
|
|
https://<registry>.azurecr.io/v2/
|
|
BasicAuth(username=<registry>, password=<password>)
|
|
```
|
|
|
|
Practical enrichment fields:
|
|
|
|
```text
|
|
azure_acr_registry
|
|
azure_acr_login_server = <registry>.azurecr.io
|
|
```
|
|
|
|
This detector has enough context in `RawV2` to be useful.
|
|
|
|
### Azure OpenAI
|
|
|
|
Detector: `AzureOpenAI`
|
|
|
|
Detector finds API keys and Azure OpenAI endpoints.
|
|
|
|
Patterns:
|
|
|
|
```text
|
|
endpoint: <service>.openai.azure.com
|
|
key: 32 lowercase hex chars near api_key/openai_key keywords
|
|
```
|
|
|
|
TruffleHog output behavior:
|
|
|
|
```text
|
|
Raw = api key
|
|
RawV2 = key:endpoint when endpoint is paired during verification or when only one endpoint exists
|
|
Redacted = shortened key
|
|
```
|
|
|
|
Verification calls:
|
|
|
|
```text
|
|
https://<endpoint>/openai/deployments?api-version=2023-03-15-preview
|
|
Header: Api-Key: <key>
|
|
```
|
|
|
|
Practical enrichment fields:
|
|
|
|
```text
|
|
azure_openai_endpoint
|
|
azure_openai_resource_name
|
|
```
|
|
|
|
TODO: If RawV2 is empty, scan nearby source context for `.openai.azure.com` to pair keys with endpoints.
|
|
|
|
### Azure DevOps PAT
|
|
|
|
Detector: `AzureDevopsPersonalAccessToken`
|
|
|
|
Detector finds a 52-character token and an organization-like string near `azure`.
|
|
|
|
TruffleHog output behavior:
|
|
|
|
```text
|
|
Raw = PAT
|
|
RawV2 = PAT + organization
|
|
```
|
|
|
|
Verification calls:
|
|
|
|
```text
|
|
https://dev.azure.com/<organization>/_apis/projects
|
|
BasicAuth(username="", password=<PAT>)
|
|
```
|
|
|
|
Risk: `RawV2` is concatenated without delimiter, so organization extraction from RawV2 is ambiguous unless the token length is known.
|
|
|
|
Practical enrichment fields:
|
|
|
|
```text
|
|
azure_devops_org
|
|
```
|
|
|
|
TODO: Parse organization from raw finding JSON/source context rather than relying on concatenated RawV2 alone.
|
|
|
|
## DockerHub
|
|
|
|
Detector: `Dockerhub`
|
|
|
|
DockerHub v2 detects modern PATs:
|
|
|
|
```text
|
|
dckr_pat_[a-zA-Z0-9_-]{27}
|
|
```
|
|
|
|
DockerHub v1 detects UUID-like legacy tokens near `docker`.
|
|
|
|
Both versions try to pair the token with nearby usernames or emails:
|
|
|
|
```text
|
|
username: user/usr/-u/id nearby value or email address
|
|
Raw = token
|
|
RawV2 = username:token when username/email is found
|
|
```
|
|
|
|
Verification calls:
|
|
|
|
```text
|
|
POST https://hub.docker.com/v2/users/login
|
|
{"username":"<username>","password":"<token>"}
|
|
```
|
|
|
|
If verified, ExtraData can include:
|
|
|
|
```text
|
|
hub_username
|
|
hub_email
|
|
hub_scope
|
|
2fa_required
|
|
```
|
|
|
|
Practical enrichment fields:
|
|
|
|
```text
|
|
dockerhub_username
|
|
dockerhub_email
|
|
dockerhub_scope
|
|
dockerhub_2fa_required
|
|
```
|
|
|
|
Risk: A token without nearby username cannot be verified by this detector, but may still be useful if a username can be found elsewhere in the same package/repo.
|
|
|
|
TODO: For unverified DockerHub PATs with empty RawV2, scan nearby context and package/repo metadata for plausible DockerHub usernames.
|
|
|
|
## Proposed Enrichment Layer
|
|
|
|
Add a post-processing enrichment layer after TruffleHog result parsing and before DB insert.
|
|
|
|
Input:
|
|
|
|
```text
|
|
finding JSON
|
|
target metadata
|
|
source file path/line if available
|
|
optional nearby source context
|
|
```
|
|
|
|
Output fields stored in DB/dashboard:
|
|
|
|
```text
|
|
provider
|
|
credential_kind
|
|
credential_confidence
|
|
required_context_missing
|
|
principal
|
|
project_id
|
|
tenant_id
|
|
organization
|
|
registry
|
|
endpoint
|
|
username
|
|
email
|
|
scope
|
|
resource
|
|
```
|
|
|
|
Suggested confidence levels:
|
|
|
|
```text
|
|
verified
|
|
structured_complete
|
|
prefix_shape_complete
|
|
token_only_missing_context
|
|
legacy_unverified
|
|
noisy_unverified
|
|
```
|
|
|
|
Priority implementation:
|
|
|
|
1. Parse GCP service-account JSON from `RawV2`.
|
|
2. Parse Azure ACR `RawV2` JSON.
|
|
3. Parse Azure OpenAI `RawV2` as `key:endpoint` when available.
|
|
4. Parse DockerHub `RawV2` as `username:token` and ExtraData when verified.
|
|
5. Add low-confidence classification for GitHub/GitLab legacy detectors instead of hard-dropping them.
|
|
6. Optionally read nearby file context for detectors where RawV2 lacks required context.
|