Skip to content
Benchmarks

REDACT-SECRET 0.1.0-BETA.14

What the benchmark shows

Synthetic inputs, the same for every scanner, scored span by span. Start from a provider or a family, or see what changed.

PROVIDER-DOCUMENTED

Three answers

  • Run 2026-10-07
  • Same 7,036 inputs for 4 scanners
  • Accounting v1.1
  • Mode published · redact-secret 0.1.0-beta.14

Does it miss real secrets?

at most 2.3%

23 of 1,468 secret spans leaked

Leaked span rate. Lower is better. 95% pessimistic bound, corpus-relative.

Does it flag safe values?

at most 2.2%

0 of 167 controls flagged

False alarm rate on provider-documented controls. Lower is better. Few controls keep the bound wide.

Does it tell near-twins apart?

at least 97.9%

808 of 817 pairs discriminated

Twin discrimination: the secret is covered and its one-character fake stays quiet. Higher is better.

TOOL-CORROBORATED

Three answers

  • Run 2026-10-07
  • Same 7,036 inputs for 4 scanners
  • Accounting v1.1
  • Mode published · redact-secret 0.1.0-beta.14

Does it miss real secrets?

at most 3.1%

9 of 548 secret spans leaked

Leaked span rate. Lower is better. 95% pessimistic bound, corpus-relative.

Does it flag safe values?

at most 1.3%

14 of 1,866 controls flagged

False alarm rate on tool-corroborated controls. Lower is better. Few controls keep the bound wide.

Does it tell near-twins apart?

not published —

230 of 233 pairs discriminated Withheld

Twin discrimination: the secret is covered and its one-character fake stays quiet. Higher is better. Too few secrets have an authored near-twin to publish a rate.

PROJECT POLICY

Three answers

  • Run 2026-10-07
  • Same 7,036 inputs for 4 scanners
  • Accounting v1.1
  • Mode published · redact-secret 0.1.0-beta.14

Does it leave policy spans readable?

at most 7.1%

58 of 1,053 secret spans leaked

Leaked span rate. Lower is better. 95% pessimistic bound, corpus-relative. These spans are this project’s redaction policy: a difference here is a difference of opinion, not a defect.

Does it flag safe values?

at most 0.9%

9 of 1,880 controls flagged

False alarm rate on project policy controls. Lower is better. Few controls keep the bound wide.

Does it tell near-twins apart?

at least 96.8%

308 of 312 pairs discriminated

Twin discrimination: the secret is covered and its one-character fake stays quiet. Higher is better.

Other scanners on the same inputs

We ran 3 other scanners on the same 1,440 provider-documented inputs. This shows what each one left readable. It does not show which scanner is better.

Read this before the numbers

  1. Our inputs, our answer key. The redact-secret team wrote every input and every expected span, using redact-secret’s own definition of a secret.
  2. redact-secret was tuned on these inputs. 56 of 64 findings from this corpus are recorded as fixed in redact-secret. The other scanners were never tuned against it.
  3. Different jobs. Most inputs fall outside at least one scanner’s rules. A readable span there shows where its rules end, not that it failed.
Other scanners on the same inputs
ScannerInputs its rules targetLeft readable on those inputsLeft readable everywhere elseLeft readable on all inputsSafe values flagged
flare-redact 1.6.1Runtime library · Published npm packageBuilt to redact secrets and personal data from text at runtime. Run secrets-only here: its personal-data and generic-assignment detectors are off.468 of 1,44038 of its 81 rules target a credential family48 of 485 spansredact-secret, same inputs: 4 of 485870 of 983 spansNo rule of its own targets these918 of 1,468 spansredact-secret, same inputs: 23 of 1,4686 of 167redact-secret, same inputs: 0 of 167
gitleaks 8.30.1Repository scanner · Directory scanBuilt to find secrets in git history, files and directories before they are committed.635 of 1,44073 of its 222 rules target a credential family92 of 646 spansredact-secret, same inputs: 4 of 646426 of 822 spansNo rule of its own targets these518 of 1,468 spansredact-secret, same inputs: 23 of 1,46810 of 167redact-secret, same inputs: 0 of 167
trufflehog 3.97.4Repository scanner · Filesystem scanBuilt to find and verify secrets in repositories and other sources. Verification is off here.722 of 1,44083 of its 892 rules target a credential family168 of 739 spansredact-secret, same inputs: 2 of 739715 of 729 spansNo rule of its own targets these883 of 1,468 spansredact-secret, same inputs: 23 of 1,4684 of 167redact-secret, same inputs: 0 of 167

Observed once each, in the official run sha256:4bec6e539482…, recorded 2026-10-07. Listed in run order; a new scanner adds a row. Rules are matched to families from each scanner’s pinned rule file (reviewed 2026-09-30).

Quoting these numbers

Don't “redact-secret leaks far fewer secrets than flare-redact.”

Do “On 468 provider-documented inputs that match flare-redact 1.6.1’s default rules, in a corpus written and used for tuning by the redact-secret team, flare-redact left 48 of 485 secret spans readable.”

Any quote names the input slice, the version, the date, and who wrote the inputs.

Other scanners on the same inputs

We ran 3 other scanners on the same 548 tool-corroborated inputs. This shows what each one left readable. It does not show which scanner is better.

Read this before the numbers

  1. Our inputs, our answer key. The redact-secret team wrote every input and every expected span, using redact-secret’s own definition of a secret.
  2. redact-secret was tuned on these inputs. 56 of 64 findings from this corpus are recorded as fixed in redact-secret. The other scanners were never tuned against it.
  3. Different jobs. Most inputs fall outside at least one scanner’s rules. A readable span there shows where its rules end, not that it failed.
Other scanners on the same inputs
ScannerInputs its rules targetLeft readable on those inputsLeft readable everywhere elseLeft readable on all inputsSafe values flagged
flare-redact 1.6.1Runtime library · Published npm packageBuilt to redact secrets and personal data from text at runtime. Run secrets-only here: its personal-data and generic-assignment detectors are off.195 of 54838 of its 81 rules target a credential family44 of 195 spansredact-secret, same inputs: 9 of 195316 of 353 spansNo rule of its own targets these360 of 548 spansredact-secret, same inputs: 9 of 54868 of 1,866redact-secret, same inputs: 14 of 1,866
gitleaks 8.30.1Repository scanner · Directory scanBuilt to find secrets in git history, files and directories before they are committed.268 of 54873 of its 222 rules target a credential family49 of 268 spansredact-secret, same inputs: 9 of 268117 of 280 spansNo rule of its own targets these166 of 548 spansredact-secret, same inputs: 9 of 548121 of 1,866redact-secret, same inputs: 14 of 1,866
trufflehog 3.97.4Repository scanner · Filesystem scanBuilt to find and verify secrets in repositories and other sources. Verification is off here.336 of 54883 of its 892 rules target a credential family111 of 336 spansredact-secret, same inputs: 9 of 336212 of 212 spansNo rule of its own targets these323 of 548 spansredact-secret, same inputs: 9 of 54862 of 1,866redact-secret, same inputs: 14 of 1,866

Observed once each, in the official run sha256:4bec6e539482…, recorded 2026-10-07. Listed in run order; a new scanner adds a row. Rules are matched to families from each scanner’s pinned rule file (reviewed 2026-09-30).

Quoting these numbers

Don't “redact-secret leaks far fewer secrets than flare-redact.”

Do “On 195 tool-corroborated inputs that match flare-redact 1.6.1’s default rules, in a corpus written and used for tuning by the redact-secret team, flare-redact left 44 of 195 secret spans readable.”

Any quote names the input slice, the version, the date, and who wrote the inputs.

Other scanners on the same inputs

Hidden by default. Project policy is this project’s masking policy: a peer’s rate here reflects scope, not accuracy, because a peer is not built to flag it. Show anyway.

Other scanners on the same inputs

We ran 3 other scanners on the same 1,006 project policy inputs. This shows what each one left readable. It does not show which scanner is better.

Read this before the numbers

  1. Our inputs, our answer key. The redact-secret team wrote every input and every expected span, using redact-secret’s own definition of a secret.
  2. redact-secret was tuned on these inputs. 56 of 64 findings from this corpus are recorded as fixed in redact-secret. The other scanners were never tuned against it.
  3. Different jobs. Most inputs fall outside at least one scanner’s rules. A readable span there shows where its rules end, not that it failed.
Other scanners on the same inputs
ScannerInputs its rules targetLeft readable on those inputsLeft readable everywhere elseLeft readable on all inputsSafe values flagged
flare-redact 1.6.1Runtime library · Published npm packageBuilt to redact secrets and personal data from text at runtime. Run secrets-only here: its personal-data and generic-assignment detectors are off.365 of 1,00638 of its 81 rules target a credential family50 of 403 spansredact-secret, same inputs: 20 of 403564 of 650 spansNo rule of its own targets these614 of 1,053 spansredact-secret, same inputs: 58 of 1,05348 of 1,880redact-secret, same inputs: 9 of 1,880
gitleaks 8.30.1Repository scanner · Directory scanBuilt to find secrets in git history, files and directories before they are committed.517 of 1,00673 of its 222 rules target a credential family212 of 554 spansredact-secret, same inputs: 54 of 554298 of 499 spansNo rule of its own targets these510 of 1,053 spansredact-secret, same inputs: 58 of 1,05352 of 1,880redact-secret, same inputs: 9 of 1,880
trufflehog 3.97.4Repository scanner · Filesystem scanBuilt to find and verify secrets in repositories and other sources. Verification is off here.519 of 1,00683 of its 892 rules target a credential family459 of 555 spansredact-secret, same inputs: 20 of 555483 of 498 spansNo rule of its own targets these942 of 1,053 spansredact-secret, same inputs: 58 of 1,05334 of 1,880redact-secret, same inputs: 9 of 1,880

Observed once each, in the official run sha256:4bec6e539482…, recorded 2026-10-07. Listed in run order; a new scanner adds a row. Rules are matched to families from each scanner’s pinned rule file (reviewed 2026-09-30).

Quoting these numbers

Don't “redact-secret leaks far fewer secrets than flare-redact.”

Do “On 365 project policy inputs that match flare-redact 1.6.1’s default rules, in a corpus written and used for tuning by the redact-secret team, flare-redact left 50 of 403 secret spans readable.”

Any quote names the input slice, the version, the date, and who wrote the inputs.

Hide