Skip to content
Benchmarks

Evaluation · Credential

How credentials are evaluated

How authored credential evidence becomes scanner observations, and how benchmarks applies its own qualification policy.

  • Evaluation profile qualification-view
  • Accounting 1.1
  • Mode published · redact-secret 0.1.0-beta.14
  • Support claims None

Evidence, execution and qualification

credential-evidence authors synthetic inputs, expected spans and controls from provider documentation, tool corroboration and project-policy decisions. Expected answers precede scanner output.

credential-eval runs pinned configurations on those bytes and records span outcomes. A peer scanner is an observation, never the expected-answer authority.

benchmarks owns the published qualification criteria and applies them to recorded evidence. The evaluator does not authorize product support. redact-secret owns product behavior and declared scope.

How a case is judged

  1. AuthorA fixture is made-up text with the secret’s byte range marked. The expected answer comes from provider documentation, tool corroboration or project policy, never from a scanner.
  2. RunEvery scanner reads the same bytes. Another scanner is never the oracle.
  3. CompareEach secret span gets one outcome. A control file is flagged or not.
  4. RecordA rate carries a Wilson interval and is withheld under a minimum count. There is no overall score.
Outcome of a secret span
WordMeans
EXACTThe reported range equals the secret.
COVEREDInside the allowed envelope around the secret.
OVERBROADRedacts more than the envelope allows.
PARTIALPart of the secret is left readable.
MISSNot reported.
Level of evidence
WordMeans
T1Provider-documented.
T2Tool-corroborated.
T3Project policy.
T0Pending review. Observed, never scored.

Who decides the expected answer

T1
The provider’s own documentation states the format.
T2
A tool or a structural check corroborates it.
T3
The project states a policy, with its rationale.
Peers
Another scanner’s output is a comparison, never the truth.

Methods that build cases

twin
Authored positive and negative pairs.
benign
Controls that must stay silent.
mutation
Seeded edits with a stated effect.
metamorphic
The same secret in another context or encoding.
differential
Compares scanners. Disagreements go to review.
holdout
Isolated frozen cases, aggregate output only.
Metric definitions and separate denominators

Each rate is read against its own population and carries a Wilson bound. None is combined with another or with a PII metric.

Metric definitions
MetricPopulationCountsBetter
leaked-span-ratescanner × expected secret spanspan is PARTIAL or MISS, of expected spansLower
false-alarm-ratescanner × must-not-flag filefile is flagged, of control filesLower
collateral-ratioscanner × secret bytebytes redacted outside the envelope, of secret bytesLower

Qualification belongs to benchmarks

The evaluator records outcomes against authored answers. Benchmarks applies published evidence floors and interval thresholds to classify support; evidence class and product capability remain separate. A policy-qualified route requires its own accepted evidence and protected holdout, never an inferred approval from scanner observations.

Current limits and policy interpretation

Provider-documented, tool-corroborated and project-policy evidence remain distinct. Documented, empirical and policy-qualified benchmark classifications use their own evidence floors and criteria. Generic credential carriers can depend on context and explicit policy; they do not grant a provider-specific claim.

The bounded credential policy contract requires carrier context, value/span/action rules and benign/twin controls authored before scanner execution to distinguish ordinary values and look-alikes. Project-policy evidence does not become provider-documented evidence; outside-contract matches remain unknown or incidental. The qualification rationale and exclusions explains why a bounded policy classification does not establish live validity or provider-specific qualification.

Policy-qualified profile

Not measured. The profile needs a protected holdout on a frozen candidate, and none is committed. Exact limitation record

Live validity

Not measured. Fixtures are made up, so no credential is checked against its provider.

Real-world rate

Not measured. An interval describes the corpus, not production traffic.

Other evidence, kept apart

Recorded. Custodian-blind evaluation, score evasion and mixed parity are separate classes and are not merged into one number. Exact limitation record

Intervals describe the authored population. Unresolved, withheld and unavailable observations retain their stated reasons. None is a production traffic rate or a pooled cross-population score.

Detailed evidence and results

The fixture-kind and evidence-level inventory moved to the credential corpus detail. This preserves the former overview bookmark.

Recorded case and variant counts by method now live in qualification method results and each method’s detail page.

How to read the numbers

N of M
A numerator over its own denominator. Two metrics never share a denominator, so there is no total.
Not measured is not zero
A dashed mark means nothing was recorded. A zero is a recorded count of none.
Status words
Pending, provisional, stable and unsupported are classifications by published rules. Every artifact says supportClaims is false: they are not claims about a product.
Mode
Published means a release. Candidate means an unreleased build. A count that depends on the build always names one.
Levels stay apart
T1, T2 and T3 are different evidence. They are shown side by side and never merged.
An interval describes the corpus
A Wilson bound says how sure the count is on these fixtures, not what happens in production.

Sources