Evaluation · Credential
How credentials are evaluated
How authored credential evidence becomes scanner observations, and how benchmarks applies its own qualification policy.
Looking for product coverage?
Evidence, execution and qualification
credential-evidence authors synthetic inputs, expected spans and controls from provider documentation, tool corroboration and project-policy decisions. Expected answers precede scanner output.
credential-eval runs pinned configurations on those bytes and records span outcomes. A peer scanner is an observation, never the expected-answer authority.
benchmarks owns the published qualification criteria and applies them to recorded evidence. The evaluator does not authorize product support. redact-secret owns product behavior and declared scope.
How a case is judged
- AuthorA fixture is made-up text with the secret’s byte range marked. The expected answer comes from provider documentation, tool corroboration or project policy, never from a scanner.
- RunEvery scanner reads the same bytes. Another scanner is never the oracle.
- CompareEach secret span gets one outcome. A control file is flagged or not.
- RecordA rate carries a Wilson interval and is withheld under a minimum count. There is no overall score.
| Word | Means |
|---|---|
EXACT | The reported range equals the secret. |
COVERED | Inside the allowed envelope around the secret. |
OVERBROAD | Redacts more than the envelope allows. |
PARTIAL | Part of the secret is left readable. |
MISS | Not reported. |
| Word | Means |
|---|---|
T1 | Provider-documented. |
T2 | Tool-corroborated. |
T3 | Project policy. |
T0 | Pending review. Observed, never scored. |
Who decides the expected answer
- T1
- The provider’s own documentation states the format.
- T2
- A tool or a structural check corroborates it.
- T3
- The project states a policy, with its rationale.
- Peers
- Another scanner’s output is a comparison, never the truth.
Methods that build cases
- twin
- Authored positive and negative pairs.
- benign
- Controls that must stay silent.
- mutation
- Seeded edits with a stated effect.
- metamorphic
- The same secret in another context or encoding.
- differential
- Compares scanners. Disagreements go to review.
- holdout
- Isolated frozen cases, aggregate output only.
Metric definitions and separate denominators
Each rate is read against its own population and carries a Wilson bound. None is combined with another or with a PII metric.
| Metric | Population | Counts | Better |
|---|---|---|---|
leaked-span-rate | scanner × expected secret span | span is PARTIAL or MISS, of expected spans | Lower |
false-alarm-rate | scanner × must-not-flag file | file is flagged, of control files | Lower |
collateral-ratio | scanner × secret byte | bytes redacted outside the envelope, of secret bytes | Lower |
Qualification belongs to benchmarks
The evaluator records outcomes against authored answers. Benchmarks applies published evidence floors and interval thresholds to classify support; evidence class and product capability remain separate. A policy-qualified route requires its own accepted evidence and protected holdout, never an inferred approval from scanner observations.
Current limits and policy interpretation
Provider-documented, tool-corroborated and project-policy evidence remain distinct. Documented, empirical and policy-qualified benchmark classifications use their own evidence floors and criteria. Generic credential carriers can depend on context and explicit policy; they do not grant a provider-specific claim.
The bounded credential policy contract requires carrier context, value/span/action rules and benign/twin controls authored before scanner execution to distinguish ordinary values and look-alikes. Project-policy evidence does not become provider-documented evidence; outside-contract matches remain unknown or incidental. The qualification rationale and exclusions explains why a bounded policy classification does not establish live validity or provider-specific qualification.
Policy-qualified profile
Not measured. The profile needs a protected holdout on a frozen candidate, and none is committed. Exact limitation record
Live validity
Not measured. Fixtures are made up, so no credential is checked against its provider.
Real-world rate
Not measured. An interval describes the corpus, not production traffic.
Other evidence, kept apart
Recorded. Custodian-blind evaluation, score evasion and mixed parity are separate classes and are not merged into one number. Exact limitation record
Intervals describe the authored population. Unresolved, withheld and unavailable observations retain their stated reasons. None is a production traffic rate or a pooled cross-population score.
Detailed evidence and results
The fixture-kind and evidence-level inventory moved to the credential corpus detail. This preserves the former overview bookmark.
Recorded case and variant counts by method now live in qualification method results and each method’s detail page.
How to read the numbers
- N of M
- A numerator over its own denominator. Two metrics never share a denominator, so there is no total.
- Not measured is not zero
- A dashed mark means nothing was recorded. A zero is a recorded count of none.
- Status words
- Pending, provisional, stable and unsupported are classifications by published rules. Every artifact says supportClaims is false: they are not claims about a product.
- Mode
- Published means a release. Candidate means an unreleased build. A count that depends on the build always names one.
- Levels stay apart
- T1, T2 and T3 are different evidence. They are shown side by side and never merged.
- An interval describes the corpus
- A Wilson bound says how sure the count is on these fixtures, not what happens in production.
Sources
- Measurement protocol
docs/specs/measurement-v4.md - Evaluation engine
docs/specs/evaluation-engine-v1.md - Support status criteria
docs/specs/support-status.md - Taxonomy
docs/specs/taxonomy.md