Skip to content
Benchmarks

Evaluation · PII

PII measurement results

The benchmark checks whether a value is found as the right kind of personal data, and whether a value that looks like personal data is sensitive where it sits. It records what happened. It does not grade a product.

  • Evaluation profile pii-v1
  • Accounting pii-v1
  • Mode Public synthetic only
  • Support claims None

Public synthetic measurement

Validated

4 public populations · 10 metrics each. The measurement artifact is validated; this is not a product support approval.

Current product support qualification

Not established

No usable protected support record is bound. The PII protected binding did not validate: PII protected support binding rejected: cost-acceptance-ledger-mismatch Public measurement remains separate.

Protected execution and audit

Not operational

The protected execution and audit paths are pending. They do not gate public synthetic measurement.

Explore measurement results

Select a population or report, then a metric or personal-data category. Each value keeps its own denominator. Public measurement does not establish product support qualification.

Selected result

b11-population-v2-oracle-plan · redact-secret-core 0.1.0-beta.13 · candidate · public synthetic

  • authored benign case suppression rate (pii-v1:benign-suppression-rate)Withheld0 / 0 effective N. zero-denominator. Measured 0, eligible 0, unresolved 0, not measured 0.

How a case is judged

  1. AuthorA case is made-up text with one value in it. The expected answer is written before any scanner runs, from a public authority such as an RFC, ISO 13616, the SSA rules or the NANP plan.
  2. RunEvery scanner reads the same bytes. The run keeps ranges and actions, never the matched text.
  3. CompareEach result is checked on two axes: what kind of value it is, and whether it is sensitive where it sits. The reported range is checked against the authored range.
  4. RecordEach metric keeps its own numerator and denominator. Nothing is added across metrics, families or domains.
Type identity
WordMeans
correctFound as the expected type.
missNot found.
invalid-correctA deliberately invalid look-alike was rejected.
invalid-acceptedA deliberately invalid look-alike was accepted.
wrong-familyFound as another family.
wrong-jurisdictionFound under another jurisdiction.
Sensitivity in context
WordMeans
correctFlagged when sensitive, left alone when not.
missA sensitive value was not flagged.
false-positiveA non-sensitive value was flagged.
unresolvedNeeds review. Stays in the denominator and never counts as a pass.

Who decides the expected answer

Identity
A reference validator (Luhn, ISO 13616 mod-97, SSA allocation, NANP structure, IP syntax) runs on the authored range, never on scanner output.
Sensitive
Needs a context rule from the family contract. A validator hit or a keyword alone is rejected as a basis.
Non-sensitive
Needs a reserved value from an authority: RFC 2606 and 6761 names, IANA ranges, published test cards, NANPA 555-0100 to 555-0199.
Not established
The contract does not decide it. It counts as review required and never as a pass.

Methods that build cases

type-validation
Checks the validator state against the authored expectation.
context-discrimination
The same value in sensitive, neutral and non-sensitive context.
pii-benign
Reserved, documentation, test, public, placeholder and context-negative values.
jurisdiction-collision
A value that validates in more than one family.
mutation
Invalidates the final digit and checks the result.
reference-differential
Compares with a reference. The reference is an observation, not truth.
The 10 pii-v1 metrics

Each validated pii-eval population keeps all ten numerator/denominator and interval or withheld results separate. Where the artifact is schema 1.2, the same ten results are also projected per family and view, with language and control-class strata and the run mode; none of it is presented as family qualification (#618).

Metric definitions
MetricPopulationCountsBetter
type-miss-ratescanner-source × authored valid-type occurrencetype state is miss, of resolved type assertions for authored valid typesLower
wrong-family-ratescanner-source × authored valid-type occurrencetype state is wrong-family, of resolved type assertions for authored valid typesLower
wrong-jurisdiction-ratescanner-source × authored jurisdictional valid-type occurrencetype state is wrong-jurisdiction, of resolved jurisdictional type assertionsLower
sensitive-miss-ratescanner-source × authored sensitive occurrencesensitivity state is miss, of resolved sensitivity assertions for authored sensitive occurrencesLower
non-sensitive-flag-ratescanner-source × authored non-sensitive occurrencesensitivity state is false-positive, of resolved sensitivity assertions for authored non-sensitive occurrencesLower
context-discrimination-ratecomplete scanner-source × authored context triosboth sensitive and non-sensitive endpoints pass, of resolved complete context triosHigher
benign-suppression-ratescanner-source × distinct authored benign casenon-sensitive assertion passes, of resolved authored benign casesHigher
jurisdiction-collision-ratescanner-source × authored jurisdiction collision casetarget family and jurisdiction assertion passes, of resolved collision type assertionsHigher
range-collateral-ratescanner-source × reported span for authored valid typerange is overbroad or partial, of exact, overbroad, or partial reported spansLower
measurable-shareall scanner-source × authored axis assertionsresolved pass or fail assertions, of all eligible authored axes including unresolved axesHigher

Recorded, not graded

The ledger records each outcome against the expected answer. It never turns outcomes into a score, a rank or a verdict about a product. A family's status is a classification by published rules: each metric's interval bound has to be on the right side of its threshold and the protected run has to be met. The protected route gives provisional at most, and every artifact says supportClaims is false.

Coverage of the recorded populations

Public measurement coverage and product qualification are separate records. Select one coverage table; counts from different populations are never added.

Selected coverage

Each table describes its own population. Missing coverage is not zero, and public coverage does not qualify product support.
Cases and variants by method · b11-population-v2-oracle-plan · redact-secret-core 0.1.0-beta.13 · Official run · Candidate
View · methodAuthored casesVariants
Oracle plan · schema-only v1Recorded as converted cases read against the schema. This is not a run of the other registered methods.146146
Oracle plan · type-validationRegistered. This population records no cases for it.Not recordedNot recorded
Oracle plan · context-discriminationRegistered. This population records no cases for it.Not recordedNot recorded
Oracle plan · pii-benignRegistered. This population records no cases for it.Not recordedNot recorded
Oracle plan · jurisdiction-collisionRegistered. This population records no cases for it.Not recordedNot recorded
Oracle plan · mutationRegistered. This population records no cases for it.Not recordedNot recorded
Oracle plan · reference-differentialRegistered. This population records no cases for it.Not recordedNot recorded
Oracle plan · authored146 occurrences in this view. Occurrences are recorded for the view, not per method. The methods above add up to these two figures.146146

Population 7b39dab9154e, version 1, public synthetic. Rows are added over families within one view only; the same case can belong to another population or view, so no figure is added across them.

Reading the columns

Oracle plan
Authored identity and sensitivity labels, one per plan case.
Qualification plan
The frozen plan of each family.
Diagnostic-balanced
A tuning view. It does not decide a release.
Benign-heavy
An evaluation view dominated by non-sensitive values.
Protected
A sealed corpus, run once. Only its size is public.

Scope

Languages with context evidence
Not recorded. Language is not jurisdiction.
Jurisdictions
Not recorded.

How to read the numbers

N of M
A numerator over its own denominator. Two metrics never share a denominator, so there is no total.
Not measured is not zero
A dashed mark means nothing was recorded. A zero is a recorded count of none.
Status words
Pending, provisional, stable and unsupported are classifications by published rules. Every artifact says supportClaims is false: they are not claims about a product.
Mode
Published means a release. Candidate means an unreleased build. A count that depends on the build always names one.
Two populations
Diagnostic-balanced and benign-heavy stay apart, so a local change is not hidden by an average.
Language is not jurisdiction
Reading Korean labels says nothing about a Korean national identifier.

Sources