Evaluation · PII
PII measurement results
The benchmark checks whether a value is found as the right kind of personal data, and whether a value that looks like personal data is sensitive where it sits. It records what happened. It does not grade a product.
Public synthetic measurement
Validated
4 public populations · 10 metrics each. The measurement artifact is validated; this is not a product support approval.
Current product support qualification
Not established
No usable protected support record is bound. The PII protected binding did not validate: PII protected support binding rejected: cost-acceptance-ledger-mismatch Public measurement remains separate.
Protected execution and audit
Not operational
The protected execution and audit paths are pending. They do not gate public synthetic measurement.
Explore measurement results
Select a population or report, then a metric or personal-data category. Each value keeps its own denominator. Public measurement does not establish product support qualification.
Selected result
b11-population-v2-oracle-plan · redact-secret-core 0.1.0-beta.13 · candidate · public synthetic
- authored benign case suppression rate (pii-v1:benign-suppression-rate)Withheld0 / 0 effective N. zero-denominator. Measured 0, eligible 0, unresolved 0, not measured 0.
How a case is judged
- AuthorA case is made-up text with one value in it. The expected answer is written before any scanner runs, from a public authority such as an RFC, ISO 13616, the SSA rules or the NANP plan.
- RunEvery scanner reads the same bytes. The run keeps ranges and actions, never the matched text.
- CompareEach result is checked on two axes: what kind of value it is, and whether it is sensitive where it sits. The reported range is checked against the authored range.
- RecordEach metric keeps its own numerator and denominator. Nothing is added across metrics, families or domains.
| Word | Means |
|---|---|
correct | Found as the expected type. |
miss | Not found. |
invalid-correct | A deliberately invalid look-alike was rejected. |
invalid-accepted | A deliberately invalid look-alike was accepted. |
wrong-family | Found as another family. |
wrong-jurisdiction | Found under another jurisdiction. |
| Word | Means |
|---|---|
correct | Flagged when sensitive, left alone when not. |
miss | A sensitive value was not flagged. |
false-positive | A non-sensitive value was flagged. |
unresolved | Needs review. Stays in the denominator and never counts as a pass. |
Who decides the expected answer
- Identity
- A reference validator (Luhn, ISO 13616 mod-97, SSA allocation, NANP structure, IP syntax) runs on the authored range, never on scanner output.
- Sensitive
- Needs a context rule from the family contract. A validator hit or a keyword alone is rejected as a basis.
- Non-sensitive
- Needs a reserved value from an authority: RFC 2606 and 6761 names, IANA ranges, published test cards, NANPA 555-0100 to 555-0199.
- Not established
- The contract does not decide it. It counts as review required and never as a pass.
Methods that build cases
- type-validation
- Checks the validator state against the authored expectation.
- context-discrimination
- The same value in sensitive, neutral and non-sensitive context.
- pii-benign
- Reserved, documentation, test, public, placeholder and context-negative values.
- jurisdiction-collision
- A value that validates in more than one family.
- mutation
- Invalidates the final digit and checks the result.
- reference-differential
- Compares with a reference. The reference is an observation, not truth.
The 10 pii-v1 metrics
Each validated pii-eval population keeps all ten numerator/denominator and interval or withheld results separate. Where the artifact is schema 1.2, the same ten results are also projected per family and view, with language and control-class strata and the run mode; none of it is presented as family qualification (#618).
| Metric | Population | Counts | Better |
|---|---|---|---|
type-miss-rate | scanner-source × authored valid-type occurrence | type state is miss, of resolved type assertions for authored valid types | Lower |
wrong-family-rate | scanner-source × authored valid-type occurrence | type state is wrong-family, of resolved type assertions for authored valid types | Lower |
wrong-jurisdiction-rate | scanner-source × authored jurisdictional valid-type occurrence | type state is wrong-jurisdiction, of resolved jurisdictional type assertions | Lower |
sensitive-miss-rate | scanner-source × authored sensitive occurrence | sensitivity state is miss, of resolved sensitivity assertions for authored sensitive occurrences | Lower |
non-sensitive-flag-rate | scanner-source × authored non-sensitive occurrence | sensitivity state is false-positive, of resolved sensitivity assertions for authored non-sensitive occurrences | Lower |
context-discrimination-rate | complete scanner-source × authored context trios | both sensitive and non-sensitive endpoints pass, of resolved complete context trios | Higher |
benign-suppression-rate | scanner-source × distinct authored benign case | non-sensitive assertion passes, of resolved authored benign cases | Higher |
jurisdiction-collision-rate | scanner-source × authored jurisdiction collision case | target family and jurisdiction assertion passes, of resolved collision type assertions | Higher |
range-collateral-rate | scanner-source × reported span for authored valid type | range is overbroad or partial, of exact, overbroad, or partial reported spans | Lower |
measurable-share | all scanner-source × authored axis assertions | resolved pass or fail assertions, of all eligible authored axes including unresolved axes | Higher |
Recorded, not graded
The ledger records each outcome against the expected answer. It never turns outcomes into a score, a rank or a verdict about a product. A family's status is a classification by published rules: each metric's interval bound has to be on the right side of its threshold and the protected run has to be met. The protected route gives provisional at most, and every artifact says supportClaims is false.
Coverage of the recorded populations
Public measurement coverage and product qualification are separate records. Select one coverage table; counts from different populations are never added.
Selected coverage
| View · method | Authored cases | Variants |
|---|---|---|
| Oracle plan · schema-only v1Recorded as converted cases read against the schema. This is not a run of the other registered methods. | 146 | 146 |
| Oracle plan · type-validationRegistered. This population records no cases for it. | Not recorded | Not recorded |
| Oracle plan · context-discriminationRegistered. This population records no cases for it. | Not recorded | Not recorded |
| Oracle plan · pii-benignRegistered. This population records no cases for it. | Not recorded | Not recorded |
| Oracle plan · jurisdiction-collisionRegistered. This population records no cases for it. | Not recorded | Not recorded |
| Oracle plan · mutationRegistered. This population records no cases for it. | Not recorded | Not recorded |
| Oracle plan · reference-differentialRegistered. This population records no cases for it. | Not recorded | Not recorded |
| Oracle plan · authored146 occurrences in this view. Occurrences are recorded for the view, not per method. The methods above add up to these two figures. | 146 | 146 |
Population 7b39dab9154e, version 1, public synthetic. Rows are added over families within one view only; the same case can belong to another population or view, so no figure is added across them.
Reading the columns
- Oracle plan
- Authored identity and sensitivity labels, one per plan case.
- Qualification plan
- The frozen plan of each family.
- Diagnostic-balanced
- A tuning view. It does not decide a release.
- Benign-heavy
- An evaluation view dominated by non-sensitive values.
- Protected
- A sealed corpus, run once. Only its size is public.
Scope
- Languages with context evidence
- Not recorded. Language is not jurisdiction.
- Jurisdictions
- Not recorded.
How to read the numbers
- N of M
- A numerator over its own denominator. Two metrics never share a denominator, so there is no total.
- Not measured is not zero
- A dashed mark means nothing was recorded. A zero is a recorded count of none.
- Status words
- Pending, provisional, stable and unsupported are classifications by published rules. Every artifact says supportClaims is false: they are not claims about a product.
- Mode
- Published means a release. Candidate means an unreleased build. A count that depends on the build always names one.
- Two populations
- Diagnostic-balanced and benign-heavy stay apart, so a local change is not hidden by an average.
- Language is not jurisdiction
- Reading Korean labels says nothing about a Korean national identifier.
Sources
- Identity oracle
docs/specs/pii-identity-oracle.md - Benign and collision evidence
docs/specs/pii-benign-collision-evidence.md - Population views
docs/specs/pii-populations.md - Metric profile
qualification/pii-v1.json