Evaluation · PII
How PII is evaluated
The benchmark checks whether a value is found as the right kind of personal data, and whether a value that looks like personal data is sensitive where it sits. It records what happened. It does not grade a product.
Looking for product capabilities?
Open PII coverage for the release-bound catalog, activation settings and limitations.
Evidence, evaluator and benchmark responsibilities
- pii-evidence on GitHub
Authors source cases, taxonomy, evidence provenance and expected-answer contracts before a scanner runs.
- pii-eval on GitHub
Runs source-bound measurement and records authored type, range and sensitivity/context assertions with separate metric denominators.
- redact-secret core on GitHub
Implements the scanner and its explicit activation configuration; scanner findings do not author the expected answer.
- Benchmarks on GitHub
Consumes validated artifacts, keeps populations separate and owns interpretation and qualification policy. It does not adopt evidence or infer a support verdict from this overview.
How a case is judged
- AuthorA case is made-up text with one value in it. The expected answer is written before any scanner runs, from a public authority such as an RFC, ISO 13616, the SSA rules or the NANP plan.
- RunEvery scanner reads the same bytes. The run keeps ranges and actions, never the matched text.
- CompareEach result is checked on two axes: what kind of value it is, and whether it is sensitive where it sits. The reported range is checked against the authored range.
- RecordEach metric keeps its own numerator and denominator. Nothing is added across metrics, families or domains.
| Word | Means |
|---|---|
correct | Found as the expected type. |
miss | Not found. |
invalid-correct | A deliberately invalid look-alike was rejected. |
invalid-accepted | A deliberately invalid look-alike was accepted. |
wrong-family | Found as another family. |
wrong-jurisdiction | Found under another jurisdiction. |
| Word | Means |
|---|---|
correct | Flagged when sensitive, left alone when not. |
miss | A sensitive value was not flagged. |
false-positive | A non-sensitive value was flagged. |
unresolved | Needs review. Stays in the denominator and never counts as a pass. |
Who decides the expected answer
- Identity
- A reference validator (Luhn, ISO 13616 mod-97, SSA allocation, NANP structure, IP syntax) runs on the authored range, never on scanner output.
- Sensitive
- Needs a context rule from the family contract. A validator hit or a keyword alone is rejected as a basis.
- Non-sensitive
- Needs a reserved value from an authority: RFC 2606 and 6761 names, IANA ranges, published test cards, NANPA 555-0100 to 555-0199.
- Not established
- The contract does not decide it. It counts as review required and never as a pass.
Methods that build cases
- type-validation
- Checks the validator state against the authored expectation.
- context-discrimination
- The same value in sensitive, neutral and non-sensitive context.
- pii-benign
- Reserved, documentation, test, public, placeholder and context-negative values.
- jurisdiction-collision
- A value that validates in more than one family.
- mutation
- Invalidates the final digit and checks the result.
- reference-differential
- Compares with a reference. The reference is an observation, not truth.
pii-v1 metric definitions and separate denominators
Each validated pii-eval population keeps all ten numerator/denominator and interval or withheld results separate. Where the artifact is schema 1.2, the same ten results are also projected per family and view, with language and control-class strata and the run mode; none of it is presented as family qualification (#618).
| Metric | Population | Counts | Better |
|---|---|---|---|
pii-v1:type-miss-rate | scanner-source × authored valid-type occurrence | type state is miss, of resolved type assertions for authored valid types | Lower |
pii-v1:wrong-family-rate | scanner-source × authored valid-type occurrence | type state is wrong-family, of resolved type assertions for authored valid types | Lower |
pii-v1:wrong-jurisdiction-rate | scanner-source × authored jurisdictional valid-type occurrence | type state is wrong-jurisdiction, of resolved jurisdictional type assertions | Lower |
pii-v1:sensitive-miss-rate | scanner-source × authored sensitive occurrence | sensitivity state is miss, of resolved sensitivity assertions for authored sensitive occurrences | Lower |
pii-v1:non-sensitive-flag-rate | scanner-source × authored non-sensitive occurrence | sensitivity state is false-positive, of resolved sensitivity assertions for authored non-sensitive occurrences | Lower |
pii-v1:context-discrimination-rate | complete scanner-source × authored context trios | both sensitive and non-sensitive endpoints pass, of resolved complete context trios | Higher |
pii-v1:benign-suppression-rate | scanner-source × distinct authored benign case | non-sensitive assertion passes, of resolved authored benign cases | Higher |
pii-v1:jurisdiction-collision-rate | scanner-source × authored jurisdiction collision case | target family and jurisdiction assertion passes, of resolved collision type assertions | Higher |
pii-v1:range-collateral-rate | scanner-source × reported span for authored valid type | range is overbroad or partial, of exact, overbroad, or partial reported spans | Lower |
pii-v1:measurable-share | all scanner-source × authored axis assertions | resolved pass or fail assertions, of all eligible authored axes including unresolved axes | Higher |
Recorded, not graded
The ledger records each outcome against the expected answer. It never turns outcomes into a score, a rank or a verdict about a product. A family's status is a classification by published rules: each metric's interval bound has to be on the right side of its threshold and the protected run has to be met. The protected route gives provisional at most, and every artifact says supportClaims is false.
Metric groups and separate denominators
pii-v1 and b11 are different quantities. pii-v1 is neutral accounting owned by pii-eval, using authored occurrences, context trios, spans or axis assertions as each metric defines. b11 is the benchmark-owned product scorer, using scored cases, twin pairs or findings. The same ten metric IDs do not give them the same denominator or meaning.
- Type identity
Type miss, wrong family and wrong jurisdiction use resolved authored type assertions; jurisdictional assertions have their own eligible population.
- Sensitivity and benign context
Sensitive miss and non-sensitive flag use their respective resolved authored sensitivity assertions. Context discrimination uses complete trios; benign suppression uses distinct authored benign cases.
- Collision and range
Jurisdiction collision uses resolved collision assertions. Range collateral uses reported exact, overbroad and partial spans; it is not a rate over every evidence case.
- Measurable share
Resolved pass or fail assertions over all eligible authored axes, including unresolved axes. It reports measurement availability, not product support.
Unresolved assertions remain visible. Withheld values state why a value cannot be published; neither state is a zero or a pass. Measurable share includes all eligible authored axes, including unresolved axes.
Complete ten-metric definitions for both protocols
The accepted interpretation policy keeps product thresholds and verdicts on b11 quantities. Those thresholds are not applied to pii-v1 observations, and no qualification transfers between protocols. Benchmark qualification profile and internal thresholds
Population and qualification boundaries
Each public synthetic population retains its own identity, methods, numerator, denominator and interval. Diagnostic-balanced, benign-heavy, separate public evidence and protected populations are never pooled. Synthetic rates are not production rates.
Language and jurisdiction are separate dimensions. Korean or English context labels do not establish support for a national identifier. A type validator does not decide sensitivity in context.
Current product qualification limitation
Current product support qualification: Not established. No usable protected support record is bound. The PII protected binding did not validate: PII protected support binding rejected: cost-acceptance-ledger-mismatch Public measurement remains separate. Protected execution and audit: Not operational. The protected execution and audit paths are pending. They do not gate public synthetic measurement.
The benchmark-owned support policy requires its recorded gates and bound protected evidence. Public measurement validity alone never establishes product qualification. Ground truth authored before scanning, validator, cost and protected-path requirements remain unavailable where the recorded contract has not established them.
Recorded results and diagnostic details
Population results, all metric selectors, method coverage, sources and execution details
Separate public evidence population and complete evidence-kind matrix