tests/test_methodology.py
parses the thresholds below and asserts them against the running scanner, so a number here
that disagreed with the software would fail the build. It is generated from the same file,
which is why it cannot quietly fall out of date.
ProvenGone Measurement Methodology
Version 1.9 · 2026-08-08 · applies to every number we publish, including our own
This document exists so our results can be checked rather than believed. If you think a ProvenGone figure is wrong, this is the document that lets you demonstrate it.
We audit a market we compete in. That is only defensible if the method is public, fixed in advance, and applied to us on the same terms as everyone else. All three below are commitments, not aspirations:
1. This methodology is published and versioned. Changes appear in the changelog with a date. 2. We are measured by it, with no exemption, and our results appear on the same board. 3. Every audited provider's result is published, including failures. No one can pay to be removed from the board.
1. What we measure
One question: for a given person and a given data broker, is a listing for that person present on the broker's public site, and does it stay absent after a removal request?
What we do not measure: anything requiring an account, a purchase, or a login. We read the same public pages any visitor sees. We make no claim about data a broker holds internally but does not publish — we cannot see it, so we do not report on it.
2. Scan states
Every individual check resolves to exactly one of three states. There is no two-state version of this, and the third state is the reason this document exists.
| State | Meaning |
|---|---|
LISTED | The page loaded and a listing matched the subject at or above the match threshold. |
NOT_FOUND | The page loaded successfully and no listing met the threshold. |
INCONCLUSIVE | We could not establish either. |
INCONCLUSIVE is recorded whenever any of the following holds:
- the request was blocked, rate-limited, timed out, or met an anti-bot challenge
- the page loaded but our parser failed, indicating the site changed shape
- the best match landed in the ambiguous confidence band (§3)
- the broker showed a subset of its matches and the subject was not among them (§5)
A failure to observe is never recorded as an absence. Collapsing INCONCLUSIVE into NOT_FOUND is the single practice that makes published removal rates across this industry unreliable, because it converts a rate limit into a removal.
3. Match scoring
A candidate listing is scored against the subject in basis points (0–1000). Every signal is named and appears in the evidence report.
Surname is a hard gate. A candidate whose surname does not match scores 0 and is discarded.
Surname comparison ignores generational suffixes (Jr, Sr, II, III) and treats hyphenated and spaced compound surnames as equivalent, since sources disagree about both. It still requires the compound parts in order — "Garcia Lopez" and "Lopez Garcia" are different people — and requires a given name to be present, because a bare surname identifies nobody.
City and state are normalised on both sides before comparison, so "Ft. Worth" matches "Fort Worth", "Saint Louis" matches "St. Louis", and "Texas" matches "TX". Sources disagree about all of these, and a formatting difference alone was enough to drop a genuine match from LISTED to NOT_FOUND. Normalization expands abbreviated prefixes and strips a trailing state from the city field; it does not merge distinct places — "Fort Smith" is not "Fort Worth".
| Signal | Points |
|---|---|
| Given name exact | +350 |
| Given name variant (documented nickname or initial) | +200 |
| State matches | +250 |
| City matches | +200 |
| Age within 2 years | +150 |
| Age differs by 10+ years | −200 |
| Middle initial matches | +50 |
| Each shared relative (capped at 300) | +150 |
Thresholds
| Band | Result |
|---|---|
| ≥ 700 | LISTED |
| 400 – 699 | INCONCLUSIVE — a plausible match we will not assert either way |
| < 400 | not this person; contributes nothing |
The 400–699 band is deliberate dead space. A same-named person in the same state, with no city or age to separate them, lands there and stays unresolved rather than being forced into a yes or a no.
4. Removal states
A removal is a claim about history, derived from a broker's full scan timeline.
| State | Requires |
|---|---|
NEVER_LISTED | No LISTED scan has ever been recorded. |
LISTED | Currently listed, no opt-out submitted. |
REMOVAL_PENDING | Opt-out submitted, still listed. |
REMOVAL_UNCONFIRMED | One clean scan since removal. Not a removal. |
REMOVAL_VERIFIED | 2+ consecutive clean scans, ≥14 days apart, uninterrupted. |
RELISTED | Listed again after a clean scan. Reported prominently. |
UNKNOWN | The most recent scan was inconclusive. |
One clean scan is never a removal. A listing can disappear from a search index temporarily and return, and re-listing is common enough that a single check proves little. REMOVAL_VERIFIED requires the absence to persist across at least 14 days with no inconclusive result breaking the streak.
Any recent inconclusive result downgrades the state to UNKNOWN, including from REMOVAL_VERIFIED. If we lose the ability to observe a removal we previously confirmed, we stop claiming it.
We do not take credit for disappearances we did not ask for. A listing can vanish from a broker's own data refresh. Where no removal request lies behind the absence, the report says so (attributed_to_request: false) and the case is excluded from the published removal rate.
Relistings are reported even after they resolve. "Removed, reappeared, removed" is a materially different history from "removed" — it says something about that broker's data hygiene, and it is the customer's best reason to keep watching.
5. Truncated result sets
Brokers paginate. When a page states it holds 221 records and renders 10, failing to find the subject among those 10 says nothing about the other 211.
Where a broker declares a total, or exposes pagination, and we have seen only part of it, the result set is marked incomplete. Then:
- a match still counts — finding someone is finding them
- a non-match cannot be
NOT_FOUND— it is recorded asINCONCLUSIVE
Some brokers never return a searchable list at all, resolving instead to a single best-match profile. Those are always treated as incomplete.
Insufficient subject detail
A name alone cannot distinguish someone from a namesake. With no city, state or age, an exact name match reaches only 350bp — below the match threshold — so every listing would score as "not this person" and the scan would report absence. Tested against real captured pages, a name-only subject was told they were absent from seven brokers that were listing them.
That absence describes our inputs, not the broker's data. Where the subject gave no distinguishing detail and a name match exists that we cannot confirm, the result is INCONCLUSIVE.
A name nobody holds still returns NOT_FOUND. Otherwise a customer who declines to share their address — an entirely reasonable position for a privacy customer — could never have a removal verified.
Unfiltered nationwide searches
A subtler version of the same problem, with no pagination marker to give it away. Several brokers' search URLs cannot be narrowed to a state — BeenVerified's is /people/{first}-{last}/. For a rare surname that returns everything they hold. For a common one it returns an arbitrary slice.
Searching "John Smith, Dallas TX" on such an endpoint returned 69 records, none from Texas — not because no John Smith in Texas is listed, but because we were shown 69 of them from Georgia, New Jersey and Missouri. Absence there is not a finding.
So where the search could not be filtered to the subject's state, the result set came back full, and no record matches that state, the result is INCONCLUSIVE.
The rule stays deliberately narrow. A small unfiltered result set is still a real answer, and if these brokers could never return NOT_FOUND, a verified removal on them would be unreachable — the same trap described under limitations below.
6. Published numbers
Three figures. The removal rate is never published alone.
Removal rate
removal rate = REMOVAL_VERIFIED / brokers where the subject was LISTED and an opt-out was submitted
Only REMOVAL_VERIFIED counts as success. UNKNOWN stays in the denominator, so failing to measure costs exactly as much as failing to remove. A provider cannot improve this number by breaking a scanner.
Measurement quality
measurement quality = (eligible − unknown) / eligible
The share of eligible cases we could actually resolve.
Coverage
coverage = brokers whose pages we retrieved / brokers attempted
Coverage exists because the removal rate alone is gameable. A broker that blocks us from the first scan is never established as LISTED, so it never enters the rate's denominator — it vanishes rather than counting against us. Without coverage, any provider could publish a flawless rate by measuring only the brokers it finds easy.
100% removal across 5 readable brokers out of 50 is a worse product than 75% across 45. A published number that cannot express that difference is not an honest one.
Distinct sources, not domains
Coverage is quoted against brokers that actually hold their own data. Some sites in this industry are affiliate fronts: Persopo's search form redirects to TruthFinder with tracking parameters, and Addresses.com links exclusively to Intelius. They are separate domains serving one dataset.
Counting a front as a separate broker inflates coverage with duplicates, and worse, a removal secured at the real source would appear to independently clear the front. We record the relationship and exclude fronts from source counts. Where we say "N brokers", N is distinct sources.
We have almost certainly not found them all. Fronts are identified by observing where a search redirects or which trackers a page links to, and that is a manual check we have run on some of the registry, not all of it. Treat the distinct-source count as an upper bound that will fall as more are found.
Shared suppression backends
Separately from fronts, several brands share one opt-out system. Observed directly by following each opt-out URL:
| Backend | Brands |
|---|---|
suppression.peopleconnect.us | Intelius, TruthFinder, InstantCheckmate, US Search |
/svc/optout/search/ | BeenVerified, Ownerly, NeighborWho |
spokeo.com/optout | Spokeo, AnyWho |
Nine brands, three requests. We report opt-out work as distinct actions, not brand count. Submitting to those nine sites and calling it nine removals would inflate our numbers in precisely the way we criticise.
It cuts in our favour too: one submission should clear every brand in the group, and because we verify presence independently on each domain, we can show whether it actually did. A group where the suppression succeeded on two brands and not the third is a finding worth publishing, and it is invisible to anyone counting submissions.
7. Reproducing our numbers
Every report is canonical JSON signed with Ed25519, carrying the URL and UTC timestamp of every individual scan behind the claim. To check one:
1. Verify the signature at /verify — it runs in your browser, offline, and needs nothing from us. 2. Read the evidence entries. Each states what was observed, when, and where. 3. Visit those URLs yourself.
A valid signature proves the report has not been altered since we issued it. It does not ask you to accept what the report says — that is what the URLs are for.
8. Known limitations
Stated here rather than discovered later.
- We only see public pages. A broker may hold data it does not publish. We do not report on what we cannot observe.
- Anti-bot outcomes are non-deterministic. The same broker can be readable in one run and challenged minutes later with nothing changed on our side. This is why retries exist and why
INCONCLUSIVEis a first-class state rather than an error. - Ambiguous identities stay ambiguous. Two people with the same name in the same state, with no age or city to separate them, will not resolve. We report that rather than guessing.
- Coverage is currently partial. Not every broker in our registry has a parser, and an unimplemented broker is recorded as
INCONCLUSIVE, never as clean. - Name shapes vary more than software expects. Taking the last token as a surname made "Robert Quillon Jr" resolve to a surname of "Jr", so every person with a generational suffix scored zero and was reported as not listed — a false absence attached to a whole class of people. Suffixes and compound-surname punctuation are now handled explicitly, but this is the category where we most expect to still be wrong.
- Age is often approximate. Some brokers publish a birth year rather than a date, so derived ages can be a year out. The ±2 year tolerance accounts for this.
- Over-caution is a failure too, and a quieter one. A guard that mistakes a genuinely empty result page for a broken parser sends every true absence to
INCONCLUSIVE. That is worse than it sounds: after a successful removal the page is empty by definition, soREMOVAL_VERIFIEDcould never be reached and a removal we actually achieved would be permanently unprovable. This happened to two brokers and was found by scanning a name that does not exist — a test no fixture could contain, since every fixture is a page that had results.
- Undetected interstitials are the failure mode we fear most. If an anti-bot page is not recognised as one, our parser finds nothing in it and the scan resolves to
NOT_FOUND— a positive claim of absence manufactured from a block page. This has happened: two brokers were serving pages titled exactly "Captcha", which our detection did not recognise, and the bug survived a 373-test suite because every test fixture is a page we had already fetched successfully. Only live scanning surfaced it. We now scan live routinely and review everyNOT_FOUNDas the assertion it is.
- This methodology can be wrong. If you can demonstrate an error, we will correct it and record the correction in the changelog below.
Changelog
| Version | Date | Change |
|---|---|---|
| 1.9 | 2026-08-08 | Reports state whether a removal request lies behind an absence, and count relistings even after they resolve. |
| 1.8 | 2026-08-08 | A name match that cannot be confirmed, from a subject who gave no distinguishing detail, is INCONCLUSIVE rather than absent. |
| 1.7 | 2026-08-08 | City and state comparison normalises abbreviations and spelled-out forms. |
| 1.6 | 2026-08-08 | Surname matching ignores generational suffixes and compound-surname punctuation. |
| 1.5 | 2026-08-08 | A non-match from a search that could not be narrowed to the subject's state is INCONCLUSIVE, not NOT_FOUND. |
| 1.4 | 2026-08-08 | Recorded the over-caution failure mode: an empty result page misread as parser drift would make verified removals unreachable. |
| 1.3 | 2026-08-08 | Recorded the undetected-interstitial failure mode and the live-scanning practice that catches it. |
| 1.2 | 2026-08-08 | Opt-out work is reported as distinct actions: brands sharing one suppression backend count once. |
| 1.1 | 2026-08-08 | Coverage is quoted against distinct sources: affiliate fronts that resell another broker's data are recorded and excluded from source counts. |
| 1.0 | 2026-08-07 | First published version. |