Does the US have a
duplicate voter problem?
Yes — but a small one, and it is a data quality problem rather than a fraud problem.
We ran entity resolution across 49.5 million publicly available voter records from 7 states and found 394,396 duplicate registrations — 0.8% of the analysed population, with no meaningful difference between the parties. Here is what the data shows.
Download the analysis as a PDF49.5 million records. One identity resolution engine.
Across 7 US states
AR, FL, GA, MI, NC, OH, PA
0.8% of the analysed population
Duplicate rate by party — no meaningful skew
Source data: public voter registration files purchased between September and November 2023 under state open records laws (Michigan $23, Pennsylvania $20, Georgia $250). Analysis published 2024. Matching uses text similarity, geographic proximity, and temporal range matching.
How entity resolution finds voter duplicates
Where duplicates are most concentrated
Every figure in this table is a duplicate registration rate — the share of active registrations on that state's roll belonging to a person with two or more active registrations, whether in the same county, another county in the same state, or another state in our sample. Rates vary widely below the state level: Arkansas averages 0.92%, but Searcy County reaches 2.06% (121 of 5,805 voters). In Florida, Hardee County reaches 2.3% (314 of 13,654) and Osceola County 1.94% (5,376 of 214,588).
Duplicate rates are one measure. Registration rates — the share of a county's population that appears on the roll at all — are a different one, and two states stood out at opposite ends.
Several counties show registration rates above 90% of the population. That is not feasible when roughly 22% of the US population is under 18, which suggests deceased registrants are not being removed from the rolls as efficiently as in other states.
At the other end, some Arkansas counties show unusually low registration rates — Lincoln County is 26.02%. That points to room for improvement in voter engagement rather than to a data quality fault.
Duplicates by registered party
Rates are the share of registrations affiliated with each party that are duplicates. Michigan is excluded because it does not disclose party affiliation; voters registered with other parties or with no affiliation are excluded from this table. Overall the gap between the two parties is 0.07 percentage points, and every individual state gap is under 0.16 points. The direction is not uniform — Republican-affiliated records show the higher rate in Arkansas, Democrat-affiliated records in most other states. Gaps this small, in inconsistent directions, do not support a partisan reading in either direction.
What we could — and could not — test
A duplicate registration is not a duplicate vote. To test whether someone actually voted twice, you need the roll to record whether each registrant voted — and of the seven states in our sample, only Ohio and Pennsylvania publish that. Those are the only two states where we could ask the question. In the other five the data is not publicly available — the state holds it, but it is not published, so an outside analysis cannot reach it.
Within that subset we found roughly 1,000 potential cases of the same person voting twice. Because we wanted to be stringent, we then applied additional deduplication rules and reviewed the remaining cases manually one by one, narrowing them to 61 cases we were confident about.
An individual registered to vote in West Chester, Pennsylvania in October 2020, shortly before the presidential election. Nine days later the same individual registered in Philadelphia. The name differed slightly; the date of birth and phone number were identical. Both registrations voted.
An individual in Normalville, Pennsylvania registered twice at the same address, 16 years apart almost to the day. The only difference between the two records was a minor discrepancy in the date of birth. Both registrations voted. This was the most common pattern — a small change to a date of birth, with name and address unchanged.
Two of the 61 span Pennsylvania and Ohio, both involving absentee votes; the rest sit within a single state. Set against the millions of registrations in those two states this is a vanishingly small number — but it was detectable, and it was only detectable because the vote history was published. Had the other five states published theirs, we would expect to find further cross-state cases.
These remain a candidate set for review by election officials rather than a legal finding. Entity resolution, followed by manual review, can establish that two records are very probably the same person; establishing that an offence occurred is a matter for the relevant authorities.
Why this is a data quality problem, not a conspiracy
Voters move states and counties — old registrations are rarely cancelled in time
Name variations across records (Jim vs James, hyphenated surnames) prevent simple deduplication
Death registries are not synchronised in real time, leaving deceased registrants on the rolls
Cross-state matching needs fuzzy matching — exact-match rules miss most genuine duplicates
The two double-voting cases we describe above illustrate the pattern: in each, the duplicate survived because a single field had drifted — a slightly different spelling of a name, or a date of birth out by a few days. Exact-match rules treat those records as different people. That is the whole problem in one line.
What already exists — and where Tilores fits
We were not the first to run this analysis. A non-profit organisation called ERIC — the Electronic Registration Information Center — already performs cross-state list comparison for its member states using entity resolution, and supplements voter rolls with additional sources such as DMV records. When we compared methodologies with them, they were finding broadly the same duplicates we were, including the double-vote cases.
Two honest conclusions follow. First, detection is not the bottleneck — acting on the results is, and that sits with each state rather than with ERIC. Second, coverage is uneven: a number of states have left the consortium in recent years, which reduces the quality of their own lists and weakens cross-state matching for the states that remain.
A state can run its own Tilores instance and participate in ERIC. They answer different questions:
Continuous resolution of your own roll, inside your own AWS account, under your own control. Gives you a live view of your data quality and something to act on between reports — plus the same engine for every other identity dataset your agency holds.
The cross-state picture that no single state can see on its own, with data-sharing agreements and supplementary sources already in place. Nothing a state-level instance can replicate alone.
Everything above was produced from the fields that appear in public voter files — name, address, date of birth, registration status, party affiliation. That is a deliberately limited view, because a public showcase can only use public data.
A state election office holds considerably more about each registrant: driver's licence numbers, full or partial Social Security numbers, prior addresses, application and signature history, and links into DMV and vital records systems. Those additional identifiers improve matching in both directions at once — they confirm genuine duplicates that public data alone cannot resolve, and they rule out false positives where two different people share a name and a rough date of birth.
Vote history is the clearest example. Five of the seven states we analysed hold it but do not publish it, which is the only reason we could not test those states for double voting. Each of them could run that test on its own data tomorrow.
The practical implication: the figures on this page should be read as a floor, not a ceiling. A state running the same engine against its own complete records would find duplicates we could not see from the outside, and would do so with materially higher confidence in every match.
What real-time entity resolution would change
Catch the same voter registered in two states when they move, not years later.
Remove deceased registrants from rolls automatically as records update, using fuzzy rather than exact matching.
Resolve name variants, typos, and format differences across 50M+ records in hours.
Voter rolls are the public example.
The problem is everywhere.
We used voter data because it is one of the few large-scale, publicly available identity datasets in the United States — not because electoral rolls are the largest instance of the problem. The same fragmentation shows up wherever an agency has been collecting records on people and organisations across separate systems and many years.
The same person enrolled more than once across programmes, agencies, or states under name and address variants that exact-match rules never catch.
Students move between schools, districts, and states, re-enrolling each time under a slightly different name, address, or identifier. Resolving those into one learner record is what makes longitudinal outcomes, funding allocation, and support programmes accurate — and it is the same matching problem as a voter who moves.
One person's records spread across systems that were never designed to share a common key — reassembled into a single view without merging the source systems.
Multiple licence or permit records for one individual, created over decades of different data entry standards and formats.
Resolve vendor entities to surface shell companies, related parties, and sanctioned entities across procurement and grant systems.
One taxpayer appearing as several across years, filing types, and associated entities — with the links between them never made explicit.
Where the data lives
Deploy into your own AWS account, including AWS GovCloud. Your records stay inside your own boundary — we do not need a copy.
The engine depends only on a key-value store, a queue, and file storage, so it runs on-premise and inside isolated networks.
Tilores is SOC 2 audited. Deployment, access, and retention questions are answered in detail during evaluation.
Cross-agency matching happens at the identity layer rather than by centralising records, so agencies keep custody of their own data.
How US agencies buy Tilores
Tilores is available to federal, state and local government agencies through our partnership with Carahsoft. If this looks relevant to your agency, the quickest route is a short technical demo against a dataset you care about — we will bring Carahsoft in for contract vehicle and procurement questions.
Circulating this internally? Download the full analysis as a PDF — designed to be forwarded to colleagues and procurement teams.
Entity resolution works on any dataset — including yours
We ran this against 50 million public records. The interesting version of this exercise is the one run against your own.