💻 Tilores Studio is now available. Run entity resolution locally on your machine.Download free

For US government agencies: Tilores is available to federal, state & local agencies through our partnership with Carahsoft

Does the US have a
duplicate voter problem?

Yes — but a small one, and it is a data quality problem rather than a fraud problem.

We ran entity resolution across 49.5 million publicly available voter records from 7 states and found 394,396 duplicate registrations — 0.8% of the analysed population, with no meaningful difference between the parties. Here is what the data shows.

Download the analysis as a PDF

The Dataset

49.5 million records. One identity resolution engine.

49.5M
Voter profiles

Across 7 US states

7
States analysed

AR, FL, GA, MI, NC, OH, PA

394,396
Duplicate registrations

0.8% of the analysed population

0.96 v 0.89
Democrat v Republican

Duplicate rate by party — no meaningful skew

Source data: public voter registration files purchased between September and November 2023 under state open records laws (Michigan $23, Pennsylvania $20, Georgia $250). Analysis published 2024. Matching uses text similarity, geographic proximity, and temporal range matching.


The Method

How entity resolution finds voter duplicates

Unresolved — Raw Records
JAMES R. WILSON
Georgia · DOB 1974-03-12 · Active
James Wilson
Florida · DOB 1974-03-12 · Active
Jim R Wilson
Georgia · DOB 03/12/1974 · Inactive
3 separate records across 2 states — no link between them
Resolved — Unified Entity
JW
James R. Wilson
entity_vt_4k9nQr · score 94
StatesGeorgia, Florida
Registration statusActive + Inactive
Name variants3 recorded
Source records3 linked

State Breakdown

Where duplicates are most concentrated

State
Duplicate registrations
Duplicate rate
Florida
148,516
1.10%
Pennsylvania
80,142
1.01%
Georgia
51,876
0.73%
Ohio
47,000
0.79%
Michigan
35,723
0.47%
North Carolina
20,323
0.32%
Arkansas
10,816
0.92%
Total
394,396
0.80%

Every figure in this table is a duplicate registration rate — the share of active registrations on that state's roll belonging to a person with two or more active registrations, whether in the same county, another county in the same state, or another state in our sample. Rates vary widely below the state level: Arkansas averages 0.92%, but Searcy County reaches 2.06% (121 of 5,805 voters). In Florida, Hardee County reaches 2.3% (314 of 13,654) and Osceola County 1.94% (5,376 of 214,588).

A separate signal: registration rates

Duplicate rates are one measure. Registration rates — the share of a county's population that appears on the roll at all — are a different one, and two states stood out at opposite ends.

Michigan — above 90%

Several counties show registration rates above 90% of the population. That is not feasible when roughly 22% of the US population is under 18, which suggests deceased registrants are not being removed from the rolls as efficiently as in other states.

Arkansas — as low as 26%

At the other end, some Arkansas counties show unusually low registration rates — Lincoln County is 26.02%. That points to room for improvement in voter engagement rather than to a data quality fault.


Party Affiliation

Duplicates by registered party

State
Democrat
Republican
Florida
1.30%
58,455 records
1.29%
65,905 records
Pennsylvania
1.12%
39,797 records
0.96%
30,928 records
Arkansas
1.10%
632 records
1.23%
1,149 records
Ohio
0.68%
6,686 records
0.67%
9,061 records
Georgia
0.58%
8,713 records
0.43%
7,449 records
North Carolina
0.37%
7,702 records
0.29%
5,792 records
Overall
0.96%
121,985 records
0.89%
120,384 records

Rates are the share of registrations affiliated with each party that are duplicates. Michigan is excluded because it does not disclose party affiliation; voters registered with other parties or with no affiliation are excluded from this table. Overall the gap between the two parties is 0.07 percentage points, and every individual state gap is under 0.16 points. The direction is not uniform — Republican-affiliated records show the higher rate in Arkansas, Democrat-affiliated records in most other states. Gaps this small, in inconsistent directions, do not support a partisan reading in either direction.


Double Voting

What we could — and could not — test

A duplicate registration is not a duplicate vote. To test whether someone actually voted twice, you need the roll to record whether each registrant voted — and of the seven states in our sample, only Ohio and Pennsylvania publish that. Those are the only two states where we could ask the question. In the other five the data is not publicly available — the state holds it, but it is not published, so an outside analysis cannot reach it.

Within that subset we found roughly 1,000 potential cases of the same person voting twice. Because we wanted to be stringent, we then applied additional deduplication rules and reviewed the remaining cases manually one by one, narrowing them to 61 cases we were confident about.

~1,000
Potential cases from matching
61
After stricter rules + manual review
Democratic Party affiliation 31
Republican Party affiliation 21
Both parties across the two records 6
No party affiliation 3
What these cases look like

An individual registered to vote in West Chester, Pennsylvania in October 2020, shortly before the presidential election. Nine days later the same individual registered in Philadelphia. The name differed slightly; the date of birth and phone number were identical. Both registrations voted.

An individual in Normalville, Pennsylvania registered twice at the same address, 16 years apart almost to the day. The only difference between the two records was a minor discrepancy in the date of birth. Both registrations voted. This was the most common pattern — a small change to a date of birth, with name and address unchanged.

Two of the 61 span Pennsylvania and Ohio, both involving absentee votes; the rest sit within a single state. Set against the millions of registrations in those two states this is a vanishingly small number — but it was detectable, and it was only detectable because the vote history was published. Had the other five states published theirs, we would expect to find further cross-state cases.

These remain a candidate set for review by election officials rather than a legal finding. Entity resolution, followed by manual review, can establish that two records are very probably the same person; establishing that an offence occurred is a matter for the relevant authorities.


Key Finding

Why this is a data quality problem, not a conspiracy

Voters move states and counties — old registrations are rarely cancelled in time

Name variations across records (Jim vs James, hyphenated surnames) prevent simple deduplication

Death registries are not synchronised in real time, leaving deceased registrants on the rolls

Cross-state matching needs fuzzy matching — exact-match rules miss most genuine duplicates

The two double-voting cases we describe above illustrate the pattern: in each, the duplicate survived because a single field had drifted — a slightly different spelling of a name, or a date of birth out by a few days. Exact-match rules treat those records as different people. That is the whole problem in one line.


Context

What already exists — and where Tilores fits

We were not the first to run this analysis. A non-profit organisation called ERIC — the Electronic Registration Information Center — already performs cross-state list comparison for its member states using entity resolution, and supplements voter rolls with additional sources such as DMV records. When we compared methodologies with them, they were finding broadly the same duplicates we were, including the double-vote cases.

Two honest conclusions follow. First, detection is not the bottleneck — acting on the results is, and that sits with each state rather than with ERIC. Second, coverage is uneven: a number of states have left the consortium in recent years, which reduces the quality of their own lists and weakens cross-state matching for the states that remain.

These are complementary, not competing

A state can run its own Tilores instance and participate in ERIC. They answer different questions:

Your own Tilores instance

Continuous resolution of your own roll, inside your own AWS account, under your own control. Gives you a live view of your data quality and something to act on between reports — plus the same engine for every other identity dataset your agency holds.

ERIC membership

The cross-state picture that no single state can see on its own, with data-sharing agreements and supplementary sources already in place. Nothing a state-level instance can replicate alone.

Important
A state running this on its own data would get substantially better results

Everything above was produced from the fields that appear in public voter files — name, address, date of birth, registration status, party affiliation. That is a deliberately limited view, because a public showcase can only use public data.

A state election office holds considerably more about each registrant: driver's licence numbers, full or partial Social Security numbers, prior addresses, application and signature history, and links into DMV and vital records systems. Those additional identifiers improve matching in both directions at once — they confirm genuine duplicates that public data alone cannot resolve, and they rule out false positives where two different people share a name and a rough date of birth.

Vote history is the clearest example. Five of the seven states we analysed hold it but do not publish it, which is the only reason we could not test those states for double voting. Each of them could run that test on its own data tomorrow.

The practical implication: the figures on this page should be read as a floor, not a ceiling. A state running the same engine against its own complete records would find duplicates we could not see from the outside, and would do so with materially higher confidence in every match.


What Good Looks Like

What real-time entity resolution would change

Cross-state deduplication

Catch the same voter registered in two states when they move, not years later.

Real-time death registry sync

Remove deceased registrants from rolls automatically as records update, using fuzzy rather than exact matching.

Fuzzy matching at scale

Resolve name variants, typos, and format differences across 50M+ records in hours.


For Government Agencies

Voter rolls are the public example.
The problem is everywhere.

We used voter data because it is one of the few large-scale, publicly available identity datasets in the United States — not because electoral rolls are the largest instance of the problem. The same fragmentation shows up wherever an agency has been collecting records on people and organisations across separate systems and many years.

Improper payments & duplicate enrolment

The same person enrolled more than once across programmes, agencies, or states under name and address variants that exact-match rules never catch.

Student records across institutions

Students move between schools, districts, and states, re-enrolling each time under a slightly different name, address, or identifier. Resolving those into one learner record is what makes longitudinal outcomes, funding allocation, and support programmes accurate — and it is the same matching problem as a voter who moves.

Beneficiary & case file consolidation

One person's records spread across systems that were never designed to share a common key — reassembled into a single view without merging the source systems.

Licensing, permits & DMV records

Multiple licence or permit records for one individual, created over decades of different data entry standards and formats.

Vendor, grant & procurement integrity

Resolve vendor entities to surface shell companies, related parties, and sanctioned entities across procurement and grant systems.

Tax & revenue records

One taxpayer appearing as several across years, filing types, and associated entities — with the links between them never made explicit.


Deployment & Security

Where the data lives

Runs inside your environment

Deploy into your own AWS account, including AWS GovCloud. Your records stay inside your own boundary — we do not need a copy.

On-premise and air-gapped

The engine depends only on a key-value store, a queue, and file storage, so it runs on-premise and inside isolated networks.

SOC 2

Tilores is SOC 2 audited. Deployment, access, and retention questions are answered in detail during evaluation.

Resolution without pooling raw data

Cross-agency matching happens at the identity layer rather than by centralising records, so agencies keep custody of their own data.


Procurement

How US agencies buy Tilores

Tilores is available to federal, state and local government agencies through our partnership with Carahsoft. If this looks relevant to your agency, the quickest route is a short technical demo against a dataset you care about — we will bring Carahsoft in for contract vehicle and procurement questions.

Circulating this internally? Download the full analysis as a PDF — designed to be forwarded to colleagues and procurement teams.


Entity resolution works on any dataset — including yours

We ran this against 50 million public records. The interesting version of this exercise is the one run against your own.