What we got wrong about US duplicate voter data
A short follow-up, and a correction to our own framing.
We recently asked whether the US has a voter data duplication problem. We downloaded data on 50 million voters from 7 states, ran it through the Tilores identity resolution engine, and found that 0.8% of registrations — nearly 400,000 — resolved to a person already registered somewhere else in our sample. In Ohio and Pennsylvania, the only two states in the sample that publish whether a registrant actually voted, we also found a small number of records where the same person appears to have voted in two places.
What we got wrong was the framing. When we found those duplicates, we assumed we were the first to measure the scale of this. We were not.

In the United States there is a non-profit, non-partisan organisation called “ERIC” (the Electronic Registration Information Center), which already does very much what we had done, working directly with member states to improve their voter registration lists.
Using entity resolution technology from our competitor, Senzing, ERIC has compared data from more than half of the US states. Importantly, ERIC also brings in additional data sources, such as the driving licence registers held by state DMVs, to provide an extra layer of validation — and does so in a privacy-preserving manner.
We had a call with them to compare methodologies. They generally catch the same duplicates we do, and had already detected the same double-vote cases. Let me be clear: they do good work with the data. But as with many things in the data world, the challenge is not only finding the problem — it is what happens downstream. Actually removing a duplicate registration is a decision for each individual state, and that sits outside ERIC’s control.
So the detection problem is, broadly, solved. That is genuinely good news, and it is the opposite of what we expected to find when we started.
The remaining issue is coverage. Over the past few years a number of states have withdrawn from the consortium. A state that leaves loses its cross-state view of its own roll — and because its records are no longer in the shared pool, the states that remain also become less able to identify duplicates involving that state. The effect is not contained to the state that leaves. Votebeat covers the debate in detail here, including the argument over whether identifying eligible-but-unregistered individuals belongs in the same programme as list maintenance. That is a legitimate question of programme design, and not one we have a view on.
Our view is a narrow and technical one: cross-state matching only works when states take part in it. Whatever the governance arrangements, a state that cannot compare its roll against other states is working with less information than a state that can.
Which is also why we think the two approaches are complementary rather than alternatives. A state can run its own entity resolution instance against its own roll — continuously, inside its own infrastructure, under its own control — and participate in a cross-state programme. The first tells you the quality of your own data today. The second tells you what you cannot see on your own.
If you work with voter registration data, or any other public register where duplicate records carry real consequences, get in touch — we are happy to talk through the methodology. US federal, state and local agencies can also reach us through Carahsoft.
See what resolved entity data does for your business — and your AI.