Six months ago we migrated from HubSpot to Salesforce. The migration itself went fine. Data mapped correctly, custom fields transferred, nothing broke. We celebrated for about three days.
Then our sales team started complaining. "Why do i have two records for the same company?" "Why is this contact listed three times?" "I just called someone and they said another rep already reached out this morning."
We pulled a report. 40,000 duplicate contacts. Out of roughly 95,000 total records. More than 40% of our database was duplicates.
And the thing is, most of those duplicates already existed in HubSpot. We just hadnt noticed because HubSpot's dedup was handling some of it silently. When we moved to Salesforce, all the silent duplicates became visible and the mess that had been building for three years landed on our desk at once.
Why CRM migrations create duplicate nightmares
The duplicate problem in CRM migrations comes from multiple sources and they compound in ways that are hard to predict.
Pre-existing duplicates. Every CRM accumulates duplicates over time. Reps create new contacts instead of finding existing ones. Marketing imports lists that overlap with existing data. Web forms create new records even when the person already exists. According to estimated that poor data quality costs organizations an average of $12.9 million annually. For a company our size, duplicates alone were probably a six-figure problem.
The dedup approach that doesn't work
Our first attempt at fixing this was Salesforce's built-in duplicate management. You set up matching rules (match on email, match on name + company) and it flags potential duplicates.
The problem: it found about 8,000 duplicates based on exact email match. Thats helpful, but it missed the other 32,000 that had different emails, slightly different names, or variations in company names. Exact matching catches the easy duplicates and misses the hard ones.
Our second attempt was a manual review project. We assigned two ops people to go through flagged duplicates and merge them. After a week they had processed about 2,000 records and were losing their minds. At that rate, the project would take five months and cost more than just living with the duplicates.
Third attempt: we bought a Salesforce dedup app from the AppExchange. $200/month. It was better than the built-in tools but still relied heavily on exact matching. It caught maybe 60% of our duplicates. The other 40% (the ones with name variations, different emails, partial information) still required manual review.
Why fuzzy matching changes everything
The breakthrough came when we stopped trying to find exact matches and started looking for fuzzy matches with confidence scores.
Instead of asking "is this record identical to that record?" we asked "how similar are these records, and how confident are we that they represent the same entity?"
A fuzzy dedup approach looks at multiple fields simultaneously:
- Name similarity (using algorithms like Jaro-Winkler that can handle "Sarah Williams" matching "S. Williams")
- Company similarity ("Acme Corp" matching "Acme Corporation Inc")
- Phone number matching (ignoring formatting differences)
- Address similarity (handling abbreviations and format variations)
- Email domain matching (two records at @acmecorp.com are more likely to be from the same company)
Each field contributes to an overall confidence score. Two records might not match on any single field exactly, but when you combine name similarity of 85%, same company domain, and a phone number thats off by one digit, the confidence that theyre the same person is very high.
This is exactly the problem I built , 44% of companies estimate they lose over 10% of annual revenue due to poor CRM data quality. Duplicates are the single biggest contributor to that loss.
If you're planning a CRM migration, budget time for dedup. If you just finished one and the numbers look suspiciously high, pull a duplicate report. You might not like what you find, but you'll be glad you looked.
SOCIAL SHARE CARD GENERATOR