How to Find Duplicate Files on Mac (Without Deleting the Wrong Copy)
Open your Downloads folder and look. Budget.pdf, Budget-1.pdf, Budget final.pdf, Budget final 2.pdf, Budget FINAL copy.pdf. Five files, maybe three real versions, and you have no idea which one you actually sent your accountant.
That's the duplicate problem in miniature, and it scales up ugly: the same photo exported twice, a song ripped from two sources, a 2 GB video you "saved a copy" of and forgot. None of it shows up as a problem in Finder, because Finder doesn't know two files are the same file — it only knows their names and sizes.
Here's how to find them for real, and the one distinction that keeps you from deleting something you'll regret.
Why "same name" and "same size" aren't good enough
The naive approach — sort by name, eyeball it — catches the obvious ones and misses everything interesting. Two files can have totally different names and identical contents. Two files can have the same name and be completely different documents. Name tells you nothing.
Size is a better filter but still a lie. Two 4.20 MB files are almost never the same file — that's just where file sizes tend to cluster. You'll waste an afternoon comparing unrelated PDFs that happen to match in bytes.
The only reliable answer is content: read the actual bytes of each file and check whether two files are identical. That's a checksum (a hash — usually MD5 or SHA), a short fingerprint generated from a file's contents. Same contents → same hash. Different contents → different hash, even by a single pixel.
The Terminal way (proof you can do it without any app)
If you want to feel the mechanics before trusting a tool, macOS ships everything you need:
cd ~/Downloads
md5 *.pdf | sort | uniq -w32 -D
md5 prints each file's hash, sort groups identical hashes together, and uniq -w32 -D shows only the lines where the first 32 characters (the whole MD5) appear more than once — i.e. true duplicates. It's not elegant, but it's honest, and it works right now.
The catch, and it's a real one: this only searches the folder you're in, matching exact duplicates only. It won't scan your whole disk, and it treats a slightly-edited copy (same photo, different filename and one re-save) as not-a-match. Useful for a folder, useless as a way to clean a machine.
The three traps that make this dangerous
Duplicate-hunting is one of the few cleanup tasks where being wrong actually hurts, because you're deleting files you do have — just in another location that mattered. Three traps:
- Duplicates inside app data. Your Photos library, iTunes/Music folder, and any app's
Application Supportdirectory contain files that look like loose duplicates but are wired into a database. Delete one "copy" and you corrupt the app. Never bulk-clean inside a library package. - The iCloud / synced-folder trap. A file that appears twice locally might be one synced item and one local copy. Deleting the wrong one can push a deletion up to your cloud and remove it from your other devices.
- "Similar" isn't "same." A tool that flags near-duplicates (different resolution, minor edit) is doing something genuinely different from exact-match deduping — and it needs your judgment on every result. Exact matches are safe to auto-group; fuzzy ones are not.
The rule I follow: exact duplicates outside app libraries are safe to clean aggressively. Everything else, review one at a time.
The tool way (and when it's worth it)
Doing a whole-disk content search by hand means walking every folder, hashing thousands of files, and cross-referencing — which is exactly the kind of tedious, error-prone grunt work a proper tool is for.
The Duplicates scan in CleanDiskGo does the content-hash pass across your disk, groups files that are byte-for-byte identical, and — the part that matters given the traps above — lets you review each group and keep one before deleting the rest. It also separates the photo dedupe case, because matching images is a different problem than matching files; if that's specifically what you're chasing, I wrote a dedicated guide on finding and deleting duplicate photos on Mac.
The reason to use a tool here isn't speed, exactly. It's that a content hash across a whole drive is genuinely slow to do by hand, and the "keep one, delete the rest" review is where a UI earns its keep over a uniq command.
What this is actually worth
Duplicates are rarely your biggest storage problem — that's usually one forgotten folder or a runaway cache, which a disk usage analysis will surface faster. Duplicate-hunting pays off most in specific places: a Downloads folder with years of re-downloads, a media folder, project directories with versioned copies.
So don't nuke your whole disk chasing dupes. Point the scan at the folders you know are messy, keep one of each group, and move on. The five Budget final files can finally become one.