Data tools · Browser calculation
CSV Duplicate Remover: Compare Selected Columns
Choose the columns that define a duplicate. Preview which records stay, then download kept and removed CSV separately. Your source input stays unchanged.
What is this CSV duplicate remover?
This tool identifies repeated records, not repeated text lines. Select the fields that define a match, decide which occurrence to keep, and inspect a record-by-record preview. It accepts quoted delimiters, escaped quotes and line breaks inside a field using the site’s strict CSV parser. All values remain strings; 0012 is different from 12.
Unlike a generic line cleaner, it can compare only an id column while retaining the selected record’s other fields. It does not merge the contents of different records or decide which version is correct. For plain one-item-per-line lists, use the line cleanup tool.
Worked example: the key changes the answer
Load the synthetic example above, or download the example CSV:
id,name,status
0012,Ada,open
0012,Ada,closed
0003,"North, desk",open
0003,"North, desk",open
,Unknown A,open
,Unknown B,open
Select only id and keep the default blank-key behavior:
| Mode | Input data record numbers kept | Removed |
|---|---|---|
| First occurrence | 1, 3, 5, 6 | 2, 4 |
| Last occurrence | 2, 4, 5, 6 | 1, 3 |
| Single-occurrence groups only | 5, 6 | 1, 2, 3, 4 |
The two records with empty IDs remain separate. If you select all columns, the different status on the second record stops it matching the first; only the repeated 0003 record is removed. First/last results stay in their original input order. Record numbers exclude the header and are not physical line numbers when cells contain line breaks.
Expected first-occurrence CSV · Expected removed CSV · Full walkthrough
Why preview before removing duplicates?
Two records with the same email, surname or product title might still represent separate people or events. A compound key such as order_id plus line_id is often more appropriate than one broad field. The tool follows your chosen rule; it does not validate business identity. Keep the original source and inspect both exports before replacing a working dataset.
The decision log is a JSON download containing input data record numbers, keep/remove decisions, retained-record references and comparison settings. It does not add tracking columns to your CSV. The log includes selected column names; treat exported files according to your own data-handling requirements.
Matching and export limits
- Exact by default: whitespace, case, Unicode representation and empty text remain meaningful. Optional whitespace trimming and JavaScript lowercase conversion affect matching only. They are not fuzzy matching, accent removal or locale-aware name matching.
- Empty keys: by default, any empty selected key prevents that record from grouping with another. Enable empty-key matching deliberately if blank really is an identity value for your task.
- Strict input: unique non-empty headers, consistent field counts, explicit comma/semicolon/tab delimiter. Malformed quotes and inconsistent records are errors, not silently repaired.
- Bounds: 1,000,000 input characters; 10,000 data records; 200 columns; 100,000 cells including headers; 100,000 characters per field; 4,000,000 characters per export. A smaller limit can be reached first.
- No workbook or multi-file processing: paste a bounded text sample. No XLSX import, joining, scheduling, automatic date sorting or approximate identity resolution is included.
- Text fidelity, not byte fidelity: retained field text stays unchanged except optional formula prefixes. Every exported field is quoted, exports use CRLF record endings, and textarea paste may normalize line breaks. The original CSV byte sequence and spreadsheet display formatting are not preserved.
Spreadsheet import tips
CSV quotes do not instruct a spreadsheet to preserve leading zeros or disable formula evaluation. Import identifier columns as text instead of relying on a double-click. Prefix protection is enabled by default and deliberately adds an apostrophe to formula-like text, including affected headers. The preview displays original input values, not those added prefixes. A prefix is not a universal guarantee across applications or save/reopen cycles; review the destination’s import behavior. See the formula-injection caveats.
The algorithm does not upload pasted content or persist it in browser storage. This is not a claim that the entire website makes no requests: site analytics and privacy are separate. Use synthetic samples rather than credentials or private customer records.
Frequently asked questions
Can I remove CSV duplicates using just one column?
Yes. Read the header, select only that key column, and preview. Differences in unselected columns do not stop a match; whole records are retained or removed, not merged.
Does keep last select the newest date?
No. First and last refer only to input record order. Sort and verify your records upstream if a timestamp or version determines which record should survive.
Is my CSV uploaded?
The deduplication code runs in this browser without uploading the pasted CSV or saving it in browser storage. Site analytics follow the separate privacy policy.
Sources and next steps
CSV parsing uses the common quoting conventions described by RFC 4180. Selected-key and first/last retention are also explicit concepts in the pandas duplicate-removal API; this tool is an independent browser implementation, not pandas running on a server. Sources checked September 5, 2026.
Convert the kept CSV to JSON · Export JSON to CSV · Browse all browser tools
Comparing two exports instead?
Use the CSV snapshot comparison when the task is to identify additions, removals and changed fields between two exports. It aligns records by unique, non-empty keys rather than input position. Deduplication handles repeated keys within one file; it does not establish that a key identifies the same entity across files. Review the worked comparison example before treating a changed count as an audit result.