Data tools · Browser calculation

CSV Duplicate Remover: Compare Selected Columns

Choose the columns that define a duplicate. Preview which records stay, then download kept and removed CSV separately. Your source input stays unchanged.

Up to 1,000,000 characters and 10,000 data records, with unique, non-empty column headers. Use synthetic or non-sensitive samples. CSV processing stays in this browser; site analytics are separate.

2. Select columns to compare

All selected columns must match. All columns are selected initially; the synthetic example selects only id. Unselected cell differences are ignored when grouping, not merged.

“Last” is not the latest timestamp. Single-occurrence mode removes every record in a duplicate group.

Off by default. A record with any empty selected key is kept separately. Matching options do not rewrite retained cell text.

Changes affected exported cells and headers intentionally. Quoting alone does not prevent spreadsheet formulas; no prefix works in every application. Check import settings.

Paste CSV and read its columns, or load the synthetic example.

    What is this CSV duplicate remover?

    This tool identifies repeated records, not repeated text lines. Select the fields that define a match, decide which occurrence to keep, and inspect a record-by-record preview. It accepts quoted delimiters, escaped quotes and line breaks inside a field using the site’s strict CSV parser. All values remain strings; 0012 is different from 12.

    Unlike a generic line cleaner, it can compare only an id column while retaining the selected record’s other fields. It does not merge the contents of different records or decide which version is correct. For plain one-item-per-line lists, use the line cleanup tool.

    Worked example: the key changes the answer

    Load the synthetic example above, or download the example CSV:

    id,name,status
    0012,Ada,open
    0012,Ada,closed
    0003,"North, desk",open
    0003,"North, desk",open
    ,Unknown A,open
    ,Unknown B,open
    

    Select only id and keep the default blank-key behavior:

    Mode Input data record numbers kept Removed
    First occurrence 1, 3, 5, 6 2, 4
    Last occurrence 2, 4, 5, 6 1, 3
    Single-occurrence groups only 5, 6 1, 2, 3, 4

    The two records with empty IDs remain separate. If you select all columns, the different status on the second record stops it matching the first; only the repeated 0003 record is removed. First/last results stay in their original input order. Record numbers exclude the header and are not physical line numbers when cells contain line breaks.

    Expected first-occurrence CSV · Expected removed CSV · Full walkthrough

    Why preview before removing duplicates?

    Two records with the same email, surname or product title might still represent separate people or events. A compound key such as order_id plus line_id is often more appropriate than one broad field. The tool follows your chosen rule; it does not validate business identity. Keep the original source and inspect both exports before replacing a working dataset.

    The decision log is a JSON download containing input data record numbers, keep/remove decisions, retained-record references and comparison settings. It does not add tracking columns to your CSV. The log includes selected column names; treat exported files according to your own data-handling requirements.

    Matching and export limits

    • Exact by default: whitespace, case, Unicode representation and empty text remain meaningful. Optional whitespace trimming and JavaScript lowercase conversion affect matching only. They are not fuzzy matching, accent removal or locale-aware name matching.
    • Empty keys: by default, any empty selected key prevents that record from grouping with another. Enable empty-key matching deliberately if blank really is an identity value for your task.
    • Strict input: unique non-empty headers, consistent field counts, explicit comma/semicolon/tab delimiter. Malformed quotes and inconsistent records are errors, not silently repaired.
    • Bounds: 1,000,000 input characters; 10,000 data records; 200 columns; 100,000 cells including headers; 100,000 characters per field; 4,000,000 characters per export. A smaller limit can be reached first.
    • No workbook or multi-file processing: paste a bounded text sample. No XLSX import, joining, scheduling, automatic date sorting or approximate identity resolution is included.
    • Text fidelity, not byte fidelity: retained field text stays unchanged except optional formula prefixes. Every exported field is quoted, exports use CRLF record endings, and textarea paste may normalize line breaks. The original CSV byte sequence and spreadsheet display formatting are not preserved.

    Spreadsheet import tips

    CSV quotes do not instruct a spreadsheet to preserve leading zeros or disable formula evaluation. Import identifier columns as text instead of relying on a double-click. Prefix protection is enabled by default and deliberately adds an apostrophe to formula-like text, including affected headers. The preview displays original input values, not those added prefixes. A prefix is not a universal guarantee across applications or save/reopen cycles; review the destination’s import behavior. See the formula-injection caveats.

    The algorithm does not upload pasted content or persist it in browser storage. This is not a claim that the entire website makes no requests: site analytics and privacy are separate. Use synthetic samples rather than credentials or private customer records.

    Frequently asked questions

    Can I remove CSV duplicates using just one column?

    Yes. Read the header, select only that key column, and preview. Differences in unselected columns do not stop a match; whole records are retained or removed, not merged.

    Does keep last select the newest date?

    No. First and last refer only to input record order. Sort and verify your records upstream if a timestamp or version determines which record should survive.

    Is my CSV uploaded?

    The deduplication code runs in this browser without uploading the pasted CSV or saving it in browser storage. Site analytics follow the separate privacy policy.

    Sources and next steps

    CSV parsing uses the common quoting conventions described by RFC 4180. Selected-key and first/last retention are also explicit concepts in the pandas duplicate-removal API; this tool is an independent browser implementation, not pandas running on a server. Sources checked September 5, 2026.

    Convert the kept CSV to JSON · Export JSON to CSV · Browse all browser tools

    Comparing two exports instead?

    Use the CSV snapshot comparison when the task is to identify additions, removals and changed fields between two exports. It aligns records by unique, non-empty keys rather than input position. Deduplication handles repeated keys within one file; it does not establish that a key identifies the same entity across files. Review the worked comparison example before treating a changed count as an audit result.

    📖 New to this? Read the full guide: Remove Duplicate CSV Rows Without Losing the Wrong Record →