Data tools · Exact string comparison
Compare CSV Files by Key: Added, Removed, Changed
Match two CSV snapshots by one or more unique key columns. See added, removed, changed and unchanged records—even when rows or columns move—and download an audit.
Compare records by identity, not by line position
Paste your original CSV as Before and the new export as After, select a unique key, and compare. A reordered export does not make every record look changed. The tool finds records with the same exact key, checks the value columns you select, and separates additions, removals, changes and unchanged matches.
Use it for a small inventory snapshot, a synthetic customer export, a release configuration table or a migration check. It is a deterministic comparison, not an AI judgment about which dataset is correct. It does not merge files, update a database, infer a primary key or prove that an upstream export is complete. If you need a visual comparison of ordinary text rather than identified records, use the text diff checker.
A reproducible before-and-after example
The Load synthetic before / after button loads the exact public files below and explicitly selects id. The data is invented for this example; it is not a customer dataset.
Both snapshots contain four data records. After uses a different column order and moves an unchanged record to the first position. The fixture also includes a quoted comma in North, desk, escaped double quotes, a multiline note, a leading-zero ID, a long identifier and Unicode text.
| Key | Before data record | After data record | Result with all non-key columns selected |
|---|---|---|---|
0012 |
1 | 2 | Changed: status is open → closed |
0003 |
2 | — | Removed |
0004 |
3 | 3 | Unchanged |
9007199254740993123 |
4 | 1 | Unchanged; moving a record is not a value change |
0005 |
— | 4 | Added |
The expected summary is 1 added, 1 removed, 1 changed, 2 unchanged and 1 changed field. There are five audit entries because the union of the two snapshots contains five unique keys. The differences CSV has nine data rows: one changed field, four fields from the removed record and four fields from the added record. Its added note value =1+1 receives an apostrophe prefix; the JSON keeps the original string.
For the complete walkthrough, including composite keys and reconciliation checks, read how to compare two CSV files by key.
Choose a key you can defend
- Read both sets of columns. The tool validates each input and checks that their header sets match. It does not guess a key from the header name
id. - Select one or more key columns. An account ID may be sufficient. For order lines, an order number plus line number may be necessary. Every selected key field must contain non-whitespace text.
- Choose the fields to check for changes. Initially, every non-key column is included. Uncheck export timestamps or other fields only when you deliberately want to exclude them from changed/unchanged classification.
- Compare, inspect and download. Check the counts and original record numbers before using the audit elsewhere. Input, delimiter or selection changes disable previous downloads until you run a fresh comparison.
Compound keys retain their field boundaries: ['ab', 'c'] and ['a', 'bc'] are different identities. The implementation encodes an array of original strings rather than joining fields with a separator that might also occur inside a value.
Matching is case-sensitive and preserves spaces. A, a and A are different keys; a whitespace-only field is rejected. There is no trimming, fuzzy matching, date interpretation, numeric coercion or Unicode normalization. A changed key appears as one removal and one addition because the tool has no evidence that those identities refer to the same entity.
Repeated keys stop the whole comparison. The error gives the side and conflicting data record numbers without arbitrarily keeping the first or last match. Inspect the underlying export, choose a genuinely unique composite key or resolve duplicates deliberately with the CSV duplicate remover. Removing duplicates is a separate data decision, not a default repair for a failed comparison.
Understand what “unchanged” means
Unchanged means that the key matched and none of the selected comparison fields changed. It does not imply byte-for-byte file equality. Unselected values may differ, record positions may move, and the CSV may use different quoting or header order.
If you uncheck every comparison field, the run is explicitly membership-only. Every matched key is unchanged; only new or missing keys affect the counts. The JSON records this empty selection and still includes the original cells from both sides. Select all non-key columns when you need value-level reconciliation across the full table.
An empty cell and the text null are different strings. Two differently formatted numeric strings such as 1.0 and 1 are different. This avoids silently changing identifiers but also means that business-specific equivalence rules must be applied and documented upstream.
JSON audit versus differences CSV
| Output | Included | Important boundary |
|---|---|---|
| On-page table | First 20 audit records; selected changes | Long values and lists are shortened for display |
| JSON preview | First 50,000 characters at most | An excerpt may not be valid standalone JSON |
| Complete JSON copy/download | All audit records, both original cell arrays, settings, counts and field changes | Original field strings; capped at 4,000,000 characters |
| Differences CSV | One row per changed selected field; one row per original column for added/removed records | Omits unchanged records; formula-like text is prefixed |
The JSON’s columns array uses Before’s header order. Both before and after cell arrays within each audit record align to that array, even when the After file’s header order differs. Top-level metadata retains both original header orders. A missing side is null; an existing empty field is "". Records appear in Before order, followed by After-only additions in After order. Numbers refer to 1-based data records, excluding the header, not physical lines: a quoted multiline field still belongs to one record.
Differences CSV uses the fixed columns status,before_record,after_record,key,column,old,new. The key is a JSON-array string in one quoted CSV cell. Addition rows have an empty old cell; removal rows have an empty new cell. Use their status, or the richer JSON representation, to distinguish absence from an actual empty string.
The CSV export quotes every field, doubles embedded quotes, and uses comma separators and CRLF record endings. Formula-like exported cells receive a mandatory apostrophe prefix, including a column name when necessary. That intentionally changes affected CSV text; JSON does not receive the prefix. OWASP’s CSV injection guidance explains why spreadsheet software can interpret some cell text as a formula and why quoting alone is not a universal mitigation. Spreadsheet behavior varies: inspect import settings, treat identifiers as text and avoid enabling external content. The tool does not evaluate formulas or guarantee safety in every spreadsheet application.
Supported input and practical limits
The shared parser supports quoted commas, doubled quotes, embedded line breaks and an initial Unicode BOM. Each side may explicitly use comma, semicolon or tab; there is no automatic delimiter detection. It accepts CRLF, LF and CR record separators, but preserves line breaks contained in parsed fields. The input is pasted text, not a byte-preserving file archive; browser paste handling may normalize line endings before parsing.
The quoting model follows the common CSV conventions documented in RFC 4180. This tool adds stricter application requirements: headers are mandatory, non-empty and unique; every data record has the same field count; both files have the same exact header set. Unclosed quotes, unexpected characters after a closing quote, NUL characters and invalid Unicode are rejected rather than silently repaired. Internal blank records are not skipped.
- Per input: 1,000,000 JavaScript string characters, 10,000 data records, 200 columns, 100,000 cells including headers and 100,000 characters per field.
- Complete JSON: 4,000,000 characters. Repeated long keys or header names can expand a report beyond its input size. An oversized audit fails as a whole; there is no partial result to download.
- Differences CSV: 10,000 data rows, 100,000 characters per exported cell after prefixing and 4,000,000 total characters. Field-level expansion can exceed these limits even for a valid JSON audit. In that case the complete JSON remains available and the CSV download is explicitly disabled.
These are character and record limits, not megabyte promises. This is a browser tool for bounded samples, without background processing, automatic schema mapping, scheduling or server-side storage. For a large export, use a reproducible local data pipeline rather than assuming that a sampled comparison covers the entire dataset.
Frequently asked questions
Can I compare CSV files when rows and columns are reordered?
Yes. Records match by the selected unique key, not row position. Header order may differ, but both files must contain the same exact header names. The JSON preserves each side’s original data record numbers.
What happens if a selected key is blank or repeated?
The comparison stops and identifies the Before or After side and affected data record numbers. Every selected key field must be non-empty, and each key combination must be unique on its side. No first-match or fuzzy pairing is attempted.
Can I ignore changes in some columns?
Yes. Only selected non-key columns determine changed or unchanged status. Selecting none compares membership only: all matched keys are unchanged even if other fields differ. The complete original before and after cells still appear in the JSON audit.
Does the comparison preserve leading zeros and long IDs?
Yes. All parsed fields remain strings, so 0012 differs from 12 and long numeric-looking identifiers are not rounded. JSON retains original field text. Import differences CSV as text to avoid a spreadsheet applying its own number or date conversion.
What is included in the JSON and differences CSV downloads?
JSON includes settings, counts, all matched and unmatched records, original data record numbers, complete cell arrays and selected field changes. Differences CSV contains field-level rows for additions, removals and changes; it omits unchanged records and prefixes formula-like cells with an apostrophe. Exports are complete or explicitly unavailable, never silently truncated.
Is my CSV uploaded or saved by this tool?
The comparison code processes pasted CSV in this browser without uploading it or saving it in browser storage. Copy and download happen only when selected. Site analytics follow the separate privacy policy; use synthetic or non-sensitive samples.
Related workflows and source notes
Read the CSV comparison walkthrough for worked checks, compare text workflows for other comparison tasks, or convert a snapshot with CSV to JSON and JSON to CSV. If matching keys fail because of duplicates, inspect them with the CSV duplicate remover before deciding on a data correction.
Source check: September 5, 2026. RFC 4180 informs the quoting discussion; OWASP informs the spreadsheet warning. The matching rules, limits and example outputs above describe this implementation, not claims that these sources certify it. See the site privacy policy for analytics and contact information.