lucid.page Compare document IDs and recorded hashes in two CSV inventories
Text size
Read time2 min

# Compare document IDs and recorded hashes in two CSV inventories

Save this Python program as compare_inventories.py to compare two local CSV inventories. It reports which document IDs were added or removed and which have a different recorded hash.

Contents

# Prepare both inventories

Each CSV needs document_id and sha256 as exact header names. Extra columns are allowed. Every record needs a nonempty document ID and a 64-character hexadecimal hash. IDs must be unique within each file.

The program accepts uppercase and lowercase hexadecimal digits and stores hashes in lowercase before comparison. It preserves document IDs as supplied. Empty or duplicate header names and records with the wrong number of cells are rejected. Blank records are skipped. A header-only file is accepted as an empty inventory.

# Read the comparison

Run python3 compare_inventories.py before.csv after.csv. The report contains four fields: added, removed, changed and unchanged_count. The three lists are sorted by document ID. Changed lists IDs present in both files whose recorded hashes differ; unchanged_count counts IDs present in both with matching hashes.

The program reads the two inventory files. It does not open the listed documents or calculate their hashes. Both inventories are read before a report is printed. A valid comparison exits with status 0. File, CSV or value errors print an inventory error to standard error and return status 2.

# Review what changed

Use the report to locate differences between your supplied inventory records. It does not verify document authenticity, explain a change or determine whether a document meets a legal requirement. This utility provides no legal advice.

# Publisher

Settled Estate publishes this inventory-comparison utility. Keep its complete Python source and MIT license together when copying it.

Source record: Inventory comparison utility; publisher: Settled Estate; publication date: undated authored software, copyright 2026; source: complete original Python below.

# Complete source

python
#!/usr/bin/env python3
"""Compare two local CSV inventories without opening the listed documents."""

import argparse
import csv
import json
import re
import sys


def read_inventory(path):
    records = {}
    with open(path, encoding="utf-8-sig", newline="") as stream:
        rows = csv.reader(stream, strict=True)
        header = next(rows, None)
        if not header or any(not name for name in header):
            raise ValueError("missing or empty column name")
        if len(set(header)) != len(header):
            raise ValueError("duplicate column name")
        if not {"document_id", "sha256"}.issubset(header):
            raise ValueError("required columns: document_id, sha256")
        id_column = header.index("document_id")
        hash_column = header.index("sha256")
        for line, cells in enumerate(rows, start=2):
            if not cells:
                continue
            if len(cells) != len(header):
                raise ValueError("row %d has the wrong number of cells" % line)
            document_id, digest = cells[id_column], cells[hash_column]
            if not document_id:
                raise ValueError("row %d has an empty document_id" % line)
            if document_id in records:
                raise ValueError("row %d repeats a document_id" % line)
            if not re.fullmatch(r"[0-9a-fA-F]{64}", digest):
                raise ValueError("row %d requires a 64-character hexadecimal sha256" % line)
            records[document_id] = digest.lower()
    return records


def main():
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("before", help="earlier local inventory CSV")
    parser.add_argument("after", help="later local inventory CSV")
    args = parser.parse_args()
    try:
        before, after = read_inventory(args.before), read_inventory(args.after)
    except (OSError, ValueError, csv.Error) as error:
        print("inventory error: %s" % error, file=sys.stderr)
        return 2
    report = {
        "added": sorted(set(after) - set(before)),
        "removed": sorted(set(before) - set(after)),
        "changed": sorted(doc for doc in set(before) & set(after)
                          if before[doc] != after[doc]),
        "unchanged_count": sum(before[doc] == after[doc]
                               for doc in set(before) & set(after)),
    }
    print(json.dumps(report, ensure_ascii=False, indent=2))
    return 0


if __name__ == "__main__":
    sys.exit(main())

# License

text
MIT License

Copyright (c) 2026 Settled Estate Tools

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
End