T. BHAIJI
Index SHEET 06 / 06

Offline datasheet comparison tool

Comparing a part against its competitors means pulling the same numbers out of vendor datasheets that agree on nothing, then arguing about them. I wrote a tool that reads the PDFs locally, lines the numbers up at the same operating point, and writes the comparison with every figure linked back to the page it came from.

VENDOR PDFsno two alikeLOCAL ONLYno datasheet leaves the machineextract1 023 linesunits172compare161valuepropdrafts the claimsBROWSER APP127.0.0.1CLIscriptableXLSXexportfixed rules · same PDF, same result
Data flow, and the boundary nothing crosses.
Sheet
06 of 06
Title
Offline datasheet comparison tool
Client
Personal tool
Local web app, nothing leaves the machine
Drawn
2026
Work done
Software engineering · PDF parsing · Local web app
Project type
Solo project
My role
Sole author: architecture, parsing, comparison, interface
Status
AS BUILTassembled, tested, in use
Python · pdfplumber · openpyxl · http.server · pytest

The comparison is mechanical, tedious, and easy to get wrong

Component selection kept coming up across my hardware projects, with MOSFET candidates here, gate drivers there and logic families somewhere else, and the work is always the same: open six PDFs, find the same parameter in six different table layouts, reconcile mV against V and mΩ against Ω, and retype it into a spreadsheet.

It is mechanical work, which means it is error-prone in a way that matters. A transcription slip in an RDS(on) column is how you pick the wrong part and find out during thermal testing.

Two numbers are also rarely comparable as printed. One vendor quotes a parameter at 25 °C and another at 85 °C, or at a different supply voltage, so lining them up honestly means knowing the operating point each figure was measured at.

The obvious fix would be to throw the PDFs at a language model. That fails twice here. Vendor datasheets are routinely under NDA, so uploading them is not a choice I get to make. And a comparison has to be reproducible. If the same PDF can produce different numbers on two runs, the output is not evidence.

Fixed rules, a local server, and a link back to every number

  1. The parser is where datasheet chaos is contained, 1,023 lines and the largest part of the project. It uses pdfplumber to pull text and tables, then applies fixed rules to locate parameters across inconsistent vendor layouts. Everything downstream receives normalised data and never sees a PDF.
  2. Unit handling is its own module because it is the part most likely to be silently wrong. Converting mΩ and Ω, mV and V, nC and pC into a canonical form is where a comparison tool earns or loses trust, and isolating it makes it directly testable.
  3. Every number links back to the page it came from. Click a cell and you get the sentence it was extracted from, with its page number. That is the part that makes the output usable in a design review, because the reviewer can check the claim rather than take it on faith.
  4. The interface is a local web app. A small standard-library HTTP server binds to 127.0.0.1 and opens the page in the browser, so the interface is HTML and CSS rather than a desktop toolkit, and the datasheets never leave the machine. The first version was a Tkinter window; the browser gave better tables and far better layout control for the same effort.
  5. On top of the comparison it drafts the value proposition for a chosen home-brand part against each competitor, so the output is an argument with sources attached rather than a grid of numbers. Results also export to Excel.
  6. Tests target extraction and unit translation, with a generator that builds synthetic datasheet PDFs. The parser is tested against known-good input with known-correct answers, rather than against real datasheets that cannot be committed to a repository.
Module layout, v0.2.0
extract.py1 023 linesPDF parsing and parameter location
server.py368local web app on 127.0.0.1
export.py262Excel and comparison output
valueprop.py238drafts the comparison claims
units.py172unit normalisation and conversion
compare.py161side-by-side comparison
library.py123part and category library
cli.py79terminal entry point
tests/618extraction and translation tests, sample generator

It does the job I built it for

Outcome
3 573lines of Python across nine modules plus tests
Loopbackonly. The server binds to 127.0.0.1 and nothing calls out
Every valuelinks back to its datasheet page
Deterministicfixed rules, so the same PDF always gives the same result

The tool is at v0.2.0 and does what I need: parameters out of vendor PDFs, normalised units, figures lined up at a common operating point, a drafted comparison against competing parts, and an Excel export. Dependencies stay deliberately thin, pdfplumber and openpyxl, with the web app built on the standard library, so installation is a pip command.

The honest limit is in the parsing. Fixed rules beat a model for reproducibility and privacy, but they only cover the layouts and parameter categories I have encoded. An unusual layout needs a rule, not a retrain. I wanted predictable and auditable over broad and occasionally wrong, but it does mean coverage grows by hand.

SCOPE AND LIMITATIONSv0.2.0, and a personal tool rather than a released product. Extraction covers the datasheet layouts and parameter categories encoded so far, so it is not a general-purpose datasheet parser. The drafted value proposition is a first draft for a human to edit, not finished copy.

Book a call

Happy to talk about a graduation internship, a vacancy, or a project you want a second opinion on. Twenty minutes is usually enough to work out whether it is worth a longer conversation.

Typical length
20 to 30 minutes
Time zone
Central European Time, Arnhem
Languages
English, Hindi
Usually free
Weekday evenings and most of the weekend

Looking for a graduation internship from February 2027.

Power electronics, embedded hardware, renewable energy or power systems. I am equally happy writing the firmware and tooling around them. Based in Arnhem, open to relocating in the Netherlands.

Based in
Arnhem, Netherlands
Available
Graduation internship from Feb 2027

© 2026 Tanishq Bhaiji · Arnhem tanishqbhaiji42@gmail.com LinkedIn CV (PDF)