Offline datasheet comparison tool
Comparing a part against its competitors means pulling the same numbers out of vendor datasheets that agree on nothing, then arguing about them. I wrote a tool that reads the PDFs locally, lines the numbers up at the same operating point, and writes the comparison with every figure linked back to the page it came from.
- Sheet
- 06 of 06
- Title
- Offline datasheet comparison tool
- Client
- Personal tool
Local web app, nothing leaves the machine - Drawn
- 2026
- Work done
- Software engineering · PDF parsing · Local web app
- Project type
- Solo project
- My role
- Sole author: architecture, parsing, comparison, interface
- Status
- AS BUILTassembled, tested, in use
The comparison is mechanical, tedious, and easy to get wrong
Component selection kept coming up across my hardware projects, with MOSFET candidates here, gate drivers there and logic families somewhere else, and the work is always the same: open six PDFs, find the same parameter in six different table layouts, reconcile mV against V and mΩ against Ω, and retype it into a spreadsheet.
It is mechanical work, which means it is error-prone in a way that matters. A transcription slip in an RDS(on) column is how you pick the wrong part and find out during thermal testing.
Two numbers are also rarely comparable as printed. One vendor quotes a parameter at 25 °C and another at 85 °C, or at a different supply voltage, so lining them up honestly means knowing the operating point each figure was measured at.
The obvious fix would be to throw the PDFs at a language model. That fails twice here. Vendor datasheets are routinely under NDA, so uploading them is not a choice I get to make. And a comparison has to be reproducible. If the same PDF can produce different numbers on two runs, the output is not evidence.
Fixed rules, a local server, and a link back to every number
- The parser is where datasheet chaos is contained, 1,023 lines and the largest part of the project. It uses pdfplumber to pull text and tables, then applies fixed rules to locate parameters across inconsistent vendor layouts. Everything downstream receives normalised data and never sees a PDF.
- Unit handling is its own module because it is the part most likely to be silently wrong. Converting mΩ and Ω, mV and V, nC and pC into a canonical form is where a comparison tool earns or loses trust, and isolating it makes it directly testable.
- Every number links back to the page it came from. Click a cell and you get the sentence it was extracted from, with its page number. That is the part that makes the output usable in a design review, because the reviewer can check the claim rather than take it on faith.
- The interface is a local web app. A small standard-library HTTP server binds to 127.0.0.1 and opens the page in the browser, so the interface is HTML and CSS rather than a desktop toolkit, and the datasheets never leave the machine. The first version was a Tkinter window; the browser gave better tables and far better layout control for the same effort.
- On top of the comparison it drafts the value proposition for a chosen home-brand part against each competitor, so the output is an argument with sources attached rather than a grid of numbers. Results also export to Excel.
- Tests target extraction and unit translation, with a generator that builds synthetic datasheet PDFs. The parser is tested against known-good input with known-correct answers, rather than against real datasheets that cannot be committed to a repository.
| extract.py | 1 023 lines | PDF parsing and parameter location |
|---|---|---|
| server.py | 368 | local web app on 127.0.0.1 |
| export.py | 262 | Excel and comparison output |
| valueprop.py | 238 | drafts the comparison claims |
| units.py | 172 | unit normalisation and conversion |
| compare.py | 161 | side-by-side comparison |
| library.py | 123 | part and category library |
| cli.py | 79 | terminal entry point |
| tests/ | 618 | extraction and translation tests, sample generator |
It does the job I built it for
| 3 573 | lines of Python across nine modules plus tests |
|---|---|
| Loopback | only. The server binds to 127.0.0.1 and nothing calls out |
| Every value | links back to its datasheet page |
| Deterministic | fixed rules, so the same PDF always gives the same result |
The tool is at v0.2.0 and does what I need: parameters out of vendor PDFs, normalised units, figures lined up at a common operating point, a drafted comparison against competing parts, and an Excel export. Dependencies stay deliberately thin, pdfplumber and openpyxl, with the web app built on the standard library, so installation is a pip command.
The honest limit is in the parsing. Fixed rules beat a model for reproducibility and privacy, but they only cover the layouts and parameter categories I have encoded. An unusual layout needs a rule, not a retrain. I wanted predictable and auditable over broad and occasionally wrong, but it does mean coverage grows by hand.
Book a call
Happy to talk about a graduation internship, a vacancy, or a project you want a second opinion on. Twenty minutes is usually enough to work out whether it is worth a longer conversation.
- Typical length
- 20 to 30 minutes
- Time zone
- Central European Time, Arnhem
- Languages
- English, Hindi
- Usually free
- Weekday evenings and most of the weekend
Looking for a graduation internship from February 2027.
Power electronics, embedded hardware, renewable energy or power systems. I am equally happy writing the firmware and tooling around them. Based in Arnhem, open to relocating in the Netherlands.
- Phone
- +31 6 8515 8402
- linkedin.com/in/tanishq-bhaiji
- Based in
- Arnhem, Netherlands
- Available
- Graduation internship from Feb 2027