Why “can Claude read PDFs?” is the wrong first question
People type “claude code read pdf” because they need an answer from a datasheet, not a lecture about file formats. Claude Code can work with files in a repository. A 68-page vendor PDF with a pinout on page 3 is much easier to check when the table has been turned into spec.md.
Cursor users hit the same wall: a PDF pane is not the same as a table in the model’s working set. The useful sequence is straightforward: parse the excerpt, commit the Markdown, mention the file, and check the result against the source.
Put the datasheet in the repo as Markdown
A chip datasheet is a test of the pipeline: absolute-maximum table, pin function table, package drawing, units in microamps. The demo uses a short excerpt of the TI LM358 family (pages 1–3 only, vendor copyright, not the full manual). Keep any source you use according to its license. After parsing, feature bullets and application lists should still be lists, and units should remain math.
Name the file after the part (lm358.md) and keep the excerpt small. Do not check in a manufacturer’s entire PDF if your license does not allow it. Keep the original PDF with the Markdown so a reviewer can check the source.
From Claude Code: @lm358.md then “list supply range and packages as a table.” From Cursor: same file, same ask. Codex and other repo agents follow the same rule — they need bytes they can tokenize as text.
Demo: pin tables and units
Left: the excerpt PDF. Right: the parsed Markdown, including feature bullets and application lists. Image hotlinks in the parse output point at KolmoPDF’s figure store when the parser emitted them.
This is not a replacement for SPICE or for TI’s own HTML. It is the missing step before an agent writes a comment that cites a current limit.
Keep the source beside the parsed file
This page is for developers who need a spec they can check in and cite. It is not a lab-meeting guide and not a medical journal club. Students presenting Attention Is All You Need belong to a different workflow.
If your PDF already has a clean text layer and simple paragraphs, you may not need this step. The gap is tables, scans, and equations that must survive into the repo.
A week in a repo that actually cites the datasheet
Monday: someone drops a vendor PDF in Slack with “can we support this part?” Tuesday: an intern pastes a screenshot into ChatGPT and gets a confident current limit that belongs to a different suffix of the family. Wednesday: Claude Code writes a driver comment that cannot be grepped against the table because the table never entered git.
The fix is not another chat window. The fix is a file. Parse the excerpt. Name it after the orderable device. Open a pull request whose only job is “add lm358.md.” Then the agent conversation becomes boring in the useful way: “according to lm358.md, BA version offset is 2 mV at 25 C — generate a constant and a test name.” If the Markdown is wrong, you fix Markdown. If the agent is wrong, you point at a line.
Cursor users should do the same even if they never type Claude. Indexing a PDF pane is not the same as indexing a table. Codex, Copilot-style repo agents, and internal wrappers all tokenize files. They need bytes they can tokenize as text.
Scanned application notes are worse than born-digital datasheets. If page 3 is a photograph of a table, a text-layer assumption fails. That is the same class of problem as a French tender scan, just with microamps instead of lots. KolmoPDF does the parse step; the button below goes to the conversion tool on kolmopdf.com.
Licensing still matters. Do not publish a manufacturer’s full book on a marketing site. Do not commit a forbidden PDF into a public GitHub. An excerpt for a demo, a customer-licensed PDF in a private repo — those are different. The three-page LM358 slice on this page is labeled as an excerpt on purpose.
API users who convert dozens of specs a night can use the parsing API. The useful result is simple: the PDF becomes a small Markdown file that can be diffed. Everything after that — driver stubs, test names, and comments that cite a current limit — belongs to your repository and review process.
Limits: what still fails after a clean parse
Schematics as drawings stay drawings. A parse will not produce a netlist. SPICE models that live only as encrypted blobs stay vendor tools. Mechanical drawings with a forest of dimensions may need the PDF open beside the Markdown, because a table of inches is not the same as a geometric constraint.
Multi-column application notes sometimes still shuffle a caption. You look at the before/after on this page and you keep the original PDF in the same directory as spec.md. The agent can be told “if Markdown and PDF disagree, stop.” That instruction is only possible when both files exist.
This workflow is for a spec you need to check in and cite. If you are preparing a research-paper presentation instead, keep the paper in a separate notes workflow so the source and the audience stay clear.
Can Claude read PDFs? Sometimes, in some products, on some days. Claude Code, in the loop that edits your tree, reads files. Give it one.
Keep the excerpt and the Markdown side by side so a reviewer can open both. If a pin name in spec.md does not appear in the PDF, stop and re-parse that page. Do not let an agent invent a package variant because a screenshot was blurry. Cursor, Claude Code, and Codex all fail the same way when the table never entered git. The useful artifact is a small Markdown file you can diff.
If you convert a stack of datasheets every week, use the parsing API. For one part tonight, use the conversion link above. Either way, the file that lands in git is Markdown with tables, not a screenshot in Slack. Claude Code can then read the file you parsed.
Checklist before you @file the agent
Confirm the Markdown table matches the datasheet excerpt. Confirm units still look like math, not smashed words. Commit spec.md next to a short PDF slice, not the full vendor book. Then mention the file in Claude Code or Cursor. If the agent cites a current limit that is not in spec.md, stop and fix the file first.

