Pull tables out of a PDF as CSV, Excel, JSON or Markdown
Tables drawn with gridlines extract reliably, because the rules tell us exactly where the cells are. Tables with no borders are detected from column alignment and are flagged as inferred in the result, so you know which ones to check against the original.
Two common reasons. If the PDF is a scan, it contains pictures of words and no text at all - run OCR first. If the content is free-flowing text that merely looks tabular to the eye, there is no reliable structure to recover, and we would rather say so than hand you a plausible-looking wrong grid.
Values that are plainly numeric are stored as numbers; anything with a leading zero, a thousands separator or a currency symbol stays text. That is deliberate: storing a zip code or a part number as a number silently drops its leading zeros, and there is no way to get them back.
In Excel each table becomes its own worksheet named for the page it came from. In CSV and Markdown they are separated and labelled - stacking tables with different column counts into one grid would produce a file that opens fine and means nothing.