Export a PDF as structured JSON, or as chunks ready for AI
One section per page, each holding typed elements: headings with their level, paragraphs, lists, and tables as rows of cells. It is the same model our other converters are built on, not a raw text dump.
Feeding a PDF to a language model or a vector store. Each chunk carries the heading it sits under, which is what keeps a paragraph cut out of a section still meaningful once it is retrieved on its own - the single most common cause of bad RAG answers.
No. Parsing and chunking run entirely in your browser. Nothing is sent anywhere, which matters when the documents are contracts or patient records.
Columns are detected and ordered before the text is emitted, and page rotation is applied, so a rotated scan or a two-column paper comes out in reading order.