Lightweight PDF parser with layout, tables, formulas and bounding boxes
beatrizalmeidaf
75 points
7 comments
October 01, 2026
Related Discussions
Found 5 related stories in 102.7ms across 8,245 title embeddings via pgvector HNSW
- OCR It – pull text out of un-copyable documents for your LLM thiagolima · 122 pts · August 24, 2026 · 47% similar
- Show HN: TeXbrain, a LaTeX editor that runs pdfTeX in the browser via WASM swimmingbrain · 74 pts · August 25, 2026 · 44% similar
- Show HN: TexLite – A lightweight self-hosted LaTeX workspace thisispi · 20 pts · August 26, 2026 · 43% similar
- Mistral OCR 4.1 spelk · 294 pts · August 13, 2026 · 43% similar
- Billion Dollar PDFs rafaepta · 61 pts · July 12, 2026 · 41% similar
Discussion Highlights (5 comments)
beatrizalmeidaf
I built this because extracting PDFs into Markdown/JSON often loses reading order, tables, formulas, figures, and their original locations. The goal is a lightweight document extraction pipeline that preserves document structure and bounding boxes while exporting to Markdown, JSON, Excel and Word. I'm also working on structure-aware semantic chunking for RAG, so retrieved chunks can retain their section, page and exact visual location in the PDF. The project is open source and I'd love feedback on the architecture, extraction quality, and useful use cases.
archeantus
Great work, thanks for sharing.
phenomen
I currently use https://github.com/firecrawl/anydoc in my PDF pipelines. In most cases it performs well. I'll test your lib to compare.
antonly
I just tried this and the results look very off. I put in the MiniSat[1] Paper, and the result appears catastrophically wrong: https://imgur.com/a/i3RHwcQ [1]: http://minisat.se/downloads/MiniSat.pdf
pcthrowaway
I literally just automated some transcription of bank and credit card statements, using `pdftotext` from the homebrew `poppler` package, and Papero is unfortunately vastly inferior to pdftotext for this format (banks in Canada seemingly only allow you to get historic statements as PDFs). Papero seemingly fails to determine any structure here.. for reference, this is the bash script I used to process credit card statements from Scotiabank. While this doesn't generate structured data (which I could certainly do with more work), it generates text files which preserve the layout, making it greppable, which is good enough for my needs right now. for f in *.pdf; do f="${f%.pdf}" if [[ -f "${f}".txt ]]; then continue fi # e.g. "Statement Period Mar 7, 2024 - Apr 4, 2024" period=$(pdftotext -f 1 -l 1 -nopgbrk -layout -x 331 -y 9 -W 277 -H 15 "${f}.pdf" - | xargs | sed 's/Statement Period //') startdate=${period%%-*} enddate=${period#*-} newname=$(gdate -d "${startdate}" +%F)_$(gdate -d "${enddate}" +%F) mv "${f}.pdf" ${newname}.pdf pdftotext -nopgbrk -f 1 -l 1 -x 70 -W 300 -y 276 -H 600 -layout "${newname}.pdf" pdftotext -nopgbrk -f 3 -l 3 -x 70 -W 300 -y 170 -H 720 -layout "${newname}.pdf" - >> ${newname}.txt done Extracting the tabular data here should be pretty straightforward as well, I just haven't needed to do it. LLMs would absolutely eat this kind of task up (writing a script that can turn it into structured data), but I don't have local LLMs set up and don't really wanna send financial records to big AI.