marker
https://github.com/datalab-to/marker
Python
Convert PDF to markdown + JSON quickly with high accuracy
Triage Issues!
When you volunteer to triage issues, you'll receive an email each day with a link to an open issue that needs help in this project. You'll also receive instructions on how to triage issues.
Triage Docs!
Receive a documented method or class from your favorite GitHub repos in your inbox every day. If you're really pro, receive undocumented methods or classes and supercharge your commit history.
Python not yet supported0 Subscribers
Add a CodeTriage badge to marker
Help out
- Issues
- Your project is on StackMap — a curated map of the AI stack
- Derive FONT_PATH from FONT_DIR at runtime and cache the font outside the package
- [Feature Request / Bug] Issues with mid-sentence image insertion, internal link escaping, and output file extension
- Text from no-ToUnicode CID fonts with lying glyph names bypasses flag_bad_blocks; garbled text and control bytes ship in markdown
- Deterministically re-OCR lying-font text damage; strip control bytes from markdown
- [BUG: Breaking]ollama, 400 error: Bad Request ,and error: LLM did not return a valid response
- Fix trailing-space output paths, duplicated OCR list numbers, empty OCR images
- fix: keep height/label pairs and clamp cluster count in bucket_headings
- [Feature Request] Chinese document parsing resilience — post-processing fallback, structured validation, per-page error isolation
- bucket_headings(): np.sort breaks the height/label pairing, and n_clusters is not clamped to distinct sizes
- Docs
- Python not yet supported