document-conversion
Projects tagged with document-conversion on GitHub.
20 projects
mdflux
PythonTurn any document into clean, AI-ready Markdown. Local-first desktop app: reads scanned PDFs, batches folders, runs offline, and uses far fewer tokens than vision models.
doc7
GoTurn documents into AI-ready Markdown with visual understanding
compdf-self-hosted
TypeScriptEdit, convert, and transform documents across PDFs, Office formats, HTML, TXT, CSV, RTF, JSON, and images with ComPDF Self-hosted, an open-source PDF platform.
anydoc
RustConvert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
markitdown
PythonPython tool for converting files and office documents to Markdown.
unity-ui-document-design-system
C#Drop-in design system for Unity 6 UI Toolkit (UIDocument + PanelRenderer, UXML + USS): tokens, fx shader material support, 42 components, 120 icons, mobile-responsive, dark-themed. One stylesheet, zero dependencies. Battle-tested in a shipping cross-play game.
papra
TypeScriptThe minimalistic document archiving platform.
ocr-it
JavaScriptChrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.
ragflow
GoRAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
letsseal
JavaScriptThe open standard for proving any file is real, unaltered and sealed
sloptrim
PythonA local detector for AI-writing patterns. Scores every prose file your agent saves. Python standard library only, no network, no model.
ip-as-logo-skill
A compact Agent Skill for highly simplified, rounded, subtly neo-skeuomorphic IP mascot logos.
RealReplicaBench
HTMLCommerceAgentBench: Benchmarking Long-Horizon Agents in High-Fidelity, Stateful, and Reproducible Replicas of Real Online Services
HomeBox
GoA continuation of HomeBox the inventory and organization system built for the Home User
lift
PythonExtract structured data from documents quickly and accurately.
Token-Saver
PythonA local Claude Desktop extension to query large PDFs with 92–98% fewer tokens. Performs local hybrid search, cites exact page numbers, and keeps your documents private on your machine
PixelRAG
Pythonhttps://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
pdfx
TypeScriptA free-floating 2D Canvas for processing multiple PDF files simultaneously
CodexQB
PythonCodexQB is a Codex plugin for evidence-backed repo comprehension, planning, QA audit, and gated implementation handoffs.
binky
SwiftBinky sorts your files.