{"slug":"docling","name":"Docling","domain":"docling.ai","verdict":"As of 2026-07-19, ChatGPT, Claude, Gemini, Grok collectively rank Docling first for self-hosted document parsing software for sensitive data (one of 10 leaderboards it appears on). Source: https://modelsagree.com/product/docling (modelsagree.com, CC BY 4.0).","best_rank":1,"categories":10,"brief":{"category":"best-self-hosted-document-parsing-software-for-sensitive-data","title":"Best self-hosted document parsing software for sensitive data","rank":1,"of":10,"top":null,"day":"2026-07-19","why":[{"t":"fully local for sensitive data","m":["ChatGPT","Claude","Gemini","Grok"],"q":"all running fully offline on CPU or GPU"},{"t":"advanced layout and table extraction","m":["ChatGPT","Claude","Gemini","Grok"],"q":"strong PDF layout analysis, table structure recovery (TableFormer), OCR integration"},{"t":"clean structured Markdown and JSON output","m":["Claude","Gemini","Grok"],"q":"clean Markdown/JSON output"},{"t":"permissive open-source licensing","m":["ChatGPT","Claude","Gemini","Grok"],"q":"its permissive license and no-phone-home design make it the safest default for regulated data"}],"gap":[],"fix":[{"t":"heavy resource and pipeline tuning needs","m":["ChatGPT","Claude","Gemini"],"q":"Complex scans and visually dense documents may require GPU-backed models and pipeline tuning"},{"t":"not a turnkey enterprise platform","m":["ChatGPT","Claude","Grok"],"q":"not a turnkey enterprise platform"},{"t":"lacks a native user interface","m":["Gemini"],"q":"Lacks a native user interface for non-technical users"}]},"entries":[{"slug":"best-self-hosted-document-parsing-software-for-sensitive-data","title":"Best self-hosted document parsing software for sensitive data","rank":1,"of":10,"score":20,"appearances":4,"modelRanks":{"ChatGPT":1,"Claude":1,"Gemini":1,"Grok":1},"reason":"Best overall balance of local-first privacy, permissive licensing, installation ease, and high-fidelity extraction of reading order, tables, formulas, code, and document hierarchy; strong PDF, Office, image, and RAG integrations make it the safest default for developers.","reasons":[{"model":"ChatGPT","reason":"Best overall balance of local-first privacy, permissive licensing, installation ease, and high-fidelity extraction of reading order, tables, formulas, code, and document hierarchy; strong PDF, Office, image, and RAG integrations make it the safest default for developers."},{"model":"Claude","reason":"IBM's open-source toolkit has become the de facto standard for on-prem document parsing — strong PDF layout analysis, table structure recovery (TableFormer), OCR integration, and clean Markdown/JSON output, all running fully offline on CPU or GPU with an MIT license and active maintenance; its permissive license and no-phone-home design make it the safest default for regulated data."},{"model":"Gemini","reason":"Operates entirely locally with Apache 2.0 licensing, leveraging IBM's Granite layout models to accurately extract multi-column text and complex tables into clean Markdown or JSON without data leakage; ranked first due to superior local layout analysis compared to Unstructured."},{"model":"Grok","reason":"IBM open-source library excelling in advanced PDF/DOCX/PPTX layout analysis, table/formula extraction, OCR, and structured Markdown/JSON output; fully local CPU execution ideal for air-gapped/sensitive data environments with strong GenAI integrations (LangChain/LlamaIndex)."}],"fixes":[{"model":"ChatGPT","fix":"Complex scans and visually dense documents may require GPU-backed models and pipeline tuning; it is a parsing toolkit, not a complete enterprise review workflow."},{"model":"Claude","fix":"Heavier and slower than lightweight parsers on large batch jobs, and its accuracy on messy scans still depends on the OCR backend you wire in — it is not a turnkey high-volume production service without your own orchestration."},{"model":"Gemini","fix":"Lacks a native user interface for non-technical users and requires significant local GPU or CPU resources to achieve acceptable processing speeds."},{"model":"Grok","fix":"Library-focused (requires integration for full production workflows/scaling); not a turnkey enterprise platform."}],"updated":"2026-07-19","rank_history":{"days":["2026-07-18","2026-07-19"],"ranks":[1,1]},"api":"https://modelsagree.com/api/v1/best/best-self-hosted-document-parsing-software-for-sensitive-data.json"},{"slug":"best-document-parsing-and-ocr-for-rag","title":"Best document parsing and OCR for RAG","rank":2,"of":7,"score":16,"appearances":4,"modelRanks":{"ChatGPT":3,"Claude":1,"Gemini":2,"Grok":2},"reason":"IBM's open-source parser became the default RAG ingestion layer for good reason — strong layout analysis, best-in-class open table extraction, reading-order handling, and native export to markdown/JSON with RAG-aware chunkers, plus first-class integrations in LangChain, LlamaIndex, and Haystack; it's free, runs fully local (data never leaves your infra), and is permissively licensed. Assumption: the typical practitioner values zero per-page cost and self-hosting over managed convenience.","reasons":[{"model":"Claude","reason":"IBM's open-source parser became the default RAG ingestion layer for good reason — strong layout analysis, best-in-class open table extraction, reading-order handling, and native export to markdown/JSON with RAG-aware chunkers, plus first-class integrations in LangChain, LlamaIndex, and Haystack; it's free, runs fully local (data never leaves your infra), and is permissively licensed. Assumption: the typical practitioner values zero per-page cost and self-hosting over managed convenience."},{"model":"Gemini","reason":"The premier local-first, open-source library (MIT licensed) for document parsing. It leverages lightweight models to deliver near-commercial-grade layout analysis and table extraction on local CPU/GPU hardware while integrating natively with major RAG frameworks."},{"model":"Grok","reason":"IBM's open-source layout intelligence (DocLayNet + TableFormer) excels at reading order, table structure, formulas, and multi-format support with rich Markdown/DocTags output ideal for precise RAG chunking; fully local/privacy-first with excellent framework integrations and competitive table extraction benchmarks."},{"model":"ChatGPT","reason":"Best-value self-hosted option for most teams: open source, privacy-friendly, format-rich, layout-aware, table-capable, and directly integrated with common RAG frameworks, with interchangeable OCR and vision backends."}],"fixes":[{"model":"ChatGPT","fix":"Default pipelines remain less reliable than the leading managed parsers on highly irregular layouts, merged tables, handwriting, and visually dense pages, while advanced configurations add operational complexity."},{"model":"Claude","fix":"Compute-heavy on large scanned corpora — its OCR path (and VLM pipeline) is far slower than cloud APIs without a GPU, and handwriting/degraded-scan accuracy trails the commercial leaders."},{"model":"Gemini","fix":"Requires developers to self-host and manage scale-out execution infrastructure, and lacks built-in visual-to-text diagram translation out of the box."},{"model":"Grok","fix":"Reduce initial model download size and improve throughput/latency on very large document batches for faster real-time agentic use."}],"updated":"2026-07-15","rank_history":{"days":["2026-06-29","2026-06-30","2026-07-08","2026-07-09","2026-07-10","2026-07-12","2026-07-13","2026-07-14","2026-07-15"],"ranks":[5,4,2,4,5,3,3,2,2]},"reasoning_shift":[{"model":"Gemini","from":"2026-07-14","to":"2026-07-15","added":[{"t":"MIT licensed","q":"MIT licensed"},{"t":"Native RAG framework integrations","q":"integrating natively with major RAG frameworks"},{"t":"Manage scale-out infrastructure","q":"Requires developers to self-host and manage scale-out execution infrastructure"}],"dropped":[{"t":"Matches or exceeds LlamaParse","q":"matching or exceeding LlamaParse on standard table and hierarchical layout reconstruction"},{"t":"CPU processing speed can lag","q":"processing speed can lag on CPU-only machines"}]},{"model":"ChatGPT","from":"2026-07-14","to":"2026-07-15","added":[{"t":"Common RAG framework integration","q":"directly integrated with common RAG frameworks"},{"t":"Interchangeable OCR and vision backends","q":"interchangeable OCR and vision backends"},{"t":"Advanced configurations add complexity","q":"advanced configurations add operational complexity"}],"dropped":[{"t":"Strong on formulas and reading order","q":"strong on tables, formulas, reading order, export, and RAG chunking"},{"t":"Slower than agentic services","q":"Slower and materially less accurate than leading agentic services"},{"t":"Weakness on degraded scans and charts","q":"charts, degraded scans"}]},{"model":"Grok","from":"2026-07-08","to":"2026-07-12","added":[{"t":"Formulas and multi-format support","q":"formulas, and multi-format support"},{"t":"Framework integrations","q":"excellent framework integrations"},{"t":"Smaller downloads for agentic use","q":"Reduce initial model download size and improve throughput/latency on very large document batches for faster real-time agentic use."}],"dropped":[{"t":"VLM support","q":"VLM support (Granite)"},{"t":"Benchmark leadership","q":"Consistently leads or ties independent benchmarks"},{"t":"RAG chunking strategies","q":"adding more native RAG-optimized chunking strategies"}]}],"api":"https://modelsagree.com/api/v1/best/best-document-parsing-and-ocr-for-rag.json"},{"slug":"best-layout-aware-document-parser-for-llm-applications","title":"Best layout-aware document parser for LLM applications","rank":2,"of":8,"score":10,"appearances":3,"modelRanks":{"ChatGPT":5,"Claude":2,"Gemini":1},"reason":"Leading open-source, local-first parser with native deep-learning layout analysis (TableFormer) that delivers SOTA table and reading-order extraction without cloud API costs or data privacy risks; assumes practitioner prioritizes privacy and cost-free local scalability; near-tied with LlamaParse on visual layout fidelity.","reasons":[{"model":"Gemini","reason":"Leading open-source, local-first parser with native deep-learning layout analysis (TableFormer) that delivers SOTA table and reading-order extraction without cloud API costs or data privacy risks; assumes practitioner prioritizes privacy and cost-free local scalability; near-tied with LlamaParse on visual layout fidelity."},{"model":"Claude","reason":"Strongest open-source option — permissive MIT license, runs fully local, with capable layout and TableFormer models, native Markdown/JSON export, reading-order recovery, and clean integration into LlamaIndex/LangChain; the default when data can't leave your infrastructure."},{"model":"ChatGPT","reason":"Best structured open-source framework rather than merely PDF-to-Markdown: advanced layout/reading-order/table/formula understanding, a rich lossless DoclingDocument representation, broad format support, local/air-gapped execution, and excellent integration with LLM/RAG stacks; near-tie with Marker if self-hosting, extensibility, and structured provenance matter more than absolute difficult-PDF accuracy. ([Docling][5])"}],"fixes":[{"model":"ChatGPT","fix":"On the nastiest visually complex documents its out-of-box extraction accuracy can trail newer VLM/agentic parsers, so quality-sensitive deployments need benchmarking and pipeline tuning."},{"model":"Claude","fix":"Trails the top commercial engines on messy scans, rotated pages, and the hardest nested tables; quality is tied to your own GPU/compute for throughput."},{"model":"Gemini","fix":"Requires local GPU acceleration for high-throughput batch processing and carries high compute memory overhead on CPU-only infrastructure."}],"updated":"2026-08-10","rank_history":{"days":["2026-08-04","2026-08-10"],"ranks":[1,5]},"api":"https://modelsagree.com/api/v1/best/best-layout-aware-document-parser-for-llm-applications.json"},{"slug":"best-document-parsing-apis-for-rag-pipelines","title":"Best document parsing APIs for RAG pipelines","rank":3,"of":7,"score":10,"appearances":4,"modelRanks":{"ChatGPT":5,"Claude":3,"Gemini":2,"Grok":4},"reason":"(Near-tie with LlamaParse) The strongest open-source, local-first engine for parsing complex structures; it utilizes efficient layout analysis models to output structured Markdown/JSON without data leaving the developer's infrastructure.","reasons":[{"model":"Gemini","reason":"(Near-tie with LlamaParse) The strongest open-source, local-first engine for parsing complex structures; it utilizes efficient layout analysis models to output structured Markdown/JSON without data leaving the developer's infrastructure."},{"model":"Claude","reason":"The strongest open-source option — IBM-backed, MIT-licensed, strong table-structure and reading-order models, native chunking-friendly output, and integrations with LlamaIndex/LangChain/Haystack; runs fully local, which makes it the only credible pick on this list for regulated data, air-gapped deployments, or zero-marginal-cost bulk parsing."},{"model":"Grok","reason":"Top open-source for local/self-hosted RAG: excellent layout/table/formula understanding, multi-format (PDF/DOCX/etc.), Markdown/JSON output with LangChain/LlamaIndex integrations; privacy-first (air-gapped), zero ongoing cost after setup; strong on technical/academic docs."},{"model":"ChatGPT","reason":"Best control-and-value option: open source, locally deployable as a REST API, broad format support, strong document hierarchy and provenance, table/layout models, multimodal enrichment, and native RAG chunk outputs"}],"fixes":[{"model":"ChatGPT","fix":"Self-hosting and model selection create operational work, while difficult scans and pathological layouts can still trail the best managed agentic parsers"},{"model":"Claude","fix":"You own the infrastructure — GPU provisioning, scaling, and upgrades — and accuracy on the gnarliest scanned/handwritten documents still trails the best commercial APIs."},{"model":"Gemini","fix":"It requires downloading model weights and hosting the service on your own CPU/GPU compute infrastructure, adding operational overhead and making it harder to scale dynamically compared to managed APIs."},{"model":"Grok","fix":"Requires local compute/GPU for best performance on complex scans; less \"plug-and-play\" managed scaling than cloud APIs."}],"updated":"2026-07-18","rank_history":{"days":["2026-07-17","2026-07-18"],"ranks":[2,5]},"api":"https://modelsagree.com/api/v1/best/best-document-parsing-apis-for-rag-pipelines.json"},{"slug":"best-document-parsing-api-for-rag-pipelines","title":"Best document parsing API for RAG pipelines","rank":3,"of":8,"score":9,"appearances":3,"modelRanks":{"Claude":3,"Gemini":2,"Grok":4},"reason":"(Near-tie with LlamaParse for layout dominance) High-performance, open-source local framework that uses advanced layout detection and the specialized TableFormer model to convert PDFs, DOCX, PPTX, and HTML into structured Markdown without API costs or data leakage.","reasons":[{"model":"Gemini","reason":"(Near-tie with LlamaParse for layout dominance) High-performance, open-source local framework that uses advanced layout detection and the specialized TableFormer model to convert PDFs, DOCX, PPTX, and HTML into structured Markdown without API costs or data leakage."},{"model":"Claude","reason":"The strongest open-source option — IBM-backed, MIT-licensed, excellent layout and table structure models (TableFormer), native chunking and export to markdown/JSON, runs fully local for privacy-sensitive RAG, and integrates with LangChain/LlamaIndex; near-tie with Azure below for teams that can self-host"},{"model":"Grok","reason":"Strong open-source performer with solid AI-powered layout understanding, Markdown/DocTags output, privacy-friendly local execution; excels in research/self-hosted RAG setups with good table handling and community momentum."}],"fixes":[{"model":"Claude","fix":"You own the infra — GPU provisioning, throughput tuning, and OCR of poor scans are your problem, and per-page latency on CPU is painful at scale"},{"model":"Gemini","fix":"Requires significant local CPU/GPU compute resources, leading to slower throughput for large-volume pipelines compared to horizontally scaled SaaS APIs."},{"model":"Grok","fix":"Generally trails proprietary VLM/agentic options in peak accuracy on the hardest enterprise docs; requires more setup/compute for best results."}],"updated":"2026-07-19","rank_history":{"days":["2026-07-18","2026-07-19"],"ranks":[3,4]},"api":"https://modelsagree.com/api/v1/best/best-document-parsing-api-for-rag-pipelines.json"},{"slug":"best-ai-document-extraction-api","title":"Best AI document extraction API","rank":3,"of":9,"score":8,"appearances":3,"modelRanks":{"Claude":5,"Gemini":3,"Grok":2},"reason":"IBM open-source MIT-licensed parser with outstanding table extraction, multi-format support, local execution, and accuracy on scanned/complex layouts; top performer in independent tests for practical AI pipelines without vendor lock-in.","reasons":[{"model":"Grok","reason":"IBM open-source MIT-licensed parser with outstanding table extraction, multi-format support, local execution, and accuracy on scanned/complex layouts; top performer in independent tests for practical AI pipelines without vendor lock-in."},{"model":"Gemini","reason":"The premier open-source and privacy-first parser (backed by IBM) that provides outstanding layout awareness, hierarchical structure, and RAG chunking without per-page API costs."},{"model":"Claude","reason":"The best open-source option — MIT-licensed, strong layout and table-structure models, clean markdown/JSON output, and free local execution that solves the privacy/data-residency problem no hosted API can; near-tie with LlamaParse, trading managed convenience for control."}],"fixes":[{"model":"Claude","fix":"You own the ops — GPU provisioning, scaling, and throughput tuning — and processing is slower than hosted APIs, so it's wrong for teams that just want an endpoint today."},{"model":"Gemini","fix":"High CPU/memory overhead and relatively slow throughput compared to lightweight parsers, requiring robust self-hosted infrastructure under heavy batch workloads."},{"model":"Grok","fix":"Requires self-management/infra for scale; less \"agentic\" out-of-box than cloud alternatives for highly variable docs."}],"updated":"2026-07-13","rank_history":{"days":["2026-06-25","2026-07-13"],"ranks":[2,3]},"api":"https://modelsagree.com/api/v1/best/best-ai-document-extraction-api.json"},{"slug":"best-pdf-understanding-api-for-multimodal-ai-applications","title":"Best PDF understanding API for multimodal AI applications","rank":3,"of":12,"score":6,"appearances":3,"modelRanks":{"Claude":5,"Gemini":2,"Grok":5},"reason":"Almost tied with LlamaParse, but excels as a local-first, highly efficient Apache 2.0 open-source alternative utilizing state-of-the-art layout and TableFormer extraction models.","reasons":[{"model":"Gemini","reason":"Almost tied with LlamaParse, but excels as a local-first, highly efficient Apache 2.0 open-source alternative utilizing state-of-the-art layout and TableFormer extraction models."},{"model":"Claude","reason":"The strongest open-source option — converts PDFs to structured markdown/JSON with solid table and layout handling, integrates with LangChain/LlamaIndex, runs locally for free with no data leaving your infrastructure; the right pick when privacy, cost-at-scale, or air-gapped deployment rules out hosted APIs."},{"model":"Grok","reason":"Strong open-source PDF/layout understanding with table/formula/reading-order support and VLM options (Granite Docling); efficient for structured output and GenAI readiness, with managed watsonx option."}],"fixes":[{"model":"Claude","fix":"You operate it yourself (GPU helps, throughput is on you), and accuracy on messy scans and complex forms lags the top commercial services."},{"model":"Gemini","fix":"Requires developers to self-host and manage GPU infrastructure to achieve fast inference on scanned documents."},{"model":"Grok","fix":"Newer in managed form; may trail pure commercial VLMs on cutting-edge multimodal reasoning for highly visual docs (not for teams avoiding IBM ecosystem)."}],"updated":"2026-07-19","rank_history":{"days":["2026-07-18","2026-07-19"],"ranks":[5,5]},"api":"https://modelsagree.com/api/v1/best/best-pdf-understanding-api-for-multimodal-ai-applications.json"},{"slug":"best-pdf-understanding-api-for-multimodal-ai-agents","title":"Best PDF understanding API for multimodal AI agents","rank":3,"of":9,"score":5,"appearances":2,"modelRanks":{"ChatGPT":5,"Gemini":2},"reason":"Leading open-source library that executes locally for total data privacy, providing deep-learning layout analysis and high-fidelity Markdown and JSON output optimized for LLMs with zero API fees; near-tie with LlamaParse for teams prioritizing self-hosted control.","reasons":[{"model":"Gemini","reason":"Leading open-source library that executes locally for total data privacy, providing deep-learning layout analysis and high-fidelity Markdown and JSON output optimized for LLMs with zero API fees; near-tie with LlamaParse for teams prioritizing self-hosted control."},{"model":"ChatGPT","reason":"Best open-source/self-hostable choice for many practitioners: strong PDF structure recovery, OCR, tables, formulas, image/chart enrichment, VLM pipelines, structured JSON/Markdown outputs, and a production REST server without mandatory vendor lock-in. ([Docling][5])"}],"fixes":[{"model":"ChatGPT","fix":"Operating models and infrastructure yourself adds complexity, and turnkey accuracy on the nastiest documents can trail specialized managed platforms."},{"model":"Gemini","fix":"Requires engineering overhead to manage local compute infrastructure and GPU scaling for high-throughput batch processing."}],"updated":"2026-08-10","rank_history":{"days":["2026-08-04","2026-08-10"],"ranks":[5,5]},"api":"https://modelsagree.com/api/v1/best/best-pdf-understanding-api-for-multimodal-ai-agents.json"},{"slug":"best-table-extraction-api-for-complex-pdfs","title":"Best table extraction API for complex PDFs","rank":4,"of":8,"score":4,"appearances":2,"modelRanks":{"Claude":5,"Gemini":3},"reason":"Premier open-source table parsing engine using specialized layout models (TableFormer) to extract complex tables into Markdown/JSON with zero API costs, earned under the assumption that data privacy and open-source self-hosting are essential.","reasons":[{"model":"Gemini","reason":"Premier open-source table parsing engine using specialized layout models (TableFormer) to extract complex tables into Markdown/JSON with zero API costs, earned under the assumption that data privacy and open-source self-hosting are essential."},{"model":"Claude","reason":"Best open-source option for complex tables via its TableFormer model — recovers spanning cells and structure impressively well, runs fully on-prem/offline for zero data-egress, and is free with an active project behind it."}],"fixes":[{"model":"Claude","fix":"Self-hosted means you own GPU infra, tuning, and scaling; throughput and edge-case robustness still trail the top commercial APIs, and there's no managed SLA."},{"model":"Gemini","fix":"Requires self-managed GPU/CPU infrastructure and DevOps overhead, with higher latency per page than managed cloud endpoints."}],"updated":"2026-08-10","rank_history":{"days":["2026-08-04","2026-08-10"],"ranks":[4,null]},"api":"https://modelsagree.com/api/v1/best/best-table-extraction-api-for-complex-pdfs.json"},{"slug":"best-table-extraction-api-for-financial-documents","title":"Best table extraction API for financial documents","rank":6,"of":7,"score":4,"appearances":3,"modelRanks":{"Claude":5,"Gemini":4,"Grok":5},"reason":"A powerful, open-source, local-first alternative that allows developers to process sensitive financial documents without cloud dependencies. It leverages a Vision-Language Model pipeline to read table structures visually, achieving high accuracy on complex, borderless financial layouts.","reasons":[{"model":"Gemini","reason":"A powerful, open-source, local-first alternative that allows developers to process sensitive financial documents without cloud dependencies. It leverages a Vision-Language Model pipeline to read table structures visually, achieving high accuracy on complex, borderless financial layouts."},{"model":"Claude","reason":"The strongest open-source option — IBM's TableFormer model recovers table structure (spans, headers) remarkably well, runs fully self-hosted so confidential financial documents never leave your infrastructure, is free at any volume, and is easily wrapped as an internal API; MIT-licensed with an active community. Assumption: ranked for teams willing to operate their own inference."},{"model":"Grok","reason":"Top open-source for accurate, structured table extraction (TableFormer) in financial/scientific PDFs; local/self-hosted control, preserves structure/hierarchies well, cost-effective for privacy-sensitive or custom finance workflows."}],"fixes":[{"model":"Claude","fix":"Not a managed API — you own GPUs/CPUs, scaling, and quality tuning, and accuracy on degraded scans trails Textract and Reducto."},{"model":"Gemini","fix":"Running in high-fidelity VLM mode is computationally expensive and slow, requiring dedicated GPU infrastructure to achieve reasonable throughput."},{"model":"Grok","fix":"Requires more setup/integration effort than managed APIs; performance varies on highly irregular or scanned docs without tuning."}],"updated":"2026-07-19","rank_history":{"days":["2026-07-18","2026-07-19"],"ranks":[7,5]},"api":"https://modelsagree.com/api/v1/best/best-table-extraction-api-for-financial-documents.json"}],"page":"https://modelsagree.com/product/docling","check":"https://modelsagree.com/check?q=Docling","updated":"2026-08-10T18:18:45.051Z","attribution":"modelsagree.com, CC BY 4.0"}