Skip to content

py-idp

General-purpose, AI-enabled Intelligent Document Processing for Python.

Six-stage pipeline (parse → classify → extract → assess → validate → HITL), 12+ LLM backends, Pydantic-schema-driven extraction, built-in eval harness, self-hosted OCR via Nanonets-OCR2-3B, auto-chunking for oversized documents, and AI-driven schema discovery.

Getting started View on GitHub Install from PyPI


Why py-idp?

pain point how py-idp helps
12+ LLM providers, 6 SDKs, no consistency one Pipeline(backend=...) interface; switch with one arg
PDFs with weird tables Docling + pdfplumber fallback; tables survive extraction
OCR for sensitive documents self-hosted Nanonets-OCR2-3B on your hardware (Apple Silicon / CUDA), no API key, offline
50-page documents that overflow the context window automatic chunking (page-based + token-based)
"We want fields X, Y, Z but haven't written the schema yet" discover_schema() turns a hint + a sample into a Pydantic class
Small models get arithmetic wrong confidence assessment flags low-confidence fields for human review; HITL feedback folds into runtime overrides
Production has to handle retries, caching, rate limits RetryingBackend, ExtractionCache, JsonFileStorage (or SqlStorage)
We need to actually measure which backend is best idp eval with side-by-side strategies, schema-valid rate, field-F1, latency

Quick install

pip install py-idp                          # core
pip install py-idp[docling]                 # IBM Docling PDF parser
pip install py-idp[anthropic]               # Anthropic Claude
pip install py-idp[openai]                  # OpenAI GPT-4o (+ China-LLM compatible gateways)
pip install py-idp[ollama]                  # local Ollama
pip install py-idp[hf-vlm]                  # self-hosted Nanonets-OCR2-3B
pip install py-idp[api]                     # FastAPI server (idp.api:app)
pip install py-idp[pdf-render]              # PDF → image rendering for multimodal backends
pip install py-idp[eval]                    # datasets + pandas for `idp eval`
pip install py-idp[dev]                     # pytest + ruff + mypy
pip install py-idp[docs]                    # this site (mkdocs-material)

Combine extras: pip install py-idp[docling,anthropic,eval,dev].

No API key needed to run the in-tree eval, the examples, or the full test suite — the MockBackend is built in.


30-second tour

import idp
from idp.pipeline import Pipeline

result = Pipeline(
    backend="mock",            # or "anthropic", "openai", "china:qwen", "ollama", "nanonets"...
    schema="Invoice",
).run(idp.Document.from_path("invoice.pdf"))

print(result.extraction)        # dict, validated against the Pydantic schema
print(result.confidence)        # per-field 0..1; fields <0.6 routed to HITL
print(result.validation)        # schema + custom predicate outcomes

See the 30-second tour for the full output.


Examples

The examples/ directory on GitHub has 5 numbered, copy-pasteable scripts that show each major use case end-to-end. Every one falls back to the in-tree MockBackend if no API key is set, so they all run in a fresh venv.

# script what it shows
01 pipeline_minimal.py the smallest possible end-to-end run
02 anthropic.py Anthropic Claude 3.5 Sonnet
03 china_qwen.py Qwen via DashScope (one of the 8 China-LLM providers)
04 hitl_loop.py programmatic HITL feedback loop
05 batch.py batch-process multiple documents and save JSON output

Full table with API-key requirements: examples/README.md.


Project status

Maintainer

Royce Lam · @rollroyces · roycelam@umich.edu