R_REDDYX.XYZ
AIPython1,597

enoch3712/ExtractThinker

EXTRACTTHINKER DOCUMENT INTELLIGENCE

ExtractThinker — an ORM‑style library for LLM integration, including LangChain and LlamaIndex, that streamlines parsing of PDF, DOCX and images, performs layout and analysis and OCR, and provides tools to build flexible document workflows.

// KEY FEATURES

  • ORM‑like API for convenient access to documents as database objects.
  • Parses PDF, DOCX, images and extracts text, tables and metadata.
  • Performs layout analysis and OCR, recognizing document structure and handwritten input.
  • Allows building document processing pipelines with integration into LLM chains (LangChain).
#ai#document-image-analysis#document-intelligence#document-parsing#document-processing#langchain#llm#machine-learning#nlp#ocr#openai#pdf
Open on GitHub →

New repositories every 30 minutes

REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.

Join on Telegram
← Full catalog·Full index

// SIMILAR REPOSITORIES

Claw Code: Fast Code192068n8n Process Automation187437DeepSeek Plugin Framework186631Agent AI Optimization186627AutoGPT: AI Agent184170Everything for Claude Code179304

← FULL CATALOG

enoch3712/ExtractThinker — EXTRACTTHINKER DOCUMENT…