AIPython★ 1,597
enoch3712/ExtractThinker
EXTRACTTHINKER DOCUMENT INTELLIGENCE
ExtractThinker — an ORM‑style library for LLM integration, including LangChain and LlamaIndex, that streamlines parsing of PDF, DOCX and images, performs layout and analysis and OCR, and provides tools to build flexible document workflows.
// KEY FEATURES
- ORM‑like API for convenient access to documents as database objects.
- Parses PDF, DOCX, images and extracts text, tables and metadata.
- Performs layout analysis and OCR, recognizing document structure and handwritten input.
- Allows building document processing pipelines with integration into LLM chains (LangChain).
#ai#document-image-analysis#document-intelligence#document-parsing#document-processing#langchain#llm#machine-learning#nlp#ocr#openai#pdf
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram