R_REDDYX.XYZ
AIPython1,524

emcf/thepipe

THEPIPE: DATA FROM DOCUMENTS FAST

ThePipe uses vision‑language models to extract data from PDF, Word and PowerPoint. It recognises tables, text and images, turning them into ready‑to‑use structures. Supports batch API.

// KEY FEATURES

  • Extracts tables and text from scanned PDF with >95% accuracy via VLM.
  • Converts complex Word documents to structured JSON, preserving styles and metadata.
  • Processes batches of PowerPoint files, extracting text from slides and image captions.
  • Provides REST API and Python SDK for integration into ETL and RAG pipelines.
#document-ai#large-language-models#microsoft-word#multimodal#pdf#powerpoint#python#scanned-pdf#scrapers#scraping#unstructured-data#vision-language-model
Open on GitHub →

New repositories every 30 minutes

REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.

Join on Telegram
← Full catalog·Full index

// SIMILAR REPOSITORIES

Claw Code: Fast Code192068n8n Process Automation187437DeepSeek Plugin Framework186631Agent AI Optimization186627AutoGPT: AI Agent184170Everything for Claude Code179304

← FULL CATALOG

emcf/thepipe — THEPIPE: DATA FROM DOCUMENTS FAST | REDDYX AI