R_REDDYX.XYZ
AIPython967

mbzuai-oryx/groundingLMM

GROUNDING LMM: SEE AND UNDERSTAND

CVPR 2024 🔥 Grounding Large Multimodal Model (GLaMM) – the first model that generates natural language answers tied to object segmentation masks. It combines scene understanding with precise textual explanations.

// KEY FEATURES

  • Generates natural-language answers linked to object masks in the image.
  • Performs object segmentation from text and returns the mask together with a description.
  • Supports joint vision‑language reasoning to answer complex scene questions.
  • Offers a simple Python‑API for integrating GLaMM into any multimedia applications.
#foundation-models#llm-agent#lmm#vision-and-language#vision-language-model
Open on GitHub →

New repositories every 30 minutes

REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.

Join on Telegram
← Full catalog·Full index

// SIMILAR REPOSITORIES

Claw Code: Fast Code192068n8n Process Automation187437DeepSeek Plugin Framework186631Agent AI Optimization186627AutoGPT: AI Agent184170Everything for Claude Code179304

← FULL CATALOG

mbzuai-oryx/groundingLMM — GROUNDING LMM: SEE AND…