AIPython★ 967
mbzuai-oryx/groundingLMM
GROUNDING LMM: SEE AND UNDERSTAND
CVPR 2024 🔥 Grounding Large Multimodal Model (GLaMM) – the first model that generates natural language answers tied to object segmentation masks. It combines scene understanding with precise textual explanations.
// KEY FEATURES
- Generates natural-language answers linked to object masks in the image.
- Performs object segmentation from text and returns the mask together with a description.
- Supports joint vision‑language reasoning to answer complex scene questions.
- Offers a simple Python‑API for integrating GLaMM into any multimedia applications.
#foundation-models#llm-agent#lmm#vision-and-language#vision-language-model
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram