AIPython★ 588
OpenMOSS/MOSS-VL
MOSS-VL: 11B MODEL FOR VIDEO
The open-weight MOSS-VL series of 11-billion-parameter models is designed for understanding long-form and streaming video in real time. It supports multimodal input and efficient stream processing.
// KEY FEATURES
- Processes videos up to several hours long without losing context.
- Performs real-time streaming video analysis with low latency.
- Understands link between video frames and text, supporting multimodal queries.
- Open weights enable fine-tuning the model for specific video-analysis tasks.
#llms#long-video-understanding#multimodal#multimodal-ai#real-time-ai#streaming-video#video-question-answering#video-understanding#video-understanding-vlm#vision#vision-language#vision-language-model
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram