groundingLMM
GLaMM (Grounding Large Multimodal Model) is an end-to-end trained LMM capable of generating natural language responses integrated with object segmentation masks, enabling visual grounding and versatile interaction with images at multiple granularity levels. It introduces the novel task of Grounded Conversation Generation (GCG), supports various downstream applications like referring expression segmentation and region-level captioning, and is underpinned by the large-scale GranD dataset.
groundingLMM is currently grouped under Vision / Multimodal, which makes it easier to evaluate through workflow fit instead of isolated features alone. Based on the available data, it leans most heavily toward Generates natural language responses seamlessly integrated with object segmentation masks. and Interactive visual assistants that understand and respond to user queries about specific image regions.. The listed license is Apache-2.0, which is useful when adoption constraints matter. It also shows measurable community traction with 963 GitHub stars.