Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model | Read Paper on Bytez