Xiaomi native omni-modal model. It supports joint understanding of text, image, video and audio, with a 1M-token long context window and multimodal Agent & tool calling capabilities.