Phi-4-reasoning-vision-15B

5.6B parameter multimodal model that unifies speech, vision, and text processing in a single architecture.