• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - Microsoft releases Phi-4-Reasoning-Vision-15B open-source model: the first compact multimodal AI with selective reasoning capability

Microsoft releases Phi-4-Reasoning-Vision-15B open-source model: the first compact multimodal AI with selective reasoning capability

KOCPC Editor by KOCPC Editor
March 6, 2026 - Updated on August 5, 2026
in AI Trends and Related News, Latest Technology News

Microsoft officially released the Phi-4-Reasoning-Vision-15B open-source model in March 2026, a multimodal reasoning model with 15B parameters that combines high-resolution visual perception with selective, task-aware reasoning capabilities. As the first Small Language Model (SLM) in the Phi-4 series to achieve both “seeing clearly” and “thinking deeply,” Phi-4-Reasoning-Vision-15B employs an innovative hybrid reasoning design that automatically switches reasoning modes based on task type, opening up entirely new possibilities for AI agent applications.

Background and Development

In recent years, multimodal large language models have evolved rapidly, progressing from initial image classification and object detection toward complex visual understanding and reasoning capabilities. However, traditional vision-language models often face a critical bottleneck: most of them can only perform passive perception tasks, such as recognizing objects in images, generating captions, or answering simple questions. When faced with tasks requiring multi-step logical reasoning, mathematical calculations, or structured analysis, these models often fall short.

Building on the success of the Phi-4 small language model series, Microsoft recognized the need to break through this technical barrier. The birth of Phi-4-Reasoning-Vision-15B is precisely designed to fill this gap, marking an important milestone for small multimodal AI as it evolves from “passive recognition” to “active reasoning.”

Core Technology Features

Phi-4-Reasoning-Vision-15B’s technical architecture is built on two core algorithms:SigLIP-2 Visual encoder and Phi-4 Reasoning Language Model. SigLIP-2 can compress images into numerical representations that neural networks can understand, preserving fine-grained visual information from images.

Adopting a unique approach mid-fusion The architecture performs multimodal information interaction only in the middle layers of the neural network. This design significantly reduces computational overhead while preserving key visual understanding and reasoning capabilities. Unlike traditional full-fusion methods, mid-fusion allows the visual encoder and language model to maintain relatively independent optimization paths.

Innovative Design of Selective Reasoning

Phi-4-Reasoning-Vision-15B’s most innovative design highlight lies in itsMixed reasoning behavior」(Hybrid Reasoning Behavior) mechanism. Traditional multimodal models typically employ a unified processing pipeline, executing the same reasoning path regardless of task complexity. While simple in design, this approach often leads to resource waste: for straightforward tasks like OCR recognition or element localization, invoking a full multi-step reasoning chain is unnecessary.

Phi-4-Reasoning-Vision-15B has revolutionized this landscape. The model features two distinctly different operating modes that can switch automatically or manually based on the task type.[2]When in reasoning mode, the model enables a complete multi-step reasoning chain for structured, deep-level thinking. When in non-reasoning mode, the model skips the lengthy reasoning chain and directly outputs results, greatly reducing latency.

Performance

Phi-4-Reasoning-Vision-15B has demonstrated impressive performance across multiple benchmark tests. According to test data released by the Microsoft research team, this model performs particularly exceptionally on mathematical and scientific reasoning tasks. In MathVista_MINI In benchmark tests, Phi-4-Reasoning-Vision-15B scored higher than Google’s Gemma-3-12b-it by 17%Fully demonstrate its leading position in the field of visual mathematical reasoning.

More impressively, Phi-4-Reasoning-Vision-15B achieves reasoning capabilities comparable to models with more than 10x the parameters at only 15B parameters. This means that on the same tasks, Phi-4-Reasoning-Vision-15B requires significantly fewer computational resources and token consumption.

 

Application Scenarios and Industry Impact

Phi-4-Reasoning-Vision-15B has extremely broad application potential, and one of the most notable application scenarios is the “Computer Agent.” Under this paradigm, the model can receive screenshots as visual input and combine them with natural language instructions to execute complex computer operation tasks.

This capability holds immense value for automated testing, UI design verification, accessibility testing, and similar scenarios. Traditional automation scripts rely on DOM structure or XPath and similar technologies, making them prone to failure when UI changes occur. Phi-4-Reasoning-Vision-15B can directly understand visual layouts and locate elements based on user language descriptions, significantly enhancing the robustness of automation solutions.

Open Source and Usability

Microsoft has officially open-sourced Phi-4-Reasoning-Vision-15B, allowing developers and researchers to Hugging Face Download and use the model for free on platforms like these. The model’s open-source strategy continues Microsoft’s open approach in the AI field in recent years, hoping to drive continuous technological advancement through the power of the community.

Microsoft Research also released detailed technical blogs, sharing valuable experiences and lessons from the model training process. These publicly available knowledge resources are of significant value to the development of the entire AI community, helping to drive continuous innovation in multimodal reasoning technology.

Conclusion

The release of Phi-4-Reasoning-Vision-15B marks a new stage of development in the multimodal AI field. With its innovative selective reasoning design, exceptional performance, and open-source availability, this model opens new directions for compact multimodal model development. As developers and enterprises gradually adopt this technology, we can expect to see more innovative applications based on visual reasoning in the future, driving AI technology toward broader real-world implementation.

Source

Source: KOCPC Chinese

Tags: aiMicrosoftOpen sourcePhi-4-Reasoning-Vision-15B

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology