DeepSeek has been generating a lot of buzz lately, and recently announced the launch of its new open-source vision multimodal model, “Janus-Pro-7B.” Janus-Pro-7B can read images provided by users and give corresponding answers based on their questions. The model also performs quite well in image generation.
DeepSeek Launches Its Own Open-Source Multimodal Model Janus-Pro-7B: Image Generation and Text Understanding Reach New Heights
DeepSeek yesterday released news about its new open-source vision multimodal model Janus-Pro-7B. According to information provided by DeepSeek, Janus-Pro is an autoregressive framework that splits the visual encoding process into multiple independent paths, addressing the limitations of previous frameworks.
Just reading the above description may feel abstract, so let’s look at a usage example of Janus-Pro provided by DeepSeek. Janus-Pro demonstrates strong comprehension abilities in scenarios such as describing image content, identifying landmark locations, recognizing text in images, and answering common-sense questions.

In addition, Janus-Pro-7B has also made improvements in image generation. Compared with the original Janus model, Janus-Pro has a better understanding of short prompts, and the image details and quality it generates are superior. Janus-Pro can also understand imaginative and creative prompts and generate logically coherent and consistent images.

Janus-Pro-7B is now open-sourced on GitHub. For those who are interested, feel free toGo to GitHub. Or watch the video below to learn how to install Janus-Pro-7B yourself:
Source: KOCPC Chinese