• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Tools and Tutorials - Microsoft Open-Sources TRELLIS.2: Generate High-Quality 3D Models from a Single Photo in 3 Seconds, with Full PBR Materials Directly Importable into Unity and Blender

Microsoft Open-Sources TRELLIS.2: Generate High-Quality 3D Models from a Single Photo in 3 Seconds, with Full PBR Materials Directly Importable into Unity and Blender

KOCPC Editor by KOCPC Editor
August 1, 2026
in AI Tools and Tutorials

AI-generated images and videos are nothing new, but 3D model generation has been stuck in an awkward stage where it “looks right at first glance, yet the details can’t hold up to scrutiny.”MicrosoftThe research institute open-sourced TRELLIS.2 at the end of 2025, which pushes the barrier all the way down to “drop in one image, get a usable 3D asset in 3 seconds.” And it’s not a toy you can only look at — it outputs a .glb file with full PBR materials that you can drop straight into Unity, Unreal Engine, or Blender.

Microsoft open-sources TRELLIS.2: Generate high-quality 3D models from a single photo in 3 seconds.

From “solid clay sculpture” to high-precision 3D assets

Existing 3D generation models have several long-standing problems: blurry details, structural distortion, and they give up when faced with open surfaces (like clothing or leaves) or transparent materials (glass, plastic). Microsoft Research Asia used a very vivid analogy in their official article: you input “transparent glass bottle,” and the AI gives you a solid lump of mud instead. The core breakthrough of TRELLIS.2 is that it invented a method called O-Voxel A brand-new 3D representation method. Traditional methods rely on “isosurface fields” to describe solid shapes, and this approach has a fatal flaw: it only recognizes closed structures and is at a loss when faced with open surfaces or internal spaces. So previously, leaves were forced into double-layered thick slabs, and cups couldn’t truly be “hollow.”

O-Voxel completely abandons the field representation approach and instead uses a sparse voxel structure, which robustly handles arbitrary topology—open surfaces, non-manifold geometry, and enclosed internal spaces are all no problem. A leaf is just a thin slice with sharp edges; a glass bottle has distinct inner and outer walls, and what’s inside is also clearly visible.

Full PBR material, not a texture slapped on afterward.

Another major upgrade is in materials. Most 3D generation pipelines first create an untextured white model and then apply textures, a process that can easily result in texture misalignment or blurriness. TRELLIS.2 instead generates geometry and PBR materials—including base color, roughness, metalness, transparency—simultaneously within the same 3D space. This means materials are natively attached from the start, not pasted on after the fact.

How do the actual results look? Metal cups reflect light depending on the viewing angle, rough fabric shows natural diffuse reflection, and objects inside transparent glass bottles are clearly visible. For game development and visual effects, this directly cuts out a huge chunk of time spent on manually adjusting materials.

40 billion parameters, 3-second image generation, powered by 16x spatial compression.

TRELLIS.2 is a large generative model with 4 billion parameters, still using the standard DiT architecture. The key to running so fast lies in its built-in SC-VAE(Sparse Compression Variational Autoencoder). This compression engine achieves 16x spatial downsampling, compressing a fully textured 1024³ 3D asset to about 9.6K latent tokens, with reconstruction that is visually almost lossless.[1]。

The performance on NVIDIA H100 is like this:

  • 512³ resolution: approximately 3 seconds (2 seconds for shape + 1 second for material)
  • 1024³ resolution: approx. 17 seconds (shape 10 seconds + material 7 seconds)
  • 1536³ resolution: Approximately 60 seconds (35 seconds shape + 25 seconds material)

What does 1536³ even mean? Every spike on the pineapple skin is distinct, and fabric patterns with natural folds—even the finest surface scratches—are fully reproduced.

100% open source, with training code also made public.

TRELLIS.2 is not the kind of “open source” that only has an inference demo. The entire project adopts MIT LicenseThe complete training code and pretrained weights are all on GitHub, which means game studios can fine-tune using their own asset library to teach the model to directly generate 3D assets that conform to the studio’s style guidelines. Indie developers don’t have to build models from scratch; they can just submit a concept image and get a base model, then refine it on top.

TRELLIS.2 GitHub Repository

The hardware requirements aren’t too crazy either: you’ll need an NVIDIA GPU with 24GB VRAM or more (official testing used A100 and H100), the environment is limited to Linux, and CUDA 12.4 is recommended.[3]There are also some on Hugging Face.Official demo. You can try playing it directly.

From game development to 3D printing, the application scenarios are very clear.

Yang Jiaolong, Chief Researcher at Microsoft Research Asia, stated the vision for TRELLIS.2 directly: “Make 3D generation as simple as text-to-image and text-to-video generation.”

 Your browser does not support video playback.

Several application directions have already emerged:

  • Game developmentRapid generation of base models for characters, props, and scenes, significantly reducing modeling time.
  • E-commerce 3D DisplayTake one product photo and generate a rotatable 3D display model.
  • 3D printingUpload photo → Generate model → Directly print into a physical object
  • Film and television special effectsRapid Prototyping and Scene Preview

The community is also looking forward to the upcoming feature updates, with the most requested ones including multi-image input for better accuracy, text-to-3D, and text-driven model modification. TRELLIS.2 has pushed image-to-3D from a “technology demo” into the “productivity tool” stage. With 4B parameters, MIT open source, 3-second generation, and full PBR, this combination has almost no rivals in the first half of 2026. If you have related needs, go try it out!

Source: KOCPC Chinese

Tags: 3D modelGithubHugging FaceMicrosoftOpen sourceTRELLIS.2

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed Xuanjie O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology