Back to news

QuantFit-VTON: a diffusion model that really does try clothes on "in your size"

QuantFit-VTON: a diffusion model that really does try clothes on "in your size"

Modern diffusion models for virtual try-on produce photorealistic images but do not actually control fit: the hem of a new garment usually just repeats the length of the original clothing in the photo, regardless of the stated size.

The authors proposed QuantFit-VTON, a model based on SDXL with a separate measurement encoder that receives explicit numerical parameters of the person and the garment (height, length, width) and embeds them directly into generation through cross-attention. Geometric losses controlling hem length and silhouette were added as well.

As a result, the model achieves photorealism on a par with the best counterparts while being the best in geometric accuracy among all compared methods.

This work is part of the dissertation research of one of our laboratory's graduates and has been published in IEEE Access. The results are available at this link.

Related news