1024px BF16 text-to-image model
According to the model card, the 4.8B-parameter model uses torch.bfloat16 precision and targets 1024px-based image generation.
Open Source Model Profile · Efficient-Large-Model
SANA1.5_4.8B_1024px_diffusers is a 4.8B text-to-image model from Efficient-Large-Model. Its model card documents BF16 precision and 1024px-based generation.
SANA1.5_4.8B_1024px_diffusers is published by Efficient-Large-Model as a text-to-image model under the sana library. The publisher introduces it as SANA-1.5, an efficient scaling of training-time and inference-time techniques. According to the model card, it is a 4.8B-parameter BF16 model built for 1024px-based generation.
According to the model card, the 4.8B-parameter model uses torch.bfloat16 precision and targets 1024px-based image generation.
The publisher describes growth from a 1.6B Sana-1.0 model to 4.8B with depth pruning and VLM-selection inference scaling.
The model card documents a SanaPipeline snippet with text-encoder BF16 handling and a 1024x1024 example prompt.
Source: Efficient-Large-Model/SANA1.5_4.8B_1024px_diffusers
Captured: Unknown. Processed: 2026-09-07T19:35:05.057223+00:00.
🐱 Sana Model Card Model We introduce SANA-1.5 ,an efficient model with scaling of training-time and inference time techniques. SANA-1.5 delivers: efficient model growth from 1.6B Sana-1.0 model to 4.8B, achieving similar or better performance than training from scratch and saving 60% training cost; efficient model depth pruning , slimming any model size as you want; powerful VLM selection based inference scaling , smaller model+inference scaling > larger model; Top-notch GenEval & DPGBench results. Detailed results are shown in the below table. Source code is available at https://github.com/NVlabs/Sana . Model Description Developed…
F001F002F003F004F005F006F007F008F009F010F011F015F016F017