Skip to content

EthenEthenEthen

Open Source Model Profile · Efficient-Large-Model

SANA1.5_4.8B_1024px

SANA1.5_4.8B_1024px is a text-to-image diffusion model from Efficient-Large-Model. According to the model card, the 4.8B BF16 model generates 1024px-based images with efficient training and inference scaling.

Publisher
Efficient-Large-Model
Task
text-to-image
Model type
Unknown
License
apache-2.0
Library
sana
Publication status
Accepted · not indexed

Model overview

SANA1.5_4.8B_1024px is published by Efficient-Large-Model as a text-to-image model. The captured record identifies the sana library with a recorded apache-2.0 license. According to the model card, developed by NVIDIA with Sana, it is a 4.8B-parameter BF16 Linear-Diffusion-Transformer for 1024px-based image generation.

Recorded capabilities

4.8B diffusion transformer

According to the model card, the model is a scalable Linear-Diffusion-Transformer text-to-image model with 4.8B parameters in torch.bfloat16 precision.

1024px generation

According to the model card, it generates 1024px-based images with multi-scale height and width.

Efficient scaling recipe

According to the model card, it scales from 1.6B to 4.8B with 60% training-cost savings, supports depth pruning, and uses VLM selection-based inference scaling.

Documented research source

According to the model card, source code is available through the NVlabs Sana repository, which the publisher recommends for research training and inference.

Use cases in the source record

  • Research artwork generation and design workflows documented as the direct intended use.
  • Hosted text-to-image inference workflows consistent with the snapshot fal-ai listing and the card's free-inference note.

Limitations and unknowns

  • No Safetensors parameter count was extracted from this record; the 4.8B figure is a publisher model-card claim.
  • GenEval and DPGBench results are described only as top-notch without extracted scores and were not independently verified.
  • The publisher limits the model to research use and excludes factual representations of people or events.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: Efficient-Large-Model/SANA1.5_4.8B_1024px

Captured: Unknown. Processed: 2026-09-07T19:35:05.050006+00:00.

🐱 Sana Model Card Model We introduce SANA-1.5 ,an efficient model with scaling of training-time and inference time techniques. SANA-1.5 delivers: efficient model growth from 1.6B Sana-1.0 model to 4.8B, achieving similar or better performance than training from scratch and saving 60% training cost; efficient model depth pruning , slimming any model size as you want; powerful VLM selection based inference scaling , smaller model+inference scaling > larger model; Top-notch GenEval & DPGBench results. Detailed results are shown in the below table. Source code is available at https://github.com/NVlabs/Sana . Model Description Developed…

F001F002F003F004F005F007F008F009F010F011F013F014F015F016