Sparse MoE efficiency
According to the model card, 17B total capacity activates only about 2B parameters per forward pass through 64 routed experts per MoE layer plus one shared expert.
Open Source Model Profile · NucleusAI
Nucleus-Image is a text-to-image MoE diffusion model from NucleusAI. According to the model card, it pairs 17B total parameters with about 2B active per pass and is released as a base model.
Nucleus-Image is published by NucleusAI as a text-to-image model. Captured Safetensors metadata reports about 16.9B parameters under apache-2.0 with diffusers support. According to the model card, it is a 32-layer sparse MoE diffusion transformer with 17B total parameters, about 2B active per pass, released as a base model without post-training optimization.
According to the model card, 17B total capacity activates only about 2B parameters per forward pass through 64 routed experts per MoE layer plus one shared expert.
According to the model card, routing uses Expert-Choice with a decoupled design separating timestep-aware assignment from timestep-conditioned computation.
According to the model card, text tokens contribute only as key-value pairs and their projections cache across denoising steps via TextKVCacheConfig in diffusers.
According to the model card, aspect-ratio bucketing runs from the outset with progressive 256 to 512 to 1024 resolution training and seven documented output sizes.
Source: NucleusAI/Nucleus-Image
Captured: Unknown. Processed: 2026-09-07T19:34:35.117136+00:00.
🌐 Website | 🖥️ GitHub | 🤗 Hugging Face | 📑 Tech Report Introduction Nucleus-Image is a text-to-image generation model built on a sparse mixture-of-experts (MoE) diffusion transformer architecture. It scales to 17B total parameters across 64 routed experts per layer while activating only ~2B parameters per forward pass, establishing a new Pareto frontier in quality-versus-efficiency. Nucleus-Image matches or exceeds leading models including Qwen-Image, GPT Image 1, Seedream 3.0, and Imagen4 on GenEval, DPG-Bench, and OneIG-Bench. This is a base model released without any post-training optimization (no DPO, no reinforcement learni…
F001F002F003F004F005F006F008F009F010F011F012F013F014F015F016F017F018F019