architecture flow
Generation settings go through one shared pipeline context instead of branching across the UI.
The local server handles the queue, seeds, uploads, previews, and long-running jobs.
ModelFactory works out the different model layouts and assembles the diffusion, encoder, and VAE pieces.
The UI is kept separate from the core pipeline.
visuals



engineering notes
LightDiffusion-Next is a local image-generation backend I wrote for speed. It started as a 3,000-line plain PyTorch script (the original LightDiffusion), then got refactored into a modular backend with a FastAPI server, a React UI, a Gradio entry for HuggingFace ZeroGPU, and Docker.
The repository, documentation, and an interactive HuggingFace Space are publicly available.
Speed
Measured on a mobile 3060, SD1.5, 1024×1024, batch 1, BF16, stock installs:
| tool | it/s |
|---|---|
| LightDiffusion + Stable-Fast | 2.8 |
| LightDiffusion | 1.9 |
| ComfyUI | 1.4 |
| SD Forge | 1.3 |
| SD WebUI | 0.9 |
The first version measured about 30% less inference time than the baselines and got into the Ready Tensor CV Projects Expo 2024.
Where the speed comes from
The repo has a source-based optimisation report listing about 35 items with the file each one lives in. The ones that matter most:
- Attention cascade. SpargeAttn, then SageAttention, then xformers, then PyTorch SDPA, whichever is present. Flux2 prefers cuDNN / Flash SDPA.
- Caches. Prompt embedding cache so a repeated prompt is not re-encoded. Cross-attention K/V projection cache for static context. DeepCache and First Block Cache reuse denoiser output between steps.
- Sampling. AYS scheduler by default (same quality in fewer steps), CFG++ samplers, CFG=1 skips the unconditional branch, optional CFG-free tapering and dynamic CFG rescale. Multi-scale latent switching does part of the denoising at lower resolution.
- Compilation and precision. Stable-Fast traces the UNet.
torch.compileas the alternative. BF16/FP16 picked per hardware. FP8 and NVFP4 weight quantisation, and load-time weight-only quantisation so Flux2 Klein fits on smaller VRAM. - Memory. Partial loading and offload policy for low-VRAM cards (down to 2 GB, or CPU only). Pinned checkpoint tensors and async transfers. Tiled VAE.
- Serving. The FastAPI server coalesces compatible requests into one batch, prefetches the next checkpoint, keeps models loaded, and returns PNG bytes from memory instead of disk.
What it runs
SD1.5, SDXL, Flux2 Klein, LoRAs, textual inversion. Hires-Fix, ADetailer (Impact Pack based), UltimateSD upscale, img2img, TAESD live previews, an optional Ollama prompt enhancer. CUDA, ROCm, and Apple MPS. It also served as the image backend for a Discord bot (Boubou) and a Newelle extension.