How to Use Stable Diffusion 3.5: 11 Steps, 90 Min [2026] - tech-insider.org
Stable Diffusion 3.5 is the model Stability AI still points to as its default getting started release for local installs, even after the company announced Stable Diffusion 4 in two tiers — Base and Ultra — on April 6, 2026, priced through its API at $0.01 per credit.
![Illustration for: How to Use Stable Diffusion 3.5: 11 Steps, 90 Min [2026] - t](/images/articles/how-to-use-stable-diffusion-3-5-11-steps-90-min-2026-tech-insider-org.webp)
- It replaced Stable Diffusion 3 Medium as the flagship release and comes in three distinct variants, each aimed at a different hardware budget: Large, Large Turbo, and Medium.
- Stability AI s own model card for Stable Diffusion 3.5 Large confirms the model targets both quality and controllability improvements over SD3.
- That means the hardware bar for running SD 3.5 comfortably keeps dropping even as the model itself keeps improving — a rare combination in AI tooling, where new releases usually demand more VRAM, not less.
Stable Diffusion 3.5 is the model Stability AI still points to as its default getting started release for local installs, even after the company announced Stable Diffusion 4 in two tiers — Base and Ultra — on April 6, 2026, priced through its API at $0.01 per credit. As of September 6, 2026, Stability AI itself is still calling Stable Diffusion 3.5 our most powerful image model yet and has confirmed no SD4 weights have actually shipped, according to HowAIWorks.ai — and a July 19, 2026 tutorial from AIToolRanked likewise still lists SD 3.5 Large, Large Turbo, and Medium as the latest official release, which is part of why this guide sticks to the version you can reliably download today. It s the newest full generation in the Stable Diffusion line confirmed to run locally, it ships in three weight classes built for three different kinds of hardware, and every variant is free to download and run on your own machine. This guide walks through installing it, choosing the right variant for your GPU, wiring up ComfyUI, and fixing the errors that trip up almost everyone on their first run. By the end you ll have a working local install capable of generating high-resolution images in seconds, plus a repeatable workflow you can extend with LoRAs, ControlNet, and the NVIDIA TensorRT optimizations that Stability AI and NVIDIA rolled out together in August 2026. Add Tech Insider once in the Google app and our stories appear in your news suggestions. What Is Stable Diffusion 3.5 and Why It Matters in 2026 Stable Diffusion 3.5 is Stability AI s open-weight text-to-image model family, and as of August 2026 it s still the version listed under the Media category on Stability AI s core models page for anyone downloading weights to run locally — even though Stability AI itself announced a successor, Stable Diffusion 4, on April 6, 2026, with an Ultra tier targeting native 4096 4096 output under the same Community License. A July 2026 model card from Best of AI puts SD4 at roughly 12 billion parameters and frames it as the major update after 3.5, but until that model s local weights are as accessible as 3.5 s, SD 3.5 remains the flagship most people can actually install — a community benchmark list updated September 10, 2026 still ranks SD 3.5 Large as the top open image model, scoring DD37.2 across its 8.1 billion parameters, according to MadeByAgents. It replaced Stable Diffusion 3 Medium as the flagship release and comes in three distinct variants, each aimed at a different hardware budget: Large, Large Turbo, and Medium. What makes SD 3.5 different from earlier Stable Diffusion releases is the MMDiT-X architecture (an improved Multimodal Diffusion Transformer), better prompt adherence, and — critically for anyone who remembers SD s early reputation for mangled hands — noticeably fewer anatomy errors. Stability AI s own model card for Stable Diffusion 3.5 Large confirms the model targets both quality and controllability improvements over SD3. The timing matters too. In August 2026, Stability AI and NVIDIA jointly announced the Stable Diffusion 3.5 NIM (NVIDIA Inference Microservice), which packages the model with TensorRT and FP8 optimizations for roughly 2x faster generation and about 40% lower memory usage on RTX GPUs. That investment in 3.5 s inference stack is notable given Stability AI is simultaneously working on an SD4-Video extension for open video generation, slated for the second half of 2026 according to an April 2026 report from Novareview Hub — even as the company wound down its older video model, pulling API support for Stable Video Diffusion effective July 24, 2026, per release notes Releasebot tracked on August 18, 2026. That contrast — retiring the old video API while still shipping fresh inference optimizations for the 3.5 image line — is a sign the company is still actively supporting the 3.5 line rather than treating it as legacy. That means the hardware bar for running SD 3.5 comfortably keeps dropping even as the model itself keeps improving — a rare combination in AI tooling, where new releases usually demand more VRAM, not less. This tutorial covers the local install path using ComfyUI, since Stability AI s own guidance is direct on the point. As the official Stable Diffusion 3.5 Medium model card puts it: For local or self-hosted use, we recommend ComfyUI for node-based UI inference, or diffusers or GitHub for programmatic use. We ll cover both the ComfyUI workflow and a Python/diffusers script, so you can pick whichever fits your project. Which Stable Diffusion 3.5 Variant Should You Pick Before installing anything, decide which variant matches your GPU. Picking the wrong one is the single most common reason people give up on local Stable Diffusion within the first hour — either the model won t load, or it loads and takes four minutes per image on a laptop GPU that was never going to handle an 8-billion-parameter model. For context, an April 2026 launch article from Novareview Hub pegs the newer Stable Diffusion 4 tiers at 12GB VRAM for Base and a full 24GB for Ultra, so SD 3.5 s spread — from 6GB up to the 24GB VRAM that CheckThat.ai s July 2026 review confirms SD 3.5 Large needs at full precision — still covers a noticeably wider range of hardware, and that same Community License keeps every variant free for commercial use at any company under $1M in annual revenue. Variant Parameters Inference Steps Best For Minimum VRAM Typical Use Case SD 3.5 Large ~8.1B 28-40 (standard) Maximum quality, prompt adherence 12-16GB (24GB recommended) Final renders, print-quality output, professional work SD 3.5 Large Turbo ~8.1B (distilled) 4 steps Fast iteration on high-end GPUs 12GB+ Rapid prompt testing, concept exploration SD 3.5 Medium ~2.5B 20-30 Consumer laptops, edge devices 6-8GB Everyday generation, resolutions from 0.25 to 2 megapixels If you re on a laptop with an 8GB RTX 4060 or similar, start with Medium — it uses the improved MMDiT-X architecture specifically so it can run out of the box on consumer hardware, per Stability AI s own release notes. If you re on a 24GB card (RTX 4090, RTX 5080, or better) and want the sharpest possible output, go with Large. If you re iterating on prompts and don t want to wait 20+ seconds per image, Large Turbo s 4-step generation gets you a preview in a couple of seconds so you can refine wording before committing to a full Large render. Prerequisites: What You Need Before You Start Gather these before step one. Skipping ahead without them is the fastest way to waste an afternoon on dependency errors. GPU: NVIDIA GPU with 6GB+ VRAM (8GB+ strongly recommended, 12GB+ for Large). AMD cards work via ROCm on Linux but expect more setup friction. OS: Windows 10/11, Ubuntu 22.04+/24.04, or macOS 14+ (Apple Silicon, via MPS backend — slower than CUDA). Python: version 3.10 or 3.11 (avoid 3.13 — several diffusion dependencies lag behind new Python releases). Git: latest stable release, for cloning ComfyUI and pulling updates. NVIDIA drivers: a current Game Ready or Studio driver with CUDA 12.1+ support. Disk space: at least 40GB free — SD 3.5 Large s safetensors checkpoint alone runs about 16GB, and you ll want room for Medium and Turbo too if you plan to compare them. Hugging Face account: free account at huggingface.co, required to accept Stability AI s community license and download gated weights. ComfyUI: latest version, cloned fresh from the official GitHub repository . PyTorch: latest stable build with CUDA support, installed per PyTorch s official install matrix for your OS and CUDA version. Time estimate: 60-90 minutes for a full clean install, most of which is spent waiting on model downloads rather than active troubleshooting — assuming you have a decent connection, since the Large checkpoint alone is a multi-gigabyte download. Step 1: Install Python and Verify Your Environment Start by confirming Python and pip are correctly installed and on your PATH. Open a terminal (PowerShell on Windows, Terminal on macOS/Linux) and run: You should see Python 3.10 or 3.11, a pip version tied to that install, and — if you have an NVIDIA GPU — a table showing your driver version and CUDA version. If nvidia-smi fails, install or update your GPU driver before continuing; nothing downstream will work without it. Create a dedicated virtual environment so Stable Diffusion s dependencies don t collide with anything else on your system: ComfyUI is the node-based interface Stability AI recommends for local SD 3.5 inference. Clone it and install its requirements: git clone https://github.com/comfyanonymous/ComfyUI.git pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121 The CUDA index URL matters — installing a CPU-only PyTorch build by accident is the number one cause of generation takes 20 minutes complaints in ComfyUI s issue tracker. If you re on an AMD card, swap that line for the ROCm-specific PyTorch wheel instead, and if you re on Apple Silicon, install the standard PyTorch build (MPS support is bundled in) rather than a CUDA wheel. Once dependencies finish installing, do a smoke test to confirm PyTorch actually sees your GPU: python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))" This should print True followed by your GPU model name. If it prints False , stop here and fix your PyTorch/CUDA install before moving to model downloads — everything after this point assumes GPU acceleration is working. Step 3: Download the Stable Diffusion 3.5 Checkpoint Stability AI distributes SD 3.5 weights through Hugging Face, and you ll need to accept the community license on the model page before downloading. Log in to Hugging Face, visit the Stable Diffusion 3.5 Large model card (or the Medium/Turbo card if you re going that route), click through the license agreement, then generate an access token under your account settings. With the token in hand, use the Hugging Face CLI to pull the checkpoint straight into ComfyUI s model directory: huggingface-cli download stabilityai/stable-diffusion-3.5-medium \ Swap the repo name to stabilityai/stable-diffusion-3.5-large or stabilityai/stable-diffusion-3.5-large-turbo depending on which variant you chose in step two. Each download is a large single file (multiple gigabytes) rather than a folder of shards, so a stalled or interrupted download usually means starting the file over rather than resuming — worth knowing before you start it on an unstable connection. Step 4: Download the Text Encoders and VAE Unlike SD 1.5 or SDXL, Stable Diffusion 3.5 uses three separate text encoders (CLIP-L, CLIP-G, and T5-XXL) to interpret your prompt, plus a separate VAE (variational autoencoder) to decode the final image. These need to go in their own ComfyUI folders: huggingface-cli download stabilityai/stable-diffusion-3.5-large \ huggingface-cli download stabilityai/stable-diffusion-3.5-large \ vae/diffusion_pytorch_model.safetensors \ If you re tight on VRAM, grab t5xxl_fp8 instead of t5xxl_fp16 — the T5-XXL encoder alone is one of the larger pieces of the pipeline, and the fp8 version cuts its footprint roughly in half with a barely perceptible quality difference for most prompts. Step 5: Launch ComfyUI and Load the SD 3.5 Workflow python main.py --listen 127.0.0.1 --port 8188 Open http://127.0.0.1:8188 in your browser. ComfyUI ships with default SD 3.5 workflow templates under Workflow → Browse Templates — load the SD3.5 template rather than building nodes from scratch on your first run. It wires up the checkpoint loader, the triple CLIP text encode node, the VAE decode node, and a KSampler with reasonable defaults already connected. In the loader nodes, point ComfyUI at the files you downloaded in steps 3 and 4. If a dropdown doesn t show your checkpoint, it usually means the file landed in the wrong folder — double-check it s sitting directly in models/checkpoints and not in a nested subfolder the loader doesn t scan. Understanding the ComfyUI Node Graph for SD 3.5 Before generating anything, it helps to understand what the default SD3.5 template is actually doing, since debugging a broken workflow later means knowing which node touches what. The graph breaks into five functional blocks: a checkpoint loader that pulls the diffusion model weights into memory, a triple CLIP text encode node that runs your prompt through CLIP-L, CLIP-G, and T5-XXL simultaneously and merges their embeddings, a KSampler that runs the actual denoising loop over your chosen step count, a VAE decode node that turns the sampler s latent output into pixels, and a save image node that writes the result to disk. The triple text encode is the part that trips up people coming from SDXL workflows, where a single CLIP encoder handled everything. In SD 3.5 s graph, if any one of the three encoder inputs is left disconnected, ComfyUI won t always throw a hard error — it may just silently degrade prompt comprehension, which is why the model isn t following my prompt is so often a wiring problem rather than a prompting problem. Before troubleshooting prompt wording, trace each of the three CLIP inputs back to a loaded file. Two settings inside the KSampler are worth understanding rather than leaving on defaults: the sampler/scheduler pair, and denoise strength. For text-to-image generation from scratch, denoise stays at 1.0 (full noise to full image). If you later use SD 3.5 for image-to-image work — refining an existing image rather than generating from nothing — dropping denoise to 0.4-0.6 preserves the input image s structure while still letting the model add detail and correct issues. With the workflow loaded, type a prompt into the positive CLIP text encode node. Stable Diffusion 3.5 responds well to natural, descriptive sentences rather than the comma-stuffed keyword lists that older SD versions preferred — a side effect of the T5-XXL encoder s stronger language understanding. Try something like: A weathered lighthouse on a rocky coastline at golden hour, dramatic clouds, cinematic lighting, photorealistic, 35mm lens Set your sampler steps according to variant: 28-40 for Large, 20-30 for Medium, and exactly 4 for Large Turbo (Turbo is distilled specifically for low step counts — pushing it past 4-8 steps doesn t meaningfully improve quality and can actually degrade it). Set CFG scale around 4.5-7 for Large/Medium and closer to 1-2 for Turbo. Click Queue Prompt. On a 24GB GPU with SD 3.5 Large, expect roughly 15-30 seconds per 1024 1024 image at 30 steps. On an 8GB card with Medium, expect closer to 20-45 seconds depending on resolution. Large Turbo on a high-end card can land under 5 seconds thanks to its 4-step design. Step 7: Generate Images Programmatically With Diffusers If you re building an application rather than clicking through a UI, Stability AI s model card also points to Hugging Face s diffusers library for programmatic use. Install it and run a minimal script: pip install diffusers transformers accelerate sentencepiece protobuf from diffusers import StableDiffusion3Pipeline pipe = StableDiffusion3Pipeline.from_pretrained( "stabilityai/stable-diffusion-3.5-medium", prompt="a weathered lighthouse on a rocky coastline at golden hour, cinematic lighting", negative_prompt="blurry, low quality, distorted", This is the path most developers use to wire SD 3.5 into an internal tool, a batch-generation pipeline, or a web backend, since it doesn t require a running ComfyUI server in the loop. Step 8: Speed It Up With TensorRT and the SD 3.5 NIM If you re on an RTX GPU and generation speed matters — batch jobs, an internal API, or just impatience — the NVIDIA collaboration announced in August 2026 is worth setting up. Stability AI and NVIDIA s joint Stable Diffusion 3.5 NIM packages TensorRT and FP8 optimization specifically for RTX hardware, and NVIDIA s own benchmarks (per Stability AI s announcement) show roughly 2x faster generation and about 40% lower memory usage compared to the unoptimized pipeline. Two ways to get there: pull the prebuilt NIM microservice container if you want a production-ready deployment, or install NVIDIA TensorRT directly and use ComfyUI s TensorRT node extensions to compile the SD 3.5 UNet into an optimized engine for your specific GPU. The container route is faster to get running; the manual TensorRT route gives you more control if you re already deep in a custom pipeline. Either way, the practical effect is the same: what took 25 seconds per image on Large now takes closer to 12-13 seconds, and VRAM headroom that used to force you down to Medium might now let you run Large comfortably. LoRAs (Low-Rank Adaptation weights) let you steer SD 3.5 toward a specific art style, character, or product look without retraining the full model. Download community LoRAs from Civitai — filter specifically for SD 3.5-tagged LoRAs, since LoRAs trained for SDXL or SD 1.5 are not compatible with the 3.5 architecture — and drop the safetensors file into models/loras . For a sense of how active that ecosystem already is, Civitai s Anima anime-style checkpoint, built on the SD 3.5 base with a compact 3.9GB footprint, had racked up 12.6 million generations by July 2026 according to CheckThat.ai — a good reference point if anime-style output is what you re after. In ComfyUI, add a Load LoRA node between your checkpoint loader and the KSampler, then set the strength (typically 0.6-1.0 to start). Too high a strength and the LoRA overrides your prompt entirely; too low and it barely registers. Iterate in increments of 0.1 until the style lands where you want it. Step 10: Add ControlNet for Structural Control ControlNet lets you guide composition using a reference image — a pose skeleton, a depth map, or a canny-edge outline — instead of relying purely on text. Install the ComfyUI ControlNet extension via the Manager, download an SD 3.5-compatible ControlNet checkpoint, and add a Apply ControlNet node feeding into your sampler alongside your text conditioning. This is the step most tutorials skip, but it s what turns generate a random image matching this prompt into generate an image matching this exact pose and this prompt — the difference between a toy and a production tool for anyone doing concept art, product mockups, or storyboard work. Once your workflow is dialed in, ComfyUI supports queuing multiple prompts and exposes an API endpoint ( /prompt ) so you can trigger generations from a script rather than the browser UI. Save your working workflow as a JSON file (Workflow → Export), then POST it programmatically: # Update the prompt text node dynamically workflow["6"]["inputs"]["text"] = "a neon-lit cyberpunk alley at night, rain-slicked streets" response = requests.post("http://127.0.0.1:8188/prompt", json={"prompt": workflow}) Node ID 6 is workflow-specific — export your own workflow first and inspect the JSON to find the correct node index for your text prompt before adapting this script. If you re running Medium on an 8GB card and still hitting out-of-memory errors, a handful of settings make the difference between a workflow that runs and one that crashes mid-generation. These aren t exotic tweaks — they re the standard toolkit for squeezing a diffusion model onto constrained hardware. Start ComfyUI with the --lowvram flag, which keeps only the actively-used portion of the model in GPU memory and shuttles the rest to system RAM between steps. It costs some speed but turns an unusable setup into a working one. If that s still not enough, add --cpu-vae to run VAE decoding on the CPU instead of the GPU — decoding is a relatively small compute step, so the speed hit is minor compared to the VRAM it frees up. python main.py --listen 127.0.0.1 --port 8188 --lowvram --cpu-vae On the diffusers side, the equivalent moves are enable_attention_slicing() and enable_sequential_cpu_offload() , both shown in the batch script later in this guide. Sequential CPU offload is the more aggressive option — it moves model components to CPU RAM when not actively computing, which can drop VRAM usage by more than half at the cost of noticeably slower generation. Reserve it for cards under 6GB where nothing else gets you a working generation at all. Resolution is the other lever worth pulling before switching model variants entirely. Dropping from 1024 1024 to 768 768 cuts VRAM demand substantially since memory usage scales roughly with the square of resolution, not linearly. Generate at 768 768 to confirm your prompt and composition work, then bump back up to full resolution only for the version you re keeping. Common Pitfalls When Setting Up Stable Diffusion 3.5 These are the mistakes that account for most first-run failures, based on the recurring patterns across ComfyUI s issue tracker and Stability AI s community forums. Installing CPU-only PyTorch. Running the plain pip install torch command instead of the CUDA-specific index URL silently installs a CPU build. Generation works but takes 15-20x longer with no error message telling you why. Skipping the Hugging Face license acceptance. SD 3.5 weights are gated. If you download without first clicking through the community license on the model page, the download will 403 even with a valid token. Mixing LoRAs across SD versions. A LoRA trained for SDXL will load without an explicit error in some ComfyUI setups, then produce garbled or unrelated output. Always confirm the LoRA s model card specifies SD 3.5 compatibility. Running Large Turbo at 20+ steps. Turbo is distilled for exactly 4 steps. Cranking the step count doesn t add quality — it wastes time and can introduce artifacts the distillation wasn t trained to avoid. Forgetting the T5-XXL text encoder. Unlike SDXL, SD 3.5 needs all three text encoders loaded. Missing the T5-XXL file causes either a crash or a severe drop in prompt comprehension, since it s the encoder doing most of the language understanding. Using an outdated ComfyUI build. SD 3.5 support was added in specific ComfyUI releases. An older clone or a stale git pull will throw unknown model type errors on load. Ignoring VAE mismatch. Loading a VAE from a different SD generation (say, an SDXL VAE) produces images with visible color banding or washed-out output, since the latent space doesn t match. Troubleshooting: Fixing the Most Common Errors Work through these in order if something breaks — most SD 3.5 problems fall into one of these buckets. CUDA out of memory VRAM insufficient for chosen variant/resolution Switch to Medium, lower resolution to 768 768, or add --lowvram launch flag torch.cuda.is_available() returns False CPU-only PyTorch installed, or driver mismatch Reinstall PyTorch with the correct CUDA index URL matching your driver s CUDA version 403 Forbidden on model download Community license not accepted on Hugging Face Visit the model page while logged in and accept the license before retrying the CLI download Unknown model type on checkpoint load Outdated ComfyUI version Run git pull in the ComfyUI directory, then reinstall requirements Images look washed out or color-shifted Wrong or missing VAE Download the SD 3.5-specific VAE and set it explicitly in the VAE loader node Prompt seems ignored, generic output Missing T5-XXL text encoder Re-download all three text encoders and verify file paths in the CLIP loader Extremely slow generation (5+ minutes/image) Running on CPU instead of GPU Verify CUDA availability with the smoke test in Step 2; reinstall GPU-enabled PyTorch LoRA has no visible effect Strength set too low, or wrong node order Increase LoRA strength incrementally; confirm the LoRA loader sits before the KSampler in the graph Black or blank image output NaN values from incompatible fp16 settings on some GPUs Try --force-fp32 launch flag, or switch to bf16 if your GPU supports it ComfyUI won t start / import errors Python version mismatch or missing dependency Confirm Python 3.10/3.11 in your venv, rerun pip install -r requirements.txt Once the basics are working, a handful of adjustments consistently improve output quality without needing a different model: Write prompts as sentences, not keyword lists. SD 3.5 s T5-XXL encoder understands grammar and relationships between objects far better than SD 1.5 or SDXL did — a cat sitting on top of a red bicycle produces more accurate composition than cat, red bicycle, sitting. Use negative prompts sparingly. SD 3.5 needs fewer negative prompt terms than older models to avoid common artifacts — over-stuffing the negative prompt can actually push output toward the exact style you re trying to exclude. Match sampler to variant. Euler and DPM++ 2M work well for Large and Medium; Turbo variants are specifically tuned around a simplified sampling schedule, so stick to the sampler ComfyUI s official template recommends rather than swapping it out. Batch at lower resolution, upscale the winner. Generate a batch of 4-8 images at 768 768 to find a composition you like, then regenerate just that seed at full 1024 1024 or higher. This cuts wasted GPU time significantly versus generating everything at full resolution. Pin your seed while iterating on prompt wording. Locking the seed and only changing prompt text isolates exactly what each word change does to the output — useful for building an intuition for how SD 3.5 interprets phrasing. Stable Diffusion 3.5 vs Other AI Image Generators in 2026 SD 3.5 isn t operating in a vacuum. A Japanese-language guide updated in September 2026 still names Stable Diffusion 3.5 as the latest Stable Diffusion generation worth running locally, pairing it with Flux.1 as the two open-weight models most commonly recommended together, per GenAI AI s blog. Here s how it stacks up against the other tools people are actively comparing it to in mid-2026. Stable Diffusion 3.5 Yes (Hugging Face) Yes, GPU required Full control, no per-image cost after setup, LoRA/ControlNet ecosystem Adobe Firefly (Image Model 5) No No, cloud-only Commercial-safe licensing, Photoshop integration Ideogram 3.0 No No, cloud-only Legible in-image text, brand/design work Google Gemini 3 (Nano Banana Pro) No No, cloud-only Fast conversational image editing inside Gemini The core tradeoff hasn t changed since SD s earliest releases: cloud tools like Firefly and Ideogram require no setup and no GPU, but every image runs through someone else s credit system. SD 3.5 asks for an hour of setup and a capable GPU, then costs nothing per image beyond electricity — which is why it remains the default choice for anyone generating images at volume, training custom LoRAs, or building an image pipeline into a larger application. Real-World Use Cases for Local Stable Diffusion 3.5 Beyond generating a single image from a browser tab, a local SD 3.5 install shows its value in workflows that would get expensive or slow on a hosted API. Game and indie studios use batch scripts like the one in this guide to generate hundreds of texture variants or concept art passes overnight, something that would run up a meaningful credit bill on a per-image API. Marketing teams train small LoRAs on a brand s existing product photography so every generated asset matches an established visual identity, rather than fighting a generic model toward brand consistency prompt by prompt. Developers building AI-assisted design tools embed the diffusers pipeline directly into their own applications, since a local model means no rate limits, no per-request latency from an external API call, and no dependency on a third party staying online. And researchers and hobbyists use the ControlNet workflow from Step 10 to explore pose and composition variations at a scale that would be tedious to prompt-engineer around one text description at a time. The common thread across all of these is volume and control. If you need one good image occasionally, a cloud tool is faster to reach for. If you need dozens, hundreds, or thousands of images with a consistent style, structural constraints, or brand identity, the setup cost of a local SD 3.5 install pays for itself quickly. Cost Comparison: Self-Hosting vs the Stability AI API If local setup isn t an option — no dedicated GPU, or you need images generated from infrastructure without one — Stability AI s hosted API is the fallback, and it s worth understanding the economics before you commit to either path. Path Upfront Cost Per-Image Cost Best For Self-host (own GPU) $0 if you already own a capable GPU Electricity only High-volume generation, full customization Self-host (cloud GPU rental) $0 ~$0.40/hr for a spot RTX 4090 instance Occasional heavy sessions without owning hardware Stability AI hosted API $0 (pay-as-you-go) ~6.5 credits/image (Large), ~4 credits (Turbo), ~3.5 credits (Medium) at roughly $10 per 1,000 credits Low-volume, no-setup use, quick prototyping At API pricing, Medium works out to roughly $0.035 per image and Large to roughly $0.065 per image. Third-party wrappers add their own markup on top of that — ZeroTwo AI s May 2026 pricing update, for instance, sets its ZeroTwo Pro tier for Stable Diffusion access at $29.99/month, with an Ultra plan running $120/month, well above what the raw API costs at moderate volume. If you re generating more than a few hundred images a month, the math tips toward self-hosting fast — a $0.40/hour cloud GPU rental for a focused generation session usually beats the equivalent API spend once volume climbs past a couple hundred images, and owning the GPU outright makes the per-image cost effectively zero. Worth noting for anyone weighing a hosted-API workflow long-term: Stability AI isn t going anywhere soon — TechCrunch reported in August 2026 that the company closed a $76M Series B, bringing total funding to $232M, up from the $181M total that Companies History tracked as recently as January 2026, with a July 2026 ToolJunction estimate putting its 2026 revenue somewhere in the $50M-$100M range. For a baseline expectation of what working correctly looks like: on SD 3.5 Medium at 1024 1024, 25 steps, CFG 5, a well-structured photorealistic prompt should produce coherent hands, correctly rendered text on signage in roughly half of attempts (still an active limitation across the field, not unique to SD 3.5), and consistent lighting direction across the frame. Large improves hand and detail coherence further, particularly in complex multi-subject scenes. If your output looks noisy, has visible grid artifacts, or shows heavy color banding, revisit the VAE troubleshooting entry above before assuming it s a prompt problem. Complete Working Project: A Minimal Batch Generation Script Putting the pieces together, here s a complete script that loads SD 3.5 Medium, runs a batch of prompts, and saves each output — a working starting point for anyone wiring SD 3.5 into a larger application. from diffusers import StableDiffusion3Pipeline "a misty pine forest at dawn, soft directional light, photorealistic", "a vintage typewriter on a wooden desk, warm afternoon light, macro lens", "a futuristic city skyline at dusk, neon reflections on wet streets", pipe = StableDiffusion3Pipeline.from_pretrained( "stabilityai/stable-diffusion-3.5-medium", pipe.enable_attention_slicing() # helps on 8GB cards # pipe.enable_sequential_cpu_offload() # uncomment on 4-6GB cards; slower but fits in less VRAM negative_prompt="blurry, low quality, watermark", generator=torch.Generator("cuda").manual_seed(42) output_path = OUTPUT_DIR / f"image_{i:03d}.png" print(f"Done. {len(prompts)} images saved to {OUTPUT_DIR}") Run it with python batch_generate.py from your activated virtual environment. On an 8GB card this batch of three 1024 1024 images should complete in under two minutes; on a 24GB card, well under one. Stability AI continues to ship updates to the SD 3.5 ecosystem — the August 2026 NIM release is one example. Keep ComfyUI current with a periodic pull, and check the Stability AI developer release notes for API and model updates that might affect a hosted-API workflow: pip install -r requirements.txt --upgrade Run this monthly at minimum. New ComfyUI releases regularly add support for newer LoRA formats, ControlNet variants, and performance fixes that specifically target the SD 3.5 pipeline. Yes. As confirmed as recently as August 17, 2026, all three variants — Large, Large Turbo, and Medium — remain available as open weights on Hugging Face at no cost, and Stability AI s hosted API still offers a free tier alongside its paid credit system for anyone who d rather skip local setup. What GPU do I need to run Stable Diffusion 3.5 locally? Medium runs on 6-8GB VRAM cards. Large and Large Turbo need 12GB minimum, with 24GB recommended for comfortable headroom at higher resolutions. Can I run Stable Diffusion 3.5 without a GPU? Technically yes on CPU, but generation times stretch to several minutes per image, which makes it impractical for anything beyond a single test render. A GPU is effectively required for real use. What s the difference between SD 3.5 Large and Large Turbo? They share the same 8.1 billion parameter base, but Turbo is distilled specifically for 4-step generation. Large produces marginally higher quality at 28-40 steps; Turbo trades a small amount of quality for dramatically faster generation. Do I need ComfyUI, or can I use Automatic1111? Stability AI s own guidance recommends ComfyUI or the diffusers library for SD 3.5 specifically, since ComfyUI s node graph handles the triple text-encoder setup more directly. Automatic1111 support for SD 3.5 has lagged behind ComfyUI s, so expect rougher edges there. Are SDXL LoRAs compatible with Stable Diffusion 3.5? No. LoRAs are trained against a specific model architecture s weights, and SD 3.5 s MMDiT-X architecture is not compatible with LoRAs trained on SDXL or SD 1.5. You need LoRAs trained specifically for SD 3.5. What resolution can Stable Diffusion 3.5 generate? Per Stability AI s official announcement, SD 3.5 is capable of generating images ranging between 0.25 and 2 megapixel resolution — roughly 512 512 up to around 1536 1536 or equivalent aspect ratios. Is Stable Diffusion 3.5 output safe for commercial use? Check the specific community license tied to the variant you download — Stability AI s licensing terms vary by use case and revenue threshold, so review the license on the Hugging Face model page for your exact commercial scenario before shipping generated images in a paid product. How to Run FLUX Locally in ComfyUI: 13 Steps, 90 Min [2026] How to Use Nano Banana Pro: 11 Steps, 80 Min [2026] How to Use Grok Imagine Image 2.0: 12 Steps, 70 Min [2026] How to Use Midjourney: 10 Steps, 75 Min [2026] Best AI Image Generator 2026: GPT Image 2 Hits 1370 Elo How to Set Up LM Studio: 13 Steps, 80 Min [2026] Nadia Dubois is the AI & Innovation Editor at Tech Insider, where she tracks the rapid evolution of artificial intelligence, from foundation models to real-world enterprise deployment. She previously covered AI and startups for La Tribune and contributed to MIT Technology Review s European coverage. Nadia specializes in generative AI, AI regulation, and the intersection of technology and European industrial policy. She holds a dual degree in Computational Linguistics and Journalism from Sciences Po Paris.
Sources
Related stories

Limits of Confidence in Diffusion
Limits of Confidence in Diffusion - Apple Machine Learning Research research area Methods and Algorithms content type paper published October 2026 Authors Russ Webb, Amitis Shidani, Alice Bizeul, Dan Busbridge Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from...

How Diffusion Controller unifies and simplifies AI image generation
How Diffusion Controller unifies and simplifies AI image generation How Diffusion Controller unifies and simplifies AI image generation Chih-wei Hsu and Moonkyung Ryu , Software Engineers, Google Research We introduce Diffusion Controller, a lightweight "steering damper" network that precisely steers image generation to achieve significantly better prompt alignment.

Stability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding - TechCrunch
Stability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040
Note from Last Week in AI (Andrey): I’m back! And i’m sorry for putting the substack on a silent pause, work got a bit too overwhelming so I fell…