Skip to main content
Guides

Set Up Qwen-Image-2.1 in ComfyUI: 13-Step AU Guide - tech-insider.org

Alibaba s Qwen team shipped Qwen-Image-2.1 on September 20, 2026, and it changes the calculus for anyone running AI image generation on their own hardware in Australia. The 20-billion-parameter model handles text-to-image generation, multi-reference editing with up to 10 source images, and native transparent PNG output from a single checkpoint.

By Precis Daily Newsroom23 min read5,042 words
Illustration for: Set Up Qwen-Image-2.1 in ComfyUI: 13-Step AU Guide - tech-in
Illustration
Key points
  • Alibaba s Qwen team shipped Qwen-Image-2.1 on September 20, 2026, and it changes the calculus for anyone running AI image generation on their own hardware in Australia.
  • The 20-billion-parameter model handles text-to-image generation, multi-reference editing with up to 10 source images, and native transparent PNG output from a single checkpoint.
  • Qwen-Image-2.1 is an open-weight diffusion model built on an optimised Multimodal Diffusion Transformer (MMDiT) architecture.

Alibaba s Qwen team shipped Qwen-Image-2.1 on September 20, 2026, and it changes the calculus for anyone running AI image generation on their own hardware in Australia. The 20-billion-parameter model handles text-to-image generation, multi-reference editing with up to 10 source images, and native transparent PNG output from a single checkpoint. No separate background-removal pass, no stitched-together workflow of three different tools. For Australian ecommerce sellers, indie game studios, and marketing teams tired of per-image API fees from hosted services, that combination is worth the two hours it takes to set up properly. This tutorial walks through the full local install on ComfyUI, from picking the right model variant for your GPU to building a repeatable product-photo pipeline you can run unattended. It also covers the part most quick-start guides skip: Alibaba tightened the license between Qwen-Image 2.0 and Qwen-Image-2.1, and that change matters if you plan to use the output commercially. Add Tech Insider once in the Google app and our stories appear in your news suggestions. Qwen-Image-2.1 is an open-weight diffusion model built on an optimised Multimodal Diffusion Transformer (MMDiT) architecture. It replaces the February 2026 Qwen-Image 2.0 release and adds three capabilities that competing open models still split across separate checkpoints: native 2K output at 2048×2048 pixels, an RGBA pipeline that generates transparency directly instead of requiring a matting step, and instruction-based editing that accepts up to 10 reference images in one pass. The practical effect is that a single model file now does the job that used to require Stable Diffusion for generation, a separate inpainting model for edits, and a background-removal tool like rembg for cutouts. Alibaba built the model to run through ComfyUI s native node system, with weights hosted on Hugging Face under both the original Qwen organisation and the Comfy-Org mirror. It runs on consumer GPUs once quantised, which is the detail that makes this a realistic home-lab project rather than a cloud-only exercise. The model sits in a lineage that moves fast even by 2026 standards. The original Qwen-Image and its Qwen-Image-Edit sibling both launched under the permissive Apache 2.0 license, and Qwen-Image 2.0 arrived in February 2026 with quality improvements but no transparency or expanded reference support. Qwen-Image-2.1 folds generation and editing back into one checkpoint, adds the RGBA and 10-image capabilities described above, and is the version this tutorial installs and configures step by step. Where this fits for Australian users specifically: product photography for online stores, marketing asset variants without a subscription to Midjourney or Adobe Firefly, and game asset generation where transparent sprites and icons are a daily need. The catch, covered in Step 11 below, is that Alibaba changed the license terms for this release, and it is not the same permissive Apache 2.0 that shipped with the original Qwen-Image. Qwen-Image-2.1 also arrives at a point where Australian small businesses are actively shopping around for cheaper alternatives to per-seat AI subscriptions priced in US dollars, which get more expensive every time the exchange rate moves against the Australian dollar. Running a model locally sidesteps that exposure entirely: once the weights are on your drive, generation cost is just electricity and, if you rent burst capacity, a card charged in whatever currency the provider bills in. This guide assumes zero prior ComfyUI experience and works through the full setup end to end, including the parts that trip up first-timers most often. Prerequisites: Hardware, Software and Account Requirements Before downloading anything, confirm your setup matches these minimums. Qwen-Image-2.1 is more forgiving than most 2K-capable image models, but it still needs a discrete GPU with enough VRAM headroom for the diffusion model, text encoder, and VAE to sit in memory at once. GPU: Nvidia card with at least 8GB VRAM for GGUF quantised builds, 12GB+ recommended for INT8, 24GB for full BF16 precision. AMD cards work through ComfyUI s ROCm build but expect slower generation. System RAM: 16GB minimum, 32GB recommended if you plan to keep a browser and other apps open while generating. Disk space: At least 25GB free. The INT8 diffusion model alone runs approximately 7.3GB, the BF16 version approximately 14GB, and you need room for the text encoder and VAE on top of that. ComfyUI: Latest stable or nightly build. Older builds do not recognise the Qwen-Image-2.1 nodes. Python: 3.10 or 3.11, matching whatever your ComfyUI installation already uses. Hugging Face account: Free account recommended for faster, resumable downloads via the CLI, though the model files are publicly accessible without login. Operating system: Windows 11, Ubuntu 22.04+ or macOS with Apple Silicon (slower, CPU/MPS fallback for some nodes). If your GPU sits below 8GB VRAM, don t skip straight to a cloud rental. Try the Q4_K_M GGUF build first (Step 2 covers this) since it runs acceptably on cards as old as an RTX 3060 12GB. Only move to rented compute if generation times become impractical for your workload. Step 1: Update ComfyUI to a Compatible Build Qwen-Image-2.1 s nodes were added to ComfyUI s core in a September 2026 update. If you installed ComfyUI any earlier than that, the workflow template simply won t load, and you ll see missing-node errors the moment you try to import it. Update first, always. # Update the Python dependencies to match pip install -r requirements.txt --upgrade # Restart ComfyUI completely (don't just refresh the browser tab) If you installed ComfyUI through the desktop app rather than from source, use the in-app updater under Settings, then fully quit and relaunch rather than reloading the interface. A partial restart is one of the most common reasons the new nodes fail to appear, because the Python backend caches the node registry on startup. Step 2: Choose Your Qwen-Image-2.1 Model Variant Alibaba and Comfy-Org distribute Qwen-Image-2.1 in three precision tiers, plus community GGUF quantisations for lower-VRAM cards. Picking the wrong one is the single biggest reason people report it s too slow or it won t even load on forums. Match your GPU to the table below before downloading anything. Variant Approx. File Size Minimum VRAM Best For BF16 (full precision) ~14 GB 24 GB Maximum quality, RTX 4090/5090, professional output INT8 convrot (recommended) ~7.3 GB 12 GB Best quality-to-VRAM ratio for most users GGUF Q6_K / Q8_0 ~6-8 GB 10-12 GB Near-INT8 quality on mid-range cards GGUF Q4_K_M (recommended low-VRAM start) ~4-5 GB 8 GB RTX 3060 12GB, RTX 4060, older cards GGUF Q2_K / Q3_K_M ~2-3 GB 6-8 GB Emergency fallback only, visible quality loss For most people building their first workflow, the INT8 convrot build is the right starting point if you have 12GB or more VRAM. It shrinks the file to roughly half the BF16 size with minimal visible quality loss on typical product and marketing imagery. If you re under 12GB, jump to the GGUF section in Step 4 before downloading anything else. Step 3: Download the Model, Text Encoder and VAE Files Qwen-Image-2.1 needs three separate files placed in three separate ComfyUI folders. Miss one, or put it in the wrong folder, and the workflow will throw a missing model error the moment you queue a generation. qwen_image_2.1_int8_convrot.safetensors models/diffusion_models/ ~7.3 GB qwen3vl_8b_int8_convrot.safetensors models/text_encoders/ ~8 GB qwen_image_2.1_vae_bf16.safetensors models/vae/ ~350 MB Use the Hugging Face CLI for a resumable download rather than dragging files through a browser, which tends to fail partway through on connections that aren t rock solid. # Install the Hugging Face CLI if you don't already have it # Download the diffusion model (INT8, recommended) huggingface-cli download Comfy-Org/Qwen-Image-2.1 \ split_files/diffusion_models/qwen_image_2.1_int8_convrot.safetensors \ --local-dir ComfyUI/models/diffusion_models huggingface-cli download Comfy-Org/Qwen-Image-2.1 \ split_files/text_encoders/qwen3vl_8b_int8_convrot.safetensors \ huggingface-cli download Comfy-Org/Qwen-Image-2.1 \ split_files/vae/qwen_image_2.1_vae_bf16.safetensors \ On a typical Australian NBN 100 connection, expect the full download to take 20-40 minutes depending on time of day and how congested your exchange is. Run it in the background while you move on to the next step. Step 4: Install the ComfyUI-GGUF Node for Low-VRAM Cards Skip this step if you downloaded the INT8 or BF16 build in Step 3. If your card sits under 12GB VRAM and you re using a GGUF quantisation instead, you need the ComfyUI-GGUF custom node, since GGUF loading isn t part of ComfyUI s core node set. # Via ComfyUI Manager (easiest): search "GGUF" in the Manager UI and install git clone https://github.com/city96/ComfyUI-GGUF Place your chosen GGUF file (Q4_K_M-HQv3.gguf as a starting point) into ComfyUI/models/unet/ , not the diffusion_models folder used by the standard checkpoints. This is the single most common mistake reported in GGUF setup threads: the file loads fine but never appears in the node dropdown because it s sitting in the wrong directory. Restart ComfyUI after moving the file so it re-scans the folder. Step 5: Load the Native Qwen-Image-2.1 Workflow Template With the models in place, open ComfyUI and load the pre-built workflow rather than wiring nodes from scratch. Go to the Workflow menu, select Browse Templates, and search for Qwen Image 2.1. ComfyUI ships an official template that already has the correct node chain: model loader, Text Encode Qwen Image 2.1 node, sampler, VAE decode, and save-image node. If the template doesn t appear in your list, your ComfyUI build is still out of date, so go back to Step 1. Loading the official template first, before attempting any customisation, saves hours of debugging node connections that a fresh install gets wrong on the first try. Step 6: Configure the Text Encode Qwen Image 2.1 Node This node is where your prompt, negative prompt, and reference images all connect. It sits under the model/conditioning/qwen_image category if you re building the graph manually rather than using the template. Set your positive prompt describing the desired output, leave the negative prompt blank or minimal since Qwen-Image-2.1 responds better to clear positive instructions than heavy negative steering. For a pure text-to-image generation, leave the image_1 through image_10 inputs empty. You ll wire those up in Step 8 when doing multi-reference editing. Double-check the model loader node is pointed at the correct checkpoint file from Step 3 or Step 4, since a mismatched filename here is the second most common cause of workflow failures after the folder-placement issue above. Step 7: Generate Your First Text-to-Image Output Set your sampler settings to the values Alibaba and Comfy-Org both recommend as the tested baseline: Euler sampler, Simple scheduler, 25 steps, and a CFG value of 1. That CFG number looks unusually low if you re used to Stable Diffusion workflows running CFG 7-12, but Qwen-Image-2.1 was trained to need minimal steering away from the base prompt, and pushing CFG higher tends to introduce artefacts rather than improve fidelity. Set resolution to 2048×2048 for native 2K output, or drop to 1024×1024 for faster iteration while you re still testing prompts. Click Queue Prompt. On an RTX 4090 with the INT8 build, expect a 2K generation to complete in roughly 15-25 seconds. On a GGUF Q4_K_M build on an RTX 3060 12GB, expect 45-90 seconds depending on your exact card and driver version. Output example: a prompt like a matte black wireless keyboard on a plain white studio background, soft top-down lighting, product photography style should return a clean, correctly-lit product shot with sharp keycap detail and accurate colour reproduction, with typography-heavy prompts (logos, packaging text) rendering legibly, which is one of Qwen-Image-2.1 s specific strengths over older diffusion models that routinely garble text. Step 8: Multi-Reference Editing With Up to 10 Input Images This is Qwen-Image-2.1 s headline feature over its predecessor. Connect up to 10 images to the image_1 through image_10 inputs on the Text Encode node, then write an instruction-style prompt describing the edit rather than a description of the desired scene from scratch. For example: combine the product from image_1 with the background from image_2, match the lighting direction from image_1. Keep your reference images at consistent resolutions where possible. Mixing a 4K reference with a 512px reference in the same edit tends to bias the output toward the higher-resolution source s detail level in ways that can look inconsistent. For product consistency across a catalogue, keep one hero reference image at full resolution and treat the rest as lower-weight style or background references. Practical use cases Australian sellers are already running with this: swapping a product into 10 different lifestyle backgrounds from a single studio shot, generating size-variant mockups (same product, different colourways) without reshooting, and combining a logo reference with a blank product mockup to generate branded merchandise previews. Prompt Engineering Tips for Better Qwen-Image-2.1 Results Qwen-Image-2.1 responds differently to prompts than the Stable Diffusion-family models most people have muscle memory for. Where SD-style prompts often stack loose keywords separated by commas, Qwen-Image-2.1 s text encoder is built on the Qwen3-VL vision-language model, and it parses full sentences and explicit instructions far more reliably than keyword soup. Write prompts the way you d brief a photographer, not the way you d tag a stock photo. For text-to-image generation, front-load the subject, then describe the setting, lighting and framing in that order. A stainless steel water bottle standing upright on a wooden kitchen bench, morning sunlight through a window on the left, shot from a slight angle, shallow depth of field outperforms a comma-stacked list of the same keywords in random order. For multi-reference edits, name the images explicitly by number ( the jacket from image_1 , the model s pose from image_2 ) rather than relying on the model to infer which reference you mean from context alone. Three habits consistently improve output quality across both generation and editing tasks. First, state the negative space explicitly when you want transparency or a plain background, since Qwen-Image-2.1 treats transparent background as an instruction rather than an inferred default. Second, keep prompts under roughly 75 words. Longer prompts don t get ignored, but the model tends to weight the first and last sentences more heavily than the middle, so critical details buried mid-paragraph sometimes get dropped. Third, when typography needs to render cleanly (packaging text, logos, UI mockups), spell the exact text in quotation marks inside the prompt rather than describing it abstractly, since the model was specifically trained for accurate text rendering and does best with literal instructions. Step 9: Generate a Transparent PNG With the Alpha Channel Qwen-Image-2.1 outputs RGBA natively, meaning the alpha channel comes straight from the model rather than a post-processing matting step. To get a transparent output, include an explicit instruction in your prompt such as isolated on a transparent background, no shadow and use the standard Save Image node, which preserves the alpha channel automatically when saving as PNG. Verify the transparency worked by opening the output PNG in any editor that shows a checkerboard pattern behind transparent pixels, such as GIMP or Photoshop. If the background renders as solid white or black instead of checkerboard, the model produced an RGB image rather than RGBA, which usually means the prompt wasn t explicit enough about transparent background or the Save Image node got swapped for a variant that flattens alpha. This single output format is what replaces a separate background-removal tool like rembg or Remove.bg in most workflows, cutting one full step out of a product-photo pipeline. Step 10: Tune CFG, Sampler and Steps for Speed vs Quality Once your first outputs are working, the fastest way to cut generation time without wrecking quality is dropping steps from 25 to 15-18 for draft iterations, then bumping back to 25-30 for your final render. Euler with the Simple scheduler is the tested default, but Euler Ancestral introduces slightly more variation between identical seeds if you want to generate multiple candidate outputs from one prompt and pick the best. Resist the urge to push CFG above 2-3 even when an output looks slightly off-prompt. Qwen-Image-2.1 s low native CFG requirement means higher values tend to oversaturate colours and introduce compression-like artefacts around edges rather than fixing prompt adherence. If an output isn t matching your prompt, rewrite the prompt to be more specific before touching CFG. Step 11: Automate Generation With a Python Script For batch work like generating 50 product variants overnight, calling ComfyUI s REST API directly beats manually queueing prompts in the browser. ComfyUI exposes a local API on port 8188 by default that accepts workflow JSON and returns generation results. data = json.dumps({"prompt": workflow}).encode("utf-8") req = urllib.request.Request(f"{COMFY_URL}/prompt", data=data) with urllib.request.urlopen(req) as response: def wait_for_completion(prompt_id: str, timeout: int = 120) -> dict: while time.time() - start Export API format) with open("qwen_image_workflow_api.json") as f: "a red ceramic mug on a marble countertop, isolated, transparent background", "a pair of running shoes on a track, isolated, transparent background", "a wireless earbud case open, isolated, transparent background", workflow["6"]["inputs"]["text"] = p # node ID depends on your exported graph print(f"Done: {p[:40]}... -> prompt_id {pid}") Export the API-format JSON from ComfyUI s Save menu after building your workflow manually once through the UI, then swap the node ID referenced in the script above to match your own graph s text-input node. Step 12: Understand the Qwen Research License Before You Publish This is the step most quick-start tutorials skip, and it matters more than any node configuration. The original Qwen-Image and Qwen-Image-Edit models shipped under Apache 2.0, a fully permissive open-source license that allowed commercial use without restriction. Qwen-Image-2.1 shipped under a different license: the Qwen Research License, which permits use and modification for non-commercial purposes only, defined specifically as research or evaluation. Practically, that means generating product images for your own online store s commercial listings, selling generated artwork, or using outputs in paid client work sits outside what the license permits without a separate commercial grant from Alibaba s Qwen team. Hugging Face discussion threads on the model page show users pushing back on this change and asking whether Alibaba will revert to Apache 2.0 as with prior releases, but as of this writing the research-only terms stand. If your use case is commercial, either stick with the Apache 2.0-licensed Qwen-Image-Edit for that specific project, treat local Qwen-Image-2.1 generation as prototyping only, or contact Qwen directly about a commercial license before shipping generated assets to customers. Complete Working Project: An Automated Product-Photo Pipeline Bringing everything together, here s a working pipeline that takes a folder of raw product photos, generates a transparent cutout for each, and composites them onto a set of background templates. This is the exact structure an Australian dropshipping or handmade-goods seller could run to produce a week s worth of listing images in one overnight batch. def queue_and_wait(workflow: dict, timeout: int = 120) -> dict: data = json.dumps({"prompt": workflow}).encode("utf-8") req = urllib.request.Request(f"{COMFY_URL}/prompt", data=data) with urllib.request.urlopen(req) as resp: prompt_id = json.loads(resp.read())["prompt_id"] with open("qwen_cutout_workflow_api.json") as f: image_files = sorted(INPUT_DIR.glob(".jpg")) + sorted(INPUT_DIR.glob(".png")) print(f"Found {len(image_files)} product photos to process") for i, image_path in enumerate(image_files, start=1): workflow = json.loads(json.dumps(workflow_template)) # deep copy per iteration workflow["12"]["inputs"]["image"] = str(image_path.resolve()) "isolate the main product from image_1, remove the background completely, " "transparent background, no shadow, preserve original colours and detail" print(f"[{i}/{len(image_files)}] Processing {image_path.name}") print("Batch complete. Check ComfyUI's output folder for the results.") Run this overnight against a folder of 30-50 raw product shots and you ll wake up to a matching set of transparent cutouts, all processed with identical settings for a consistent look across your catalogue. Remember the license restriction from Step 11 applies here too: this pipeline is straightforward to build and run, but commercial deployment of its outputs needs that separate license clearance. Extend the script to fit your own catalogue structure by adding a second pass that composites each cutout onto a set of background templates, using the multi-reference editing covered in Step 8 instead of a plain transparent output. Point image_1 at the cutout and image_2 at a lifestyle background, then swap the prompt text to something like place the product from image_1 naturally into the scene from image_2, match lighting and shadow direction. Looping that second pass across a handful of background templates turns one studio photo into a full spread of lifestyle variants without a second photoshoot, which is a realistic way for a small print-on-demand or handmade-goods store to cover a season s worth of listing photos from a single studio session. Common Pitfalls When Setting Up Qwen-Image-2.1 Downloading BF16 on a card that can t fit it. A 14GB model file plus text encoder and VAE overhead will not fit on a 12GB card. Check the VRAM table in Step 2 before you start a 14GB download you can t actually load. Placing GGUF files in the wrong folder. GGUF checkpoints go in models/unet/, not models/diffusion_models/ where the standard INT8/BF16 files live. This single mix-up accounts for most model not appearing in dropdown reports. Pushing CFG too high. Coming from Stable Diffusion habits, it s tempting to bump CFG to 7 or 8 for better prompt adherence. On Qwen-Image-2.1 this produces oversaturated, artefact-heavy output. Stay near CFG 1-2. Assuming Apache 2.0 still applies. Anyone who used the original Qwen-Image commercially and assumes the same terms carry over to 2.1 is working under the wrong license. Re-read Step 12 before any commercial deployment. Skipping the ComfyUI update. The Qwen-Image-2.1 nodes only exist in September 2026 and later builds. An outdated install won t show a helpful error, it ll just fail to load the workflow template at all. Mixing wildly different reference-image resolutions. A 4K hero shot paired with a thumbnail-sized reference in the same multi-image edit biases quality unevenly across the output. Troubleshooting Qwen-Image-2.1: Common Errors Fixed Workflow template won t load, missing node errors ComfyUI build predates the September 2026 Qwen-Image-2.1 node addition Run git pull and pip install -r requirements.txt upgrade, then fully restart Model doesn t appear in the loader dropdown File saved to wrong folder, or ComfyUI hasn t rescanned since download Confirm INT8/BF16 files sit in models/diffusion_models/ and GGUF files in models/unet/, then restart ComfyUI CUDA out of memory error mid-generation Model variant too large for available VRAM Drop to a smaller GGUF quant (Q4_K_M) or close other GPU-using applications Output has solid background instead of transparency Prompt didn t explicitly request transparency, or wrong Save node used Add transparent background, no shadow to the prompt and confirm the standard Save Image node is connected Generated text/typography is garbled Steps set too low for a text-heavy prompt Raise steps to 30 and keep CFG near 1 rather than increasing CFG to fix text Multi-reference edit ignores one or more input images Reference images not correctly wired to image_1 through image_10 slots Re-check each Load Image node connects to the correct numbered input on the Text Encode node Generation is extremely slow (5+ minutes per image) Running BF16 or high GGUF quant on insufficient VRAM, causing CPU offload Switch to INT8 or a lower GGUF quant matched to your actual VRAM GGUF node category doesn t appear at all ComfyUI-GGUF custom node not installed or failed to install dependencies Reinstall via ComfyUI Manager or manually run pip install -r requirements.txt inside the ComfyUI-GGUF folder Text encoder fails to load Qwen3-VL 8B encoder missing or placed in the wrong folder Verify it sits in models/text_encoders/ and the filename matches exactly what the workflow references Advanced Tips: LoRAs, Batching and Cloud GPU Economics Once the base workflow is stable, a few refinements are worth adding. LoRAs trained for the original Qwen-Image architecture generally do not carry over cleanly to 2.1 s updated MMDiT structure, so check any LoRA s model card for explicit 2.1 compatibility before loading it rather than assuming backward compatibility. Community fine-tunes built specifically against the 2.1 base are starting to appear on Hugging Face and Civitai as of late September 2026. For batching, queue multiple prompts through ComfyUI s built-in queue rather than launching separate processes, since the model stays loaded in VRAM between generations and avoids the reload penalty. If your local GPU genuinely can t keep up with your batch size, renting is a reasonable option, and the economics differ meaningfully between providers. Provider RTX 4090 (per hour) A100 80GB (per hour) Notes RunPod Community Cloud US$0.34 US$1.39 Third-party host hardware, no formal SLA, cheapest tier RunPod Secure Cloud US$0.69 higher RunPod-operated data centres with SLA guarantees Vast.ai (unverified listings) from US$0.34 from US$0.50 Marketplace of independent hosts, cheapest on consumer cards Vast.ai (verified datacentre) higher higher More reliability, priced closer to RunPod Secure Cloud At roughly US$0.34 an hour (around A$0.52 at typical exchange rates), even a full weekend of heavy batch generation on a rented RTX 4090 costs less than a single tank of petrol. For anyone still deciding between buying hardware and renting, an RTX 4090 currently runs around A$7,399 new in Australia as of early September 2026, while the newer RTX 5090 sits between roughly A$4,249 and A$8,999 depending on retailer and stock, which puts the breakeven point for renting well past a thousand hours of generation time for most home users. What Self-Hosting Actually Saves You Over a Subscription The math changes depending on how many images your workflow actually needs each month. A small business generating a handful of product shots a week barely notices the cost of a hosted subscription. A studio churning through hundreds of variants a day is a different story, and that s where the local setup pays for itself fastest. Monthly Image Volume Hosted Subscription (approx.) Qwen-Image-2.1 Self-Hosted Practical Winner Under 200 images/month ~A$15-30/month entry tier Electricity only, ~A$2-5/month on an existing GPU Either works, subscription is simpler to start 200-1,000 images/month ~A$45-90/month mid tier Electricity plus occasional cloud burst, ~A$10-20/month Self-hosted starts pulling ahead 1,000-5,000 images/month ~A$150-300+/month, often rate-limited One GPU purchase amortised, or steady RunPod/Vast.ai rental Self-hosted, clearly cheaper at volume 5,000+ images/month Enterprise API pricing, scales linearly with usage Dedicated GPU or small rented fleet, cost per image drops sharply Self-hosted, by a wide margin These figures assume you already own or have access to a capable GPU, or you re renting one on-demand rather than leaving it running 24/7. The comparison also ignores the license restriction covered in Step 12: a hosted subscription s price already includes commercial usage rights, while Qwen-Image-2.1 s research license means the self-hosted column only applies cleanly to non-commercial and prototyping work unless you ve secured separate commercial terms from Alibaba. Qwen-Image-2.1 vs Midjourney, Nano Banana Pro and Flux Qwen-Image-2.1 s pitch is different from the hosted image generators most Australians already know. It s not competing on raw stylistic polish against Midjourney s subscription-based aesthetic engine, it s competing on cost, control and specific technical features those services don t offer at all. Model Access Native Transparency Multi-Reference Editing Commercial Use Qwen-Image-2.1 Free, self-hosted (open weights) Yes, native RGBA Yes, up to 10 images Requires separate license (research-only by default) Midjourney Subscription, hosted only No, requires third-party removal Limited, single-image reference Yes, included in paid plans Nano Banana Pro API/subscription, hosted No Limited Yes, included in paid plans Flux (open-weight variants) Free, self-hosted (varies by license) No, standard RGB output Varies by variant Depends on specific Flux license tier The practical takeaway: if your workflow needs transparent cutouts and multi-image editing baked into one model, Qwen-Image-2.1 is currently the most capable open option doing both natively. If you need unrestricted commercial rights out of the box without contacting a vendor, a paid hosted service or a genuinely Apache-licensed alternative remains the simpler path. Speed also plays into the decision: hosted services queue your request on shared infrastructure, so response time varies with overall platform load, while a local Qwen-Image-2.1 install on your own GPU gives you a fixed, predictable generation time regardless of how many other people are using the same service at that moment. The model weights are free to download and run for non-commercial research and evaluation. Commercial use requires a separate license grant directly from Alibaba s Qwen team, which is a change from the Apache 2.0 terms on the original Qwen-Image. 8GB minimum with a GGUF Q4_K_M quantisation, 12GB for the recommended INT8 build, and 24GB for full BF16 precision. Most users get the best quality-to-hardware ratio from the INT8 build on a 12GB+ card. Can I run Qwen-Image-2.1 without ComfyUI? Yes, via the Hugging Face Diffusers library directly in Python, though ComfyUI s native node support and visual workflow make it considerably easier to manage multi-reference editing and batch pipelines without writing custom inference code. What s different between Qwen-Image-2.1 and Qwen-Image 2.0? Version 2.1, released September 20, 2026, added native 2K generation, RGBA transparency output, and expanded multi-reference editing support to 10 input images. It also switched from Apache 2.0 to the more restrictive Qwen Research License. It runs through ComfyUI s ROCm build on supported AMD cards, though generation speed and stability trail Nvidia CUDA support, and community documentation for AMD-specific troubleshooting is thinner than for Nvidia setups. Why does my transparent PNG still show a white background? Usually the prompt didn t explicitly request a transparent background, or the wrong save node was used. Add an explicit instruction like transparent background, no shadow to the prompt and confirm you re using the standard Save Image node. Can I use LoRAs from the original Qwen-Image with version 2.1? Not reliably. The updated MMDiT architecture in 2.1 means most older LoRAs don t transfer cleanly. Check each LoRA s model card for explicit 2.1 compatibility before loading it. Is renting a cloud GPU worth it instead of buying hardware? For occasional or trial use, yes. Community-tier RTX 4090 rentals run around US$0.34 an hour on both RunPod and Vast.ai, which is far cheaper than buying a A$7,000+ card unless you re generating images daily at volume. What s the fastest way to test Qwen-Image-2.1 before committing to a full local install? Rent a pre-configured RTX 4090 or A100 instance on RunPod or Vast.ai for under an hour, load the official ComfyUI template, and run a handful of prompts before deciding whether to invest the time in a permanent local setup on your own hardware. Nadia Dubois is the AI & Innovation Editor at Tech Insider, where she tracks the rapid evolution of artificial intelligence, from foundation models to real-world enterprise deployment. She previously covered AI and startups for La Tribune and contributed to MIT Technology Review s European coverage. Nadia specializes in generative AI, AI regulation, and the intersection of technology and European industrial policy. She holds a dual degree in Computational Linguistics and Journalism from Sciences Po Paris.

Sources

Summarized from the linked originals.

Related stories

Illustration for: How Botika runs full-stack generative AI on Modal
Products & Tools

Botika builds agentic e-commerce teams that automate visual production for global fashion brands. Powered by custom models researched and trained in-house, Botika handles the entire pipeline - from 4K image generation to real-time personalization at scale.

Modal Blog5 min