Flux 2 Klein 9B
Maximum fidelity and photographic detail. Ideal for background swaps, photo restoration, and pristine native-resolution finishes.
Create, edit and refine with a growing ecosystem of image models. Choose the engine for the task, then shape the result in a familiar workspace.
Create images from text, edit targeted areas, and refine finishes with full control on your hardware.
Expand brief ideas into complete scenes, direct compositions, and structure parameters with zero cloud.
Read input photos, analyze artistic details, and extract palettes or styles to guide new creations.

Each one has a clear purpose and unlocks a different creative possibility.
Maximum fidelity and photographic detail. Ideal for background swaps, photo restoration, and pristine native-resolution finishes.
Fast and ultra-lightweight. Perfect for testing ideas on the fly and sketching in seconds without stressing your GPU.
Generates from text, refines portraits, and edits specific parts of the image with ease.
Apply changes and edits in seconds. Direct, zero-delay image modifications for agile workflows.
Explore diverse visual aesthetics, from anime and illustration to conceptual looks with strong personality.
Combines 2 to 4 images into a single cohesive scene using spatial color zoning before harmonious generation.
Describe what object to isolate and get an exact, automatic cutout mask powered by text.
Surgical background removal that preserves fine hair strands, transparency, and delicate edges.
Fast, consistent background isolation for product photography, e-commerce, fashion, and everyday assets.
Transform scenes and environments while keeping your subject's exact facial identity intact.
Transfer color palettes, lighting, and mood from any reference image directly to your artwork.
Instantly transfer the vibe, texture, and aesthetic of an inspiration image to your render.
Upscale resolution with crisp fidelity, recovering real textures without hallucinations.
Reposition key and rim lighting in 3D across any portrait or render while preserving subject identity.
Local CPU algorithm (KMeans) converting renders into authentic retro pixel art with pixel-perfect precision.
The execution engine running locally on your hardware. 100% private, zero queues, and unlimited offline renders.
Local intelligent assistant that refines prompts, describes scenes, and recommends styles without cloud queries.
Paste CivitAI generation blocks. Automatically maps base models, LoRAs, and sampling settings for local execution.
Zero technical jargon! Choose your priority for generating images.
Quickly experiment with ideas without stressing your GPU.
Near-original visual fidelity rendered in just seconds.
For the final render where every pore, light reflection, and texture matters.
Several Create flows produce nothing until a language model has written something first. Assist expands your one line into a full scene. Scenes writes a series that hangs together. Layout emits the JSON that places every element on the canvas. That model runs in LM Studio, on your machine, beside everything else — and which one you pick changes how fast the app feels and how different two runs of the same idea come out.
The recommended default for Assist expansions, rich scene direction, and creative dialogue.
Full 9B creative versatility on the smallest VRAM footprint, leaving maximum room for image models.
Engineered for strict JSON coordinates, storyboard panel grids, and precise visual layouts.
Deliberates step by step before outputting. Ideal for complex character logic and strict world rules.
High-density prose and subtle atmospheric nuance, tuned for workstations with 16 GB+ VRAM.
Full uncompressed precision with maximum vocabulary depth for dedicated 24 GB workstations.
| Model | On disk | Comfortable on | What it's good at | Speed | Thinking |
|---|---|---|---|---|---|
qwythos-9b Q4 |
5.2 GB | 8 GB | Everything. The one we'd default to | Fast | Can be switched off |
ornith-1.5-9b Q4 |
5.4 GB | 8 GB | The same, on the smallest memory footprint we measured | Fast | Can be switched off |
gemma-4-12b QAT |
6.5 GB | 12 GB | Strict JSON and templates. Repeats itself more between runs | Fast | Can be switched off |
ministral-3-14b reasoning |
7.7 GB | 12 GB | Only if you specifically want a model that thinks first | Slow on batches | Always on |
qwen3.8-27b 3.7bpw |
11.7 GB | 16 GB | Nothing the 9B doesn't already do faster | Slow | Can be switched off |
qwen3.8-27b Q4_K_M |
15.7 GB | 24 GB | Won't fit beside your desktop on a 16 GB card | Spills to CPU | — |
Measured September 2026 on a 16 GB card with a browser and an editor open, across the four kinds of request the flows actually make. Speed is relative on purpose: your card, and what else you have running, move the absolute numbers more than the model does.
Don't grab the largest model. A 27B takes four times longer and yields less variety than a 9B. For prompt expansion and scenes, a 9B is faster and more creative.
Reasoning models think before answering and multiply generation time by five without improving output quality in these flows.
Drop an image into Nomad Studio and something has to read it. Describe Image turns a picture into a prompt. Bulk Vision writes one prompt per image across a whole folder. Character and style references, and the Layout scene builder, all start by looking at what you gave them. That job goes to a vision model — a separate setting from the one that writes, and the one people forget to choose.
The recommended default for vision. Excels at describing complex scenes at the exact length requested.
Delivers the deepest and most nuanced visual reads in our benchmarks, with zero lag on modern GPUs.
Already your creative writing engine. Point vision here too for rock-solid reads with zero extra downloads.
Matches the 12B in read accuracy on the smallest memory footprint. Leaves maximum headroom for image renders.
Massive linguistic depth for batch analysis and complex visual descriptions on 16 GB+ cards.
Featherweight option for rapid preliminary tagging on entry-level GPUs when memory is severely tight.
| Model | On disk | Comfortable on | What it's good at | Hard image | Speed |
|---|---|---|---|---|---|
gemma-4-12b QAT |
6.7 GB | 12 GB | The one we'd default to. Best at writing the length you asked for | Reliable | Fast |
ministral-3-14b reasoning |
8.5 GB | 12 GB | The most detailed reads of the whole table, and no slower for it | Reliable | Fast |
qwythos-9b Q4 |
6.1 GB | 8 GB | Already your writing model. One model, both jobs | Reliable | Fast |
ornith-1.5-9b Q4 |
6.2 GB | 8 GB | The same, on the smallest footprint that still never missed | Reliable | Fast |
qwen3.6-27b MTP · Q3 |
11.7 GB | 16 GB | Nothing the 12B doesn't do in half the time | Reliable | Slow |
gemma-4-e4b Q4 |
5.9 GB | 8 GB | Describes an image well, then goes blank when asked to copy a style | Fails | Fast |
qwen3.5-4b MTP |
3.1 GB | 6 GB | Works, until it doesn't, with nothing on screen to tell you which | Coin flip | Very fast |
Measured September 2026 through the app itself, not against the model directly. Three references of rising difficulty — a face filling the frame, a wide rooftop scene with a small figure, a cluttered interior — across both jobs the app asks of a vision model: describe this, and describe this in the style of that. "Hard image" is the wide scene with a style to copy, which is where the differences show. Reliable means three good reads out of three. "Fails" means three out of three came back as nine words of style and no description at all. The coin flip is exactly that: four usable reads in eight attempts.
Import a ComfyUI workflow and expose supported controls in the studio. Available controls depend on its nodes; its models and dependencies must be installed in your environment.
Positive and negative prompt nodes, samplers, seed and noise inputs, latent nodes, guiders, checkpoint and UNET loaders, CLIP last-layer nodes, quality presets, image loaders and LoRA loaders.
ComfyUI's API export flattens subgraphs. The studio reconstructs their boundaries from the parentId:childId convention — and when exactly one external input of the same type can replace a subgraph's output, it marks that subgraph as an optional on/off toggle with the rewiring precomputed.
LoRAs are injected into an imported workflow through a sentinel key, normalised and extracted back out again — with a catalogue that pulls identifiers, details and previews from CivitAI.
Model files are checked on disk across every candidate folder before a run starts, and missing files can be downloaded from the import screen.
An imported upscale workflow is inspected for ImageScaleBy versus ImageScale, and the right parameter is exposed automatically.
A built-in batch benchmark compares quality and speed of "describe image" across vision models, with progress polling and a downloadable report.
Nothing here is a closed catalogue. Import a checkpoint or a workflow and it becomes a first-class flow with the same controls as everything else.