Read this first: prompt keywords are the smallest lever
Most guides on this topic hand you a wall of quality tags and imply that better keywords produce a better upscale. That is not how Stable Diffusion upscaling works, and if you have been stacking 8K ultra HD, masterpiece, extreme fine detail and getting mush, this is why.
Here is the honest ranking of what controls the quality of an SD upscale, strongest first:
| Rank | Lever | Why it dominates |
|---|---|---|
| 1 | Denoising strength | Decides whether the model enhances your image or repaints it. One slider, enormous effect. |
| 2 | Upscaler model choice | 4x-UltraSharp, Real-ESRGAN and SUPIR produce visibly different textures on identical input. |
| 3 | ControlNet Tile | Conditions every tile on the original pixels. The difference between "sharper" and "hallucinated". |
| 4 | Tiling and pass count | 2x twice beats 4x once. Tile size and overlap decide whether you get seams. |
| 5 | Base model / checkpoint | An SD 1.5 photo checkpoint and Flux respond to the same prompt completely differently. |
| 6 | Prompt keywords | Real, but a trim on top of the above — not a substitute for any of them. |
Keywords still matter, and the tested stacks are below. But if your upscale is soft, warped, or has grown details that were never in the source, the fix is in rows 1-4. No keyword combination repairs a denoise of 0.6.
The other thing to be clear about: prompt keywords never change pixel dimensions. SD generates at its base resolution — 512×512 for SD 1.5, 1024×1024 for SDXL and Flux — and a separate upscaling step is what produces a larger file. Quality keywords increase detail density within a generation. They do not produce a 4K file.
Which interface you are in
The prompts and numeric values in this guide are identical across all three. Only the plumbing differs.
| Interface | Status in 2026 | Upscaling approach |
|---|---|---|
| ComfyUI | Where most new upscaling work lands first | UltimateSDUpscale node, tiled KSampler, SUPIR nodes, ControlNet Tile |
| Forge / reForge | Maintained A1111-style UI | Ultimate SD Upscale script, Extras tab, built-in ControlNet |
| Automatic1111 | Still works; development has slowed | Same as Forge — img2img + Ultimate SD Upscale |
If you are starting fresh in 2026 and expect to do a lot of upscaling, ComfyUI is the better investment — multi-pass upscale graphs are reusable in a way that a form-based UI never quite manages. If you already have an A1111 setup that works, nothing here requires you to move.
Quality tags: what they are and when they stop working
"Quality tags" are the short reputation words people append to prompts — masterpiece, best quality, highly detailed, 8k, ultra HD. They exist because SD 1.5 and the anime checkpoints descended from it were trained on booru-style tagged datasets where those literal strings appeared on highly-rated images. The model learned the correlation. Typing masterpiece genuinely steers toward the kind of image that got tagged that way.
That correlation weakens as you move forward through model generations:
| Base model | Response to stacked quality tags | What to write instead |
|---|---|---|
| SD 1.5 and its descendants | Strong. Tag stacking works as advertised. | Use the full stacks below. |
| SDXL / Pony / Illustrious | Moderate. Helps, with diminishing returns past ~8 tags. | Trim the stack; keep the camera and material words. |
| SD 3.5 | Weak. Trained on natural-language captions. | Describe the image in a sentence, then add 3-4 tags. |
| Flux dev / schnell | Minimal. Tag spam can crowd out the real description. | Write plain descriptive prose. Skip the tag wall. |
The practical rule: quality tags are a dialect, and you have to speak your checkpoint's dialect. Pasting an SD 1.5 tag wall into a Flux workflow is the single most common reason people report that "quality keywords do nothing."
Positive prompt keyword stacks
Use these during the diffusion pass of an upscale. They are written for SD 1.5 and SDXL photo checkpoints — trim to the first six or seven terms on SDXL, and rewrite as prose on Flux.
Photography upscaling (photorealistic):
RAW photo, ultra-high resolution, tack sharp, highly detailed, photorealistic, professional photography, extreme fine detail, rich color depth
Portrait pixel restoration:
RAW photo, tack sharp, pore-level skin detail, natural skin texture, highly detailed face, photorealistic, professional portrait photography, clean sharp eyes
Landscape upscaling:
RAW photo, ultra-high resolution, tack sharp throughout, landscape photography, highly detailed vegetation and terrain, photorealistic, cinematic dynamic range, extreme fine detail
Product photo restoration:
commercial product photography, razor-sharp edges, highly detailed surface textures, photorealistic, accurate color reproduction, clean studio lighting, extreme fine detail
A note on 8K ultra HD: it is in almost every list on the internet, including the earlier version of this one. It does something on SD 1.5 and close to nothing on SDXL and Flux. It is harmless, but do not treat it as load-bearing — and never as a reason to skip an actual upscaler.
Negative prompt stacks
Universal upscaling negative (SD 1.5 / SDXL):
blurry, out of focus, low resolution, low quality, pixelated, jpeg artifacts, noise, grain, poorly drawn, deformed, watermark, text overlay
Portrait restoration (add to universal):
bad anatomy, distorted face, oversmoothed skin, plastic skin, airbrushed, unnatural skin texture, bad eyes, extra fingers, deformed hands
Architecture / product (add to universal):
distorted lines, incorrect perspective, warped geometry, lens barrel distortion, chromatic aberration
Two things worth knowing about negative prompts:
-
Shorter is better than longer. The 60-term negative prompts that circulate on Reddit are largely cargo cult. On SDXL especially, an overstuffed negative visibly flattens contrast and desaturates colour, because you are pushing the sampler away from a huge region of latent space. The lists above are deliberately trimmed.
-
On Flux, the negative prompt does nothing. Flux dev and schnell are guidance-distilled and run at CFG 1, where there is no negative branch to apply. You need a true-CFG sampler node (CFG above 1 plus a real negative conditioning path) before a negative prompt has any effect, and that roughly doubles generation time. If you are on Flux and your negative prompt seems ignored — it is being ignored.
Settings that actually control the result
img2img / tiled upscale pass
| Setting | Recommended | Why |
|---|---|---|
| Denoising strength | 0.2 – 0.35 | The main lever. Below 0.2 barely changes anything; above 0.4 the model reinterprets. |
| Sampling steps | 25 – 40 | Past ~40 you are paying time for negligible gain. |
| CFG scale | 4 – 7 (SDXL), 7 – 9 (SD 1.5) | SDXL wants lower CFG than SD 1.5. High CFG on an upscale bakes in artifacts. |
| Sampler | DPM++ 2M Karras / DPM++ 3M SDE | Consistent, low-drift on low-denoise passes. |
| Scale per pass | 2× | Two 2× passes beat one 4× pass, reliably. |
| Tile size | 512 – 1024 | Smaller tiles = less VRAM, more seam risk. |
| Tile overlap | 64 – 128 | Raise this first if you see grid seams. |
| ControlNet Tile | On, weight 0.5 – 0.8 | Conditions each tile on the source. Prevents hallucination. |
Denoise cheat sheet by task:
| Task | Denoise | Notes |
|---|---|---|
| Clean image, add sharpness | 0.15 – 0.25 | Faces stay identical. |
| General photo upscale | 0.25 – 0.35 | The default working range. |
| Blurry source, needs reconstruction | 0.35 – 0.45 | Faces start to drift; check identity. |
| Damaged / heavily degraded | 0.45 – 0.6, or use SUPIR | Expect reinterpretation, not restoration. |
Old photo restoration additions
- Denoise 0.35 – 0.5 to allow artifact repair
- Add to positive:
restored photograph, damage repaired, reconstructed detail, period-accurate color - Add to negative:
scratches, fading, color cast, torn, creases, dust - Consider SUPIR instead of a standard img2img pass — it is purpose-built for degraded input
Upscaler models: which one for which job
This is lever #2, and it changes output more than any keyword you can type.
| Upscaler | Best for | Character |
|---|---|---|
| 4x-UltraSharp | General photos, the safe default | Crisp without over-sharpening. Community favourite for good reason. |
| Real-ESRGAN x4plus | Photography, reliable baseline | Slightly softer, very few artifacts. |
| Real-ESRGAN x4plus_anime_6B | Anime, illustration, flat colour | Keeps line art clean; do not use on photos. |
| 4x_NMKD-Siax / Remacri | Skin, fabric, organic texture | Remacri is punchier; Siax is gentler on faces. |
| SUPIR | Badly damaged or very low-res sources | Diffusion-based restoration. Heavy VRAM, best-in-class results on hard input. |
| DAT / SwinIR | Fine detail preservation | Transformer-based, slower, excellent on architecture and text. |
| ESRGAN 4x (original) | Legacy fallback | Superseded by everything above. Keep only for compatibility. |
| Latent (bicubic etc.) | Never, for photos | Needs high denoise to look right, which defeats the purpose. |
Most of these are downloadable from OpenModelDB and drop into models/ESRGAN/ (A1111/Forge) or models/upscale_models/ (ComfyUI).
Complete recipes
Old photo restoration
Positive:
RAW photo, photo restoration, tack sharp, highly detailed, photorealistic, damage repaired, reconstructed detail, period-accurate color palette, extreme fine detail
Negative:
blurry, low resolution, scratches, fading, color cast, damage, artifacts, noise, cartoon, illustration, distorted face, bad anatomy
Settings: denoise 0.4 · ControlNet Tile 0.7 · upscaler SUPIR, or 4x-UltraSharp if VRAM is limited · 2× per pass
Blurry photo to sharp 4K
Positive:
RAW photo, tack sharp throughout, extreme fine detail, photorealistic, crisp focus, professional photography quality, rich color depth
Negative:
blurry, soft focus, out of focus, low resolution, pixelated, jpeg artifacts, noise, grain, poorly drawn
Settings: denoise 0.3 · ControlNet Tile 0.7 · upscaler 4x-UltraSharp · two 2× passes to reach 4K
Portrait face restoration
Positive:
RAW photo, tack sharp, pore-level skin detail, natural skin texture, highly detailed eyes and hair, photorealistic, professional portrait photography, true-to-life color
Negative:
blurry, low resolution, bad anatomy, distorted face, oversmoothed skin, plastic skin, airbrushed, bad eyes, jpeg artifacts, watermark
Settings: denoise 0.22 — keep it low, faces drift fast · ControlNet Tile 0.8 · upscaler 4x_NMKD-Siax · check identity after every pass
Anime / illustration upscale
Positive:
masterpiece, best quality, highly detailed, clean line art, vibrant colors, sharp linework, detailed background
Negative:
blurry, low quality, worst quality, jpeg artifacts, sketch, messy lines, watermark, signature
Settings: denoise 0.3 · upscaler Real-ESRGAN x4plus_anime_6B · quality tags work properly here — this is the dialect they were trained on
Troubleshooting
Faces change identity after upscaling. Denoise too high. Drop to 0.2-0.25 and raise ControlNet Tile weight. If it still drifts, upscale the background at higher denoise and composite the face back at low denoise.
Visible grid seams. Increase tile overlap to 128, enable seam fix in Ultimate SD Upscale, or reduce scale per pass.
New objects appearing — extra windows, extra fingers, invented text. ControlNet Tile is off or weighted too low. This is exactly the failure it exists to prevent.
Output looks over-sharpened and crunchy. CFG too high for the base model, or you have stacked both an aggressive GAN upscaler and a high-denoise diffusion pass. Lower CFG first.
CUDA out of memory. Reduce tile size to 512, enable tiled VAE, lower the per-pass scale factor. Tile size is the main VRAM knob in tiled upscaling.
Quality keywords seem to do nothing. Check your base model against the dialect table above. On Flux and SD 3.5 this is expected behaviour, not a bug.
Upscale is slow. Steps above 40 and 4× single passes are the usual culprits. Two 2× passes at 30 steps is typically faster and better than one 4× pass at 60.
Stable Diffusion vs ChatGPT for this job
These are genuinely different tools, and the honest comparison matters more than a keyword list:
| Stable Diffusion | ChatGPT | |
|---|---|---|
| True pixel upscaling | Yes — dedicated upscaler models | No — re-renders at fixed resolution |
| Instruction following | None. Steer with settings, not sentences. | Yes. "Sharpen only the face" works literally. |
| Negative prompt | Yes (except Flux) | No |
| Preservation control | Denoise + ControlNet Tile | Written instructions |
| Batch processing | Excellent | Manual, one at a time |
| Setup cost | Local GPU, hours of configuration | None |
| Best at | Repeatable, controlled, high-volume upscaling | One-off restoration described in plain words |
The practical split: if you are upscaling fifty images to a consistent look, SD is the right tool and this guide is your reference. If you are restoring one family photo and want to describe what you need in a sentence, ChatGPT is faster and the results are frequently better — because instruction-following is exactly what SD lacks.
For the ChatGPT side of that workflow, the tested prompt templates are here: ChatGPT 4K Photo Enhancement Prompts.
Related resources
- ChatGPT 4K Photo Enhancement Prompts — 10 copy-paste templates for enhance, upscale and restore
- ChatGPT Upscale Image Prompt Guide — the upscaling-specific companion
- ChatGPT 8K Ultra HD Prompts — maximum-quality generation prompts
- Technical Prompt Keywords for High-Fidelity 8K Upscaling — cross-tool keyword reference
- AI Upscale Image to 4K: Sample Prompts — generative vs dedicated upscaling, compared
- Browse all image prompt templates — 140+ ready-to-use prompts by category

