Unleashed Forge
Updates every 2 s while this tab is open.
Text to image
Wan 2.2 video on a rented GPU
No cloud server is running.
ComfyUI engine
Qwen Image Edit — one picture, one instruction
The face is cropped from the main picture and sent as a lock automatically. You can still add 1 extra photo (pose, outfit). The main Picture stays the canvas.
Say what changes; everything else stays. 4 fast ≈ 15 s. 8 ≈ 30 s. Full = 20 steps, cfg 4 ≈ 70 s. 40 = no Lightning, cfg 1 ≈ 2 min — same edit strength as 4/8, more skin detail. Seed -1 = random (try again if clothes stay). A working seed is in the filename and here after each run.
Turnaround of the same character: one edit per view, then a sheet. Uses the 2511 Multi-Angles LoRA at strength 1. No extra face reference (that was forcing both sides to look like the front). The instruction box is ignored.
TRELLIS.2 — a picture or a prompt becomes a 3D object for the Quest
One object or one character, seen whole, works best. With a picture the prompt is ignored; without one, Flux paints the object from the prompt first.
Standard ≈ 1–2 min. Fine = finer grid (1536 for objects, 384 octree for people), sharper edges, hair and fingers hold together, 2–3× longer. The background is cut out automatically. AI paint redraws the sides and back from the photo so nothing stays grey, about a minute more. Object = TRELLIS.2 (things). Person = HumanNOVA (people; a half-body photo gives the best face), then the photo is painted onto the front.
Tap a view. Erase wipes pixels that should not go on the model. Paint puts the original photo back only where you brush — the rest stays as you left it. Then Preview and OK · paint. AI paint first redraws the whole model from the front cut-out (the sides and back match it, no grey), then your photos go on top; about a minute more.
Videos, images and voice lines
Voice lines for videos — cloned with F5-TTS
Speed applies to F5 only. F5: tone follows the reference clip and punctuation. IndexTTS-2: pick an emotion above (slower, ~1 min a line). Audio8: 0.6B, Cantonese + 10 languages, tone follows the reference.
Clones the reference voice saying your line (F5-TTS). The result is set as Voice automatically — pin-played in Animate, cloned in Keep face.
MiniMax H3 video
Loading a different checkpoint takes a moment.
Tap a LoRA card to add it; its trigger words appear here — tap one to add it to the prompt. 0.6-0.8 keeps the checkpoint detail.
…
Rounded to multiples of 8. A ratio button keeps the long edge and resizes the short one.
ADetailer finds faces after generate and redraws each one at 512px — that is what keeps a full-body shot from collapsing the face. GFPGAN is a cheap filter that smears tiny faces; leave it off when ADetailer is on. Show face quietly adds "looking at viewer" and pushes cropped/turned-away heads into the negative.
Batch runs the images one after another, so it needs no extra VRAM — it just takes that many times longer. Each one is saved as it finishes, so Stop keeps everything already done.
Seed -1 is random. Reuse a seed from History to repeat an image.
8 GB GPU: 1.5x max from 1024 square. Denoise 0.35-0.45 keeps the face.
Every generation from this phone and from Forge on the PC.
Flux: drop .gguf or bare-transformer .safetensors files into sd-models\checkpoints\flux on the PC and they appear here (safetensors run as fp8).
Only flux LoRAs are listed — the SDXL ones do nothing on flux. FLUX.1-Turbo-Alpha at 1.0 lets 8 steps look like 20.
Same cards as Forge: tap a card to weight it, hold and drag to reorder. No negative — flux ignores it at CFG 1 — and skip the score_9 Pony tags; plain descriptions steer flux best.
Rounded to multiples of 16.
Euler / Simple, CFG locked to 1 — flux is guidance-distilled; the Guidance slider is its steering. 3.5 is the sweet spot.
ReActor swaps the face onto every image as it generates, then restores it with GFPGAN — no separate Roop pass. Face models are built on the PC: open Forge, ReActor panel → Tools → Face Models, and Blend several photos of the same person for the best likeness.
Upload an existing photo. Forge skips redrawing the whole picture and only ADetailer inpaints detected faces at 512px — the rest of the image stays as-is. Uses the current checkpoint and the ADetailer detector / denoise from Settings.
instant_id_face_embedding → ip-adapter_instant_id_sdxl · best 0.8–1.0
instant_id_face_keypoints → control_instant_id_sdxl · best 0.6–0.8. Leave the photo empty to reuse the Unit 0 photo.
Steers generation toward the face in your photo — identity from Unit 0, head pose from Unit 1. Works with SDXL / Pony checkpoints on the Forge engine.
IP-Adapter: the model borrows the look of this image — rendering, colours, texture, the "3D donghua" feel — without copying its subject. SDXL / Pony / Illustrious checkpoints.
0.5–0.7 takes the style and keeps your prompt in charge; above 0.9 the reference starts to dictate composition too.
Works with Flux? No — the adapter is for SDXL-family checkpoints. Combine with a style LoRA and the 3D-donghua tag group for the closest match.
The name is the trigger word: write it in the H3 prompt when the LoRA is on.
10-30 photos of one person: different clothes, angles, places and lighting; the face sharp. Only people you have the right to use.
About 3-5 s per step on this card, plus ~10 min of encoding. Forge and ComfyUI are closed while it trains.
LoRAs land in ComfyUI's loras/h3 and show up as pills on the video page.
Stop keeps the disk: models stay, you pay only storage (a few dollars a month), and it restarts in about a minute. Destroy wipes it — nothing is billed after, the next one installs from scratch (~10 min).
Max $/h on the right. Single GPU, ≥48 GB RAM, ≥80 GB disk, fast download. Verified hosts are vetted by Vast — prefer them: the machine owner has physical access to anything you upload.
More passes strengthen the likeness but can look waxy. Turn off "swap every face" to change only the largest face.
Restores detail after the swap. Below about 0.6 it blends with the original skin instead of repainting it.
Negative is younger, positive older. Leave both at 0 to skip the age model entirely, which is faster.
Switching opens that address. If it does not answer you will get the browser's own error, so keep the current one until the new address is proven.
Automatic follows your iPhone appearance setting.
8 s is beyond the model's 5 s training length: 480p only, may run out of memory, and the picture can drift near the end — Continueis the safer way to go long. Continueon a finished clip adds another segment from its last frame and saves the joined video as a new file (the original stays). Forge is stopped while the clip renders (Wan 2.2 needs the whole card) and restarted when it is done. 480p · 3 s takes a few minutes; 720p several times longer. Forge pauses during generation and restarts when done.