BFS Head V1.1: Qwen Image 2.1 In-Context Face & Head Swap (Surgical Skin-Detail Ablation)
Full Resolution
Generation Specs RTX 4070
Sampler euler
Scheduler simple
Steps 12
CFG Scale 2.5
Seed 85244748616565
Target VRAM 12GB
VRAM 10GB - 12GB
GPU Hardware Compatibility
Optimal Ground Truth

Native GPU execution at full throughput with zero memory swapping.

Qwen Image 2.1 (DiT) Min 10GB VRAM 832×1248

BFS Head V1.1: Qwen Image 2.1 In-Context Face & Head Swap (Surgical Skin-Detail Ablation)

Open-source head-swapping architecture powered by Qwen Image 2.1 DiT and Alissonerdx's BFS V1.1 LoRA. Features surgical img_mlp.gate_up weight ablation to eliminate plastic skin softening while preserving 100% 3D head pose and complex ambient lighting integration.

Blueprint Summary RTX 4070 Verified

Reproducible ComfyUI workflow for BFS Head V1.1: Qwen Image 2.1 In-Context Face & Head Swap (Surgical Skin-Detail Ablation) using Qwen Image 2.1 (DiT) at 832×1248 resolution. Requires minimum 10GB VRAM with sampler euler and scheduler simple (12 steps). Includes 1-click terminal model sync and canvas JSON graph.

Node Graph Pipeline 10 Total Nodes
Verify in Resolver
01 Load Models
UnetLoaderGGUF
4 loaders (DiT, CLIP, VAE)
02 Adapters
LoraLoaderModelOnly
1 adapter
03 Conditioning
TextEncodeQwenImage21
1 prompt encodings
04 KSampler
KSampler
DiT latent denoising
05 Decode & Save
VAEDecode
Latent to pixel space

Model & Asset Setup 1-Click Script

Run in your ComfyUI root:

curl -fsSL https://comfy.yellorn.com/api/scripts/bfs-head-swap-v1-1-qwen-image-2-1.sh | bash

Positive Prompt

head_swap: start with <image1> as the base image, keeping its lighting, environment, and background. remove the head from <image1> completely and replace it with the head from <image2>, but maintain the aspect ratio of the head from <image1>, maintain the direction of the eye, head rotation, skin color, micro expressions from <image1>. high quality, sharp details, 4k

LoRA Adapter Stack 1 Adapters

RTX 4070 Calibrated Weights
01 bfs_head_v1.1_qwen_2.1.safetensors +1

Required Models 5 Models

Disk Space Required: 13.9 GB (5 models · DiT/Base: 8.5 GB · Text Encoder: 4.5 GB · LoRAs: 568 MB · VAE: 350 MB)
models/loras/ 248 MB HuggingFace
bfs_head_v1.1_qwen_2.1.safetensors Weight: 1.0 (Blocks 16-31 img_mlp.gate_up pruned to preserve natural skin micro-texture)
models/loras/ 320 MB HuggingFace
p_qwen_image_2.1_8step_v0.1.safetensors Weight: 1.0 (Pruna AI 8-step Lightning acceleration LoRA)
models/diffusion_models/ 8.5 GB HuggingFace
qwen_image_2.1_int8_convrot.safetensors Qwen Image 2.1 DiT INT8 quantized core
models/text_encoders/ 4.5 GB HuggingFace
qwen3vl_8b_int8_convrot.safetensors Qwen3-VL Vision-Language Text Encoder (CLIP loader type: qwen_image)
models/vae/ 350 MB HuggingFace
qwen_image_2.1_vae_bf16.safetensors Qwen Image 2.1 Official BF16 VAE

Field Notes RTX 4070 Benchmark

Benchmark: 87.5s on RTX 4070 (10.6 GB VRAM peak, Euler simple, 12 steps, 832x1248)

Reverse-Engineered & Synthesized In-Context Swapping Architecture: Built upon Alibaba's Qwen Image 2.1 Diffusion Transformer (DiT) paired with Alissonerdx's breakthrough BFS Head V1.1 LoRA (248 MB). Unlike legacy embedding-pasting tools (InsightFace, ReActor, RoOP) that suffer from boundary seams and flat plastic skin, BFS performs native in-context latent regeneration. Key breakthrough: The V1.1 release executes surgical weight pruning on the second-half network (blocks 16-31), disabling 16 img_mlp.gate_up projections that previously caused artificial skin smoothing, thereby recovering 33% bilateral micro-texture (pores, freckles, stubble) without retraining. The early MLP blocks (0-15) and all 128 attention projections are byte-identical to V1, preserving full 3D head rotation, eye gaze, and subtle micro-expressions. Hardware ground truth verified on dedicated RTX 4070: 10.60 GB peak VRAM, 87.5s generation time for full resolution native in-context editing. Input Contract: Image 1 must be the Target Scene/Body (composition & lighting anchor), while Image 2 is the Reference Identity.

Frequently Asked Questions FAQ

What GPU and VRAM are required to run BFS Head V1.1: Qwen Image 2.1 In-Context Face & Head Swap (Surgical Skin-Detail Ablation)?

This workflow requires a minimum of 10GB VRAM (recommended 12GB VRAM). Tested and verified on NVIDIA GeForce RTX 4070 (12GB VRAM) at 832x1248 resolution.

How do I resolve missing custom nodes for this workflow?

You can drop the workflow JSON into our client-side Missing Node Auto-Resolver at https://comfy.yellorn.com/resolve/ to detect missing nodes and generate install commands, or run the 1-click terminal setup script provided below.

What conditioning and VAE architecture does BFS Head Swap require?

BFS Head Swap V1.1 uses Qwen Image 2.1 DiT with dual VAE conditioning and in-context facial geometry alignment. It performs surgical skin-detail ablation without requiring external face-swapping models.

Action completed