STATISTICALLY WRONG - Preserving intentional distortion in a generative image pipeline
Trained a LoRA on 84 photographs of my face, then used ComfyUI and ControlNet to force that likeness into frames from my own clay stop-motion—testing where physical form, identity, and generative interpretation begin to break apart.
01 — PHYSICAL SOURCE
02 — LoRA DATASET
03 — IDENTITY MODEL
04 — CLAY TRANSFER
What I make
I'm a sculptor and filmmaker working primarily in clay. I build faces frame by frame and shoot them as stop-motion, then push the footage between the physical and digital until it's hard to tell which register you're in. The recurring subject is memory and the unfinished: forms surfacing from a void, a face that never fully resolves, the fingerprint left visible as proof a hand was here.
This case study grew out of that practice and a larger investigation into AI image models. I'm interested in what happens when a deliberately imperfect physical object enters a system trained on an enormous aggregate of existing images. If the distortion is intentional, can I make the model preserve it—or will it interpret that specificity as an error to correct?
The Experiment
Train a model on my own face, then use a frame from my clay stop-motion to control its form.
The clay is intentionally distorted. The test is whether I can preserve that distortion while forcing my likeness and human skin into it—or whether the model will interpret those decisions as errors and pull the image toward a more familiar face.
The source material
The input is real clay I sculpted and shot as a stop-motion sequence: a rough, deliberately distorted face emerging from a blob, roughly 280 frames.
Training My Likeness
Using FluxGym, I trained a LoRA on 84 photographs of my own face in a range of neutral expressions.
BuildinG the pipeline
FLUX.1-dev — Base model
The underlying image model. It provides the broader visual vocabulary the experiment is working against.
T5 + CLIP — Text conditioning
Translate my written prompt into information the model can use to steer the image.
My LoRA — Identity
Trained in FluxGym on 84 photographs of my face. This pushes the output toward my likeness.
Clay Frame — Structure
A frame from my stop-motion enters through image-to-image, while Canny ControlNet extracts its edges and helps hold the sculpture's distorted silhouette and proportions.
So the system is balancing three competing instructions: my likeness, the physical structure of the clay, and the base model's learned expectations of what a human face should look like.
Where it broke
This is what I actually wanted to document, because the failures are more revealing than the successes.
Each used the same positive text prompt:
“ohwx man, extreme close-up photograph, living human skin, glossy oily skin, visible pores, soft studio lighting, shallow depth of field, subsurface scattering”
It idealized. Without enough structural control, the model pulled the distorted clay toward a symmetrical, conventionally beautiful face. The deformation wasn't lost randomly; it was being replaced by a more statistically familiar solution. (model strength=1, denoise=.85, no control net node)
It read sculpture instead of skin. At low denoise, the structure survived but the material transformation didn't. The result remained essentially gray clay with superficial skin detail. (model strength=.95, denoise=.45, no control net node)
Edge control isn’t form control. Raising denoise produced more convincing skin, but the Canny ControlNet only constrained the image's edges—not how those edges had to exist in three-dimensional space. The model reinterpreted the clay's contours as a flat opening with an eye peering through. (model strength=.95, denoise=.95, control net strength=.5)
It over-translated. Push the model and denoise strength too far, and the clay's surface begins to resolve as damaged skin rather than natural flesh. (model strength=.95, denoise=.95, control net strength=.7)
Identical settings didn't mean identical behavior. Changing only the seed produced substantially different interpretations of the same clay frame. The parameters define a range of behavior, not a single deterministic result.
finding the narrow band
The most interesting range sits in a narrow band between two failures: enough structural lock to keep the clay's intentional distortions, enough freedom for the material to read as flesh.
(model strength=1, denoise=.65, control net strength=.7)
What the experiment exposed
I own the specific inputs—my face, my sculpture, my photographs—but they operate through a base model trained on a visual corpus I can't inspect.
That became most visible when the model resisted the sculpture's intentional distortion. I knew exactly what form I wanted, but had only text, image conditioning, and model parameters with which to communicate it. What I considered an artistic decision could become statistically legible to the model as an error to correct.
That tension is more interesting to me than whether the model can successfully reproduce my face: how do you communicate intentional wrongness to a system trained on visual probability?
Next Tests
The next test is consistency over time. A single frame can find the balance between the clay's structure and the model's interpretation; a sequence has to maintain that balance across changing forms without identity, material, or detail drifting between frames.
The goal is to determine whether the same controls can preserve the sculpture's intentional distortions across an animated sequence—not just produce one successful image.