My first attempt was one photo and a short prompt. It gave me a generic pretty face that wasn't mine at all.
The fix wasn't a better photo. It was writing my face out in words. Everything below is the rest of what I learned getting from that to this.
Disclosure: the Higgsfield link in this post is my creator referral link. Signing up through it supports my work, and it costs you nothing extra.
01 ~ The inputs
Pick your references
Boring photos make the best references. The model copies what it can actually see ~ it cannot invent the shape of your jaw from a blurry photo taken at a party, so give it flat, well lit, unremarkable pictures of your face.



Nothing special about these. Golden hour, my phone, a Spain jersey, no ring light. That is the point.
Two or three angles, not one
Front, three-quarter, and one where your head is turned. One photo gives the model a mask. Three gives it a head.
Even daylight
Golden hour or open shade. Hard flash flattens the features it needs most, and dim light gives it nothing.
Same hair era
All three from the same week if you can. Mixing a bob and long hair makes the output average them into something that is neither.
Nothing covering your face
No sunglasses, no hand on the chin, no heavy filter. Keep the jewelry you always wear, because that is a free identity anchor.
02 ~ The part nobody does
Write an identity lock
Describe your own face like a police sketch artist. Most image models weight your typed words far more heavily than your uploaded photo ~ so if you don't describe your face, the model fills the gap with its own idea of a face, and you get a stranger who is vaguely you.
IDENTITY LOCK - the subject is the woman in the uploaded reference photographs. Match her face exactly: [age] [descent], [skin tone and quality], [face shape], [eye shape and color], [nose], [lips], [hair color, length and how it is styled right now]. [Body build]. Real skin texture - visible pores, faint freckling, no smoothing, no beauty-filter plastic. [Any jewelry you always wear]. Face unmistakably hers.
Features, not compliments
High cheekbones, low lid crease, small straight nose. Not pretty, not cute ~ those words mean nothing to the model and it will default to a stock face.
Say the unflattering parts
Visible pores, faint freckling, no smoothing. Real skin is the single fastest way to stop looking generated.
Keep your jewelry in it
My gold chain, thin bracelet and the black band on my wrist are in every prompt. Small consistent objects hold an identity together across generations.
03 ~ The skeleton
Eight blocks, every single time
Every long prompt I write has the same skeleton. Once you see the structure you stop writing prompts from scratch ~ you swap the blocks. Change wardrobe and setting and the same skeleton gives you a subway shot, a Coney Island shot, or something with no spider in it at all.
Your face, written out in words. Reference photos alone are not enough, because most models weight the text far more heavily than the image.
RuleDescribe features, not vibes. 'High cheekbones, low lid crease, small straight nose' beats 'pretty girl' every time.
04 ~ Verbatim
The full prompt
Exactly what I ran. The only line you must change is the identity lock, because right now it describes me. Everything else works as is.
Hyper-realistic action photograph, full-frame DSLR, 35mm lens, f/2.8, 1/1000s shutter, vibrant saturated color grade, golden-hour sunlight with rich warm and cool contrast, cinematic realism, fine film grain, tack-sharp, high micro-detail, 8K. One consistent lighting environment across the whole frame. IDENTITY LOCK - the subject is the woman in the uploaded reference photographs. Match her face exactly: a young woman in her early twenties of Southeast Asian descent, warm light-tan skin with a natural sun-flushed glow, softly rounded face with high cheekbones and a delicate jaw, dark warm-brown almond eyes with a low lid crease, straight small nose, full natural lips, dark brown-black hair in a low ponytail with loose strands falling forward around her face. Slim, petite, athletic build. Real skin texture - visible pores, faint freckling, no smoothing, no beauty-filter plastic. Small fine gold chain necklace, thin gold bracelet, black band on the wrist. Face unmistakably hers, seen from a three-quarter downward angle. POSE - LOOKING DOWN, COILED CROUCH, not standing, not facing camera: she is perched in a tight low crouch on the narrow edge of a rooftop ledge / fire-escape railing, weight forward over the drop, one gloved hand gripping the ledge between her feet, the other hand hanging loose with a web line dangling from her fingertips. Head tilted DOWN, gaze locked on the street far below, chin tucked, eyes lowered - completely unaware of the camera, no eye contact, no smile. Loose strands of hair fall forward past her cheek. Shoulders coiled, spine curved, calves tensed - full kinetic potential energy, about to drop. Camera positioned slightly above and to the side, so we read the top of her head, the line of her cheekbone and jaw in three-quarter, and the vertigo of the street below her. Full body visible, placed low and off-center with the drop opening beneath her. Hands clean and correctly formed. MASK OFF - hood up, framing her face; the white spider mask with large magenta-pink teardrop reflective lenses is held loosely in her hanging hand, dangling by its edge, faint web texture visible on the fabric. Her real face fully exposed and sharp. WARDROBE - black textured scaled bodysuit torso with a deep white-edged V-neckline; pink-and-white web-patterned sleeves and gloves; grey sweatshirt tied at the hips; baggy charcoal-denim cargo shorts over the suit; pink-and-white web-pattern leggings; teal satin ballet pointe shoes with crisscross ribbon laces, toes gripping the ledge. Black over-ear headphones around her neck. Hood and loose fabric stirring in the wind. SETTING - vibrant, colorful Brooklyn / New York at golden hour: seen from high on a Brooklyn rooftop, the Manhattan skyline glowing across the East River, the Brooklyn Bridge catching orange sunset light, cool blue shadows filling the street canyon below, colorful graffiti murals on brick, neon bodega signs starting to glow, tiny cars and figures far below for scale, bright blue-and-gold sky with dramatic clouds. Rich, punchy, real NYC energy - vivid but not garish. WEBBING - MASSIVE, everywhere, fine dense silk (NOT ropes/cables): enormous amounts of tangled spiderweb engulfing the frame, made of fine, thin, densely layered silk threads and gauzy web sheets - thousands of delicate strands crossing at every angle, layered veils draping from the fire escape, water towers, lampposts and rooftop edges, sheer curtains of thin silk stretched across the drop below her, dense fine webbing arcing overhead into a chaotic canopy, single gossamer threads catching the low sun. Fine silk glowing in the golden light with iridescent glints. Layered at every depth. Uneven organic chaos - dense matted patches mixed with single floating strands. Explicitly fine spider silk, NOT thick ropes, NOT braided cables, NOT chunky strands. Crisp readable threads, NOT gray fog. NO clean empty space. NOT symmetrical, NOT tidy. NO spider, NO creature. COMPOSITION - off-center, strong vertical diagonal, dizzying sense of height with the street falling away beneath her, subject tack-sharp against a softly falling-off background. Clear focal hierarchy. Candid observational framing, as if photographed by someone who spotted her from an adjacent rooftop - not a posed portrait. Photojournalistic realism - crisp tactile textures (fine spider silk under tension catching light, rusted railing, brick, scaled suit fabric, denim, satin pointe shoes). Unified golden-hour light so nothing looks composited. Vibrant saturated New York color, grounded, tense, high-detail. Aspect ratio 4:5, vertical.
looking at camera, eye contact, front-facing, posed smile, standing still, static flat pose, symmetrical centered composition, mask covering face, thick ropes, braided cables, chunky rope web, thin sparse web, clean empty space, gray fog web, foggy mush web, dull flat color, washed out, muddy textures, generic city, inaccurate landmark, blurry subject, low detail, symmetrical web, spider, creature, stiff posing, polished glossy look, airbrushed skin, beauty filter, garish oversaturated neon overload, sticker/collage look, pasted elements, mismatched lighting, cartoon, 3D render, CGI, plastic skin, distorted hands, extra fingers, wrong face, generic face, watermark, text artifacts, AI artifactsRead the negative prompt twice. It is not filler. Half the work of making this look real happens in there, deleting the posed smile and the eye contact and the plastic skin before the model ever reaches for them.
05 ~ Settings
Run it and expect to reroll
Upload all three photos
Not one. Two or three angles of the same face, same hair era, gives the model enough to triangulate.
Tell it what each image is for
Say it plainly: use the photographs for her face and build, use the text for everything else.
Set 4:5 and 2K
Vertical for feed and stories. 2K if you want it to survive a crop.
Expect three or four rolls
Hands and the taut web line are where it breaks. Regenerate rather than trying to fix it in the prompt.
Hands are still the hard part
One hand gripping a ledge and one holding a mask is a lot to ask. When the hands come out wrong, roll again rather than editing the prompt. The prompt is not the problem, the dice are.
06 ~ The actual secret
Four lines that stop it looking like AI
If you take nothing else from this page, take these. They are cheap to add and they do more for realism than any camera setting.
01
Kill the eye contact
A subject beaming straight down the lens is the number one tell. Every prompt I write says looking down or side profile, no eye contact, no smile, unaware of the camera. It goes in the negative prompt too.
02
Ask for real skin
Visible pores, faint freckling, no smoothing, no beauty-filter plastic. Then negate airbrushed skin and beauty filter. Models default to a retouched magazine face and it looks fake instantly.
03
Give the camera an owner
Candid observational framing, as if someone caught her from an adjacent rooftop. The moment the frame has a reason to exist, it stops looking like a generated poster.
04
One light source, always
One consistent lighting environment across the whole frame. Mixed lighting is what makes an image look composited, like the subject was cut out and dropped in.
07 ~ Proof
Swap two blocks, get a different shoot
Same identity lock, same camera block, same negative prompt. I changed the setting and the pose, and the rooftop became a subway platform. That is the whole argument for writing prompts in blocks.


08 ~ Your turn
Go make your own version
Copy the prompt, swap my description for yours, run it. My link unlocks unlimited generations on the top models for 24 hours ~ long enough to burn through every roll it takes to get the hands right.
Claim 24h unlimitedMade in Madrid · shot on a phone · set in Brooklyn

Love,
Abie



