AI VIDEO
Fünf Memes, ein Spaziergang, kein Schnitt.
Ich laufe eine Pariser Straße entlang und treffe nacheinander fünf Gesichter, die jeder aus dem Netz kennt. Kein Schnitt, keine Montage: acht Referenzbilder liegen in Runway, ein einziger Prompt teilt ihnen ihre Rolle und ihr Zeitfenster zu, und Seedance 2.5 rendert daraus 23 Sekunden Handyvideo am Stück — inklusive Ton.
- ChatGPT
- Gemini
- Runway
- 45 Min
- Fortgeschritten
ERGEBNIS
Das ist dabei rausgekommen.
- 8Referenzbilder
- 1Video-Prompt
- 23sein Take, kein Schnitt
- 9:161080p aus Seedance 2.5
-
Der Schläfen-Tipp0:00 – 0:04
-
Der Blick zurück0:05 – 0:09
-
Das Lachen am Cafétisch0:10 – 0:14
-
Der Hund, der nicht wegsieht0:15 – 0:19
-
„Hi.“ — „Okay.“0:20 – 0:23
DAS BRAUCHST DU
Womit das gebaut ist.
ChatGPT
GeminiRunway
-
@image1 · VloggerCharakterbogen, zwei Panels
-
@image2 · LocationVordergrund komplett frei
-
@image3 · Schläfen-TippGeste und Grinsen
-
@image4 · Drei-Personen-SzeneNur das Paar, die Dritte bin ich
-
@image5 · Der LacherStudio wird ignoriert
-
@image6 · Der HundDer Blick, nicht das Wohnzimmer
-
@image7 · Das LächelnHemd, Rucksack, Haltung
-
ohne Slot · das KindHochgeladen, nie adressiert
ANLEITUNG
So baust du es nach.
-
Einen Charakterbogen von dir bauen
Ein normales Foto von dir in ChatGPT hochladen und Prompt 1 einsetzen. Heraus kommt ein Bogen mit zwei Panels: links Brustbild, rechts Ganzkörper, weißer Hintergrund. Das Outfit im Prompt ist das Outfit im Video — bei mir steht dort noch weißes T-Shirt und Jeans, benutzt habe ich am Ende die Version im dunklen Pullover. Trag ein, was du wirklich tragen willst, und beschreib es im Video-Prompt später noch einmal wörtlich.
-
Die Location generieren
Prompt 2 in Nano Banana 2 (Gemini) geben. Wichtig sind drei Stellen darin: die Caféterrasse rechts, der freie Vordergrund und das Negative-Prompt gegen lesbare Schriftzüge. Der freie Vordergrund ist der Grund, warum später Platz für die Begegnungen ist; lesbare Ladennamen wären in jedem Frame ein Wort, das das Videomodell neu erfindet.
-
Die Vorlagen hochskalieren
Die Vorlagen, die du treffen willst, liegen meistens als kleines, weich komprimiertes JPEG vor. In Runway einmal durch Upscale Image schicken, dann steht das Gesicht in einer Auflösung da, mit der das Videomodell etwas anfangen kann. Bilder mit Rand oder Balken vorher zuschneiden — der Rand ist sonst Teil der Referenz.
-
Alles als Referenzen laden und die Slots vergeben
In Runway links auf Tool, oben den Video-Tab, dann Reference. Alle Bilder in einem Rutsch hochladen, der Zähler zeigt danach 8/30. Die Reihenfolge dort bestimmt, welches Bild im Prompt @image1, @image2 und so weiter ist — der Block REFERENCE ASSIGNMENT ganz oben im Video-Prompt muss exakt dazu passen.
-
Video-Prompt einsetzen und die @-Chips verbinden
Den kompletten Video-Prompt ins Textfeld kopieren. Danach jedes @imageN löschen und an derselben Stelle ein @ tippen: Runway öffnet die Liste der Referenzen, du wählst das Bild, und aus dem Text wird ein Chip mit Vorschaubild. Bleibt irgendwo @image5 als reiner Text stehen, ist es kein Verweis, sondern eine Zeichenkette, die das Modell ignoriert.
-
Rendern und die Geografie prüfen
Unten rechts Seedance 2.5, darüber 1080p, Format 9:16, Dauer auf 25 Sekunden. Nach dem Rendern zuerst die Straße prüfen: Läuft er die ganze Zeit denselben Gehweg entlang, oder springt der Ort zwischendurch? Springt er, hat fast immer eine Referenz ihren eigenen Hintergrund mitgebracht — dann den Satz über die ignorierten Hintergründe verschärfen, statt die Szene umzuschreiben.
CREATOR KIT
Alles, was du brauchst.
Photorealistic character reference sheet on a plain seamless white background, two panels side by side of the same person @image1: LEFT — chest-up portrait facing the camera, relaxed natural expression; RIGHT — full-body standing pose, head to toe, arms relaxed at their sides. Keep the exact face, hair and skin from @image1 in both panels — same features, same hair length, texture and colour, same skin tone and texture, same natural asymmetry. Do not beautify, do not slim, do not change the age, do not smooth the skin. Ignore the background, lighting, clothing and camera angle of @image1. Wardrobe in both panels: a plain white crew-neck t-shirt, straight-leg mid-blue jeans, clean white low-top sneakers. Soft even studio lighting, natural skin and fabric texture, sharp focus. No props, no text, no labels.
Photorealistic photograph of a wide Paris street on a bright afternoon, shot straight down the sidewalk from pedestrian eye height, the pavement running away from the camera into the depth of the avenue. On the RIGHT: a café terrace running most of the block, small round tables and woven bistro chairs set out on the pavement, a few seated customers small and unremarkable in the middle distance, a dark green fabric awning stretched over the terrace, a low metal railing along its edge. Directly above that awning, the second floor of the building, with a stone-balustraded balcony running along it and two tall French windows standing open onto it — the balcony clearly visible and unobstructed. On the LEFT: a row of parked cars along the curb, a two-lane road beyond them, and the opposite side of the street closing the frame. Both sides: pale limestone Haussmann facades, wrought-iron balconies on every floor, zinc roofs, tall shuttered windows. Shopfront awnings in dark green and burgundy with no legible lettering anywhere. The pavement is wide pale stone slabs. The avenue runs to a bright open point at the far end. The immediate foreground of the sidewalk is completely clear and unoccupied. Lighting: high afternoon sun from the right, out of frame, striking the limestone and throwing short hard shadows across the pavement toward the left. Bright natural daylight at 5600K, open sky as a soft cool fill in the shade under the awning, warm bounce coming back off the pale stone. Sky highlights slightly clipped. Optics: 84 degree FOV wide, rectilinear, camera at eye height on the centre of the sidewalk, everything sharp front to back, f/8, no distortion of the vertical lines of the facades. Realism: real stone texture, weathering at the base of the walls, real glass reflections, worn paving, fine natural grain. Photographic, not rendered. Negative: legible text, shop names, street signs, brand logos, watermark, people in the foreground, empty deserted street, HDR oversaturation, tilted horizon, warped architecture, fisheye distortion.
Vertical 9:16 smartphone vlog. ONE SINGLE UNINTERRUPTED HANDHELD TAKE, approximately 25 seconds. No cuts, no montage, no hidden edits. Everything happens continuously while the male vlogger @image1 walks through the Paris street @image2 and encounters a sequence of strange people. REFERENCE ASSIGNMENT: @image1 = vlogger @image2 = Paris street @image3 = man touching his temple @image4 = three-person street reference @image5 = laughing moustached man @image6 = black dog with wide shocked eyes @image7 = smiling man in blue shirt @image3–@image7 are FACE, EXPRESSION, POSE AND GESTURE references only. Preserve the recognizable appearance and key expression/gesture of each reference, but ignore their original backgrounds, lighting, framing and camera angle. Rebuild them naturally inside @image2 and relight them with the same Paris daylight. @image1 remains visually consistent throughout: same face, hair, facial hair, skin, proportions and outfit shown in the reference. Dark knitted sweater, loose black trousers, brown shoes. He is the only person appearing throughout the video. LOCATION: The entire take happens on the SAME continuous sidewalk from @image2. Wide pale stone pavement, road and parked cars on the left, Haussmann buildings, café terrace and dark green awnings on the right. All encounters are positioned sequentially along his walking route, a few metres apart. The geography never changes. No environment from any other reference image appears. BACKGROUND: Normal pedestrians, traffic and café customers continue naturally. Nobody else treats the encounters as unusual. CAMERA: Authentic spontaneous smartphone vlog, not cinematic. Natural hand micro-shake, walking bounce, imperfect framing, slightly floaty stabilization, quick reactive pans, deep focus. Bright afternoon daylight, short shadows, slightly clipped bright highlights. Auto exposure visibly adapts between sunlight and café shade. Brief autofocus hunting after fast pans. Front camera = wide arm's-length selfie. Rear camera = slightly tighter. Digital zoom visibly loses quality and adds shake. The phone itself never appears because the phone IS the camera. AUDIO: Raw smartphone audio only: traffic, footsteps, café chatter, cutlery, wind and dialogue. No music, subtitles, captions, graphics or watermark. The vlogger should feel genuinely unaware of what he is about to encounter. His reactions are small, spontaneous and believable, never theatrical. 0:00–0:05 — FIRST ENCOUNTER FRONT CAMERA, selfie mode. @image1 walks forward casually while talking to his audience: “Okay, ich wollte euch eigentlich nur kurz zeigen—” While he speaks, @image3 approaches from the opposite direction and passes naturally beside him. @image3 wears the black leather jacket, black clothing, gold watch and overall appearance from the reference. As they draw level, @image3 looks directly toward the vlogger’s phone with the same thin knowing smirk from the reference. Without stopping, he raises his index finger and taps/points deliberately at his temple with the recognizable gesture. He holds the smug look toward the lens for a beat while continuing to walk past. The vlogger stops talking mid-sentence. He watches him pass in the phone screen and turns his head slightly after him: “…okay?” @image3 keeps walking away and never returns. 0:05–0:10 — SECOND ENCOUNTER SAME TAKE, NO CUT. Still walking forward, @image1 approaches the couple from @image4 coming toward him from the opposite direction. IMPORTANT: For this scene, use ONLY the boyfriend and girlfriend from @image4. The vlogger @image1 himself replaces the passing third person from the reference. Keep the FRONT CAMERA active so @image1 remains large in the foreground while the approaching couple becomes visible behind and beside him in the selfie frame. They pass each other naturally. Immediately after @image1 passes the couple, the boyfriend turns his head and upper body dramatically BACK toward @image1 and keeps looking at him. At the exact same moment, the girlfriend notices. She turns toward her boyfriend with the same shocked, offended, disapproving expression and body language from @image4. For approximately one second, the geometry becomes the recognizable three-person composition: @image1 walking ahead, boyfriend staring back at @image1, girlfriend staring angrily at boyfriend. @image1 notices this happening behind him through his selfie screen. He glances over his shoulder: “Bro?” The boyfriend is STILL looking. The girlfriend is STILL staring at him. @image1 gives the camera a confused look and keeps walking. The couple remains behind and never appears again. 0:10–0:15 — THIRD ENCOUNTER SAME TAKE. @image1 flips naturally to the REAR CAMERA and continues forward beside the café terrace. Suddenly loud uncontrollable laughter comes from the RIGHT. He reacts to the sound and pans toward it. At a café table beside the railing sits @image5. Use the same moustache, slicked-back hair, grey knitted sweater and recognizable laughing expression from the reference. Ignore the original studio completely. He sits naturally at a Paris café table with only a coffee and glass of water. He is completely consumed by laughter: mouth open, eyes narrowed, shoulders shaking, upper body rocking forward and backward. No dialogue from him. Only his characteristic uncontrollable laughter. @image1 slowly zooms closer. “Was ist denn so lustig?” The man continues laughing and gives no answer. People at the next table keep drinking coffee and talking as if nothing unusual is happening. Half-second confused hold. Then @image1 continues walking. 0:15–0:20 — FOURTH ENCOUNTER NO CUT. As the camera moves away from the laughing man, @image1 notices something lower down beside the next café table. He tilts the rear camera DOWN. @image6, the black dog from the reference, is sitting calmly beside a café chair on the pavement, attached naturally to a simple leash leading toward an unseen café customer. Preserve the dog's dark coat, long ears, eyes, muzzle and overall appearance from @image6. At first the dog looks away. As @image1 approaches, the dog slowly turns its head toward the camera. The moment it sees him, its face settles into exactly the extremely wide-eyed, slightly open-mouthed shocked expression from @image6. Then it becomes almost completely motionless. No barking. No jumping. No aggressive movement. No invented trick. Just an absurdly intense frozen stare directly toward the camera. @image1 slows down. Small digital zoom onto the dog's face. Long uncomfortable eye contact. @image1: “Warum guckst du mich so an?” The dog remains frozen in the exact same expression. @image1 slowly lowers the camera and walks on. 0:20–0:25 — FINAL ENCOUNTER SAME TAKE. Rear camera swings forward again. A few metres ahead near the outer edge of the café terrace stands @image7. Preserve the same face, shaved head, bright blue button-up shirt, backpack, posture and enormous friendly smile from the reference. He is standing naturally on the Paris sidewalk. As @image1 approaches, the man turns his attention toward him. @image1 casually greets him: “Hi.” Brief natural pause. The man maintains the same huge smile and replies simply: “Okay.” Nothing else. @image1 keeps the camera pointed at him. Approximately one second of awkward silence. The man keeps smiling. @image1: “…okay.” He walks past him, then quickly flips to the FRONT CAMERA. @image1's confused face fills the frame while he continues walking. He quietly says: “Was ist denn heute hier los?” The handheld movement continues naturally for the final moment. END. IMPORTANT CONTINUITY: One uninterrupted physical walk. No cuts. No teleporting. No sudden location changes. No duplicate reference characters. No character morphing. No changing outfits. No backgrounds from the reference images. No meme captions or text. Each encounter happens naturally farther along the SAME sidewalk. The comedy must come from contrast: @image1 behaves like a completely normal vlogger walking through Paris, while these bizarrely familiar expressions and situations simply happen around him and everyone else acts as if they are normal. The final result should feel like genuine accidental smartphone footage, NOT five staged scenes edited together.
======================================== RUNWAY - SEEDANCE 2.5 - PARIS-MEME-VLOG 8 Referenzbilder, 1 Prompt, 23 Sekunden ======================================== DIE SLOTS --------- @image1 -> Charakterbogen von dir (Portraet + Ganzkoerper) @image2 -> Die Location, Vordergrund frei @image3 -> 0:00 Begegnung 1, Geste im Vorbeigehen @image4 -> 0:05 Begegnung 2, Zwei-Personen-Szene @image5 -> 0:10 Begegnung 3, am Cafetisch @image6 -> 0:15 Begegnung 4, tiefer Kamerawinkel @image7 -> 0:20 Begegnung 5, kurzer Dialog In Runway an genau diesen Stellen "@" tippen und das Bild aus der Liste waehlen. Aus dem Platzhalter wird ein Chip mit Vorschaubild. REIHENFOLGE ----------- Die Nummer kommt aus der Upload-Reihenfolge im Reference-Tab, nicht aus dem Dateinamen. Erst hochladen, dann den REFERENCE-ASSIGNMENT-Block oben im Video-Prompt darauf anpassen. WAS NICHT IM PROMPT STEHT ------------------------- Jedes zusaetzlich hochgeladene Bild wird trotzdem gelesen. Bei mir lag ein achtes Bild ohne Slot in der Session - es ist in Begegnung 2 mitgelaufen.
SETTINGS
Die Einstellungen im Überblick.
| Modell | Seedance 2.5 |
|---|---|
| Referenzen | 8 Bilder, 7 davon im Prompt adressiert |
| Format | 9:16 |
| Auflösung | 1080p |
| Dauer | 25s angefragt, 23s geliefert |
| Ton | aus dem Modell: Dialog, Straße, Café |
| Kamera | Front und Rück im selben Take |
WORAUF ES ANKOMMT
Darauf musst du achten.
Referenzen sind Rollen, keine Kulissen
Der entscheidende Absatz steht ganz oben: Die Vorlagen sind ausdrücklich nur FACE, EXPRESSION, POSE AND GESTURE references, ihre Hintergründe, ihr Licht und ihr Bildausschnitt werden ignoriert und im Pariser Tageslicht neu aufgebaut. Ohne diesen Satz bringt der Lacher sein Fernsehstudio mit und der Hund sein Wohnzimmer.
Was du hochlädst, spielt mit
In meiner Session lag ein achtes Referenzbild, das im Prompt keinen Slot bekommen hat. Es ist trotzdem aufgetaucht: In Begegnung 2 steht im Prompt ein Paar, im fertigen Video läuft neben dem Mann das Kind aus dem unbenutzten Bild. Das Modell liest den Reference-Tab, nicht nur die @-Chips — wer sauber choreografieren will, lädt nichts hoch, was er nicht adressiert.
Ein Take muss dreimal dranstehen
Der Prompt sagt im ersten Satz ONE SINGLE UNINTERRUPTED HANDHELD TAKE, wiederholt an jedem Übergang SAME TAKE, NO CUT und schließt mit einem eigenen Absatz über die Kontinuität. Fünf Begegnungen sind für ein Videomodell die Einladung, fünf Clips zu schneiden — die Wiederholung ist das, was dagegenhält.
Zeitfenster statt Szenenliste
Jede Begegnung hat einen Stempel — 0:00–0:05, 0:05–0:10 und so weiter — plus eine Angabe, wo sie auf dem Gehweg passiert und dass die Figur danach nie wiederkommt. Damit wird aus fünf Ideen ein Ablauf. Ohne die Stempel schiebt das Modell die Höhepunkte zusammen und lässt am Ende zwanzig Sekunden Laufen übrig.
Fünf Gesichter, die jeder sofort erkennt.
Ein Charakterbogen von dir, eine Location, fünf Vorlagen — den Rest macht der Prompt oben.
FAQ


