grok-imagine-video-1.5
Non-official Grok 1.5 image-to-video generation requiring exactly one source image, with 1–15 second duration and up to 1080P output.
Prompt to video, in context
15-second cinematic historical period-drama scene, realistic, intimate, elegant, restrained, filled with quiet longing. A woman remains SEATED at a writing desk beside a rainy leaded window, writing a personal letter to someone she once loved. Aristocratic interior, aged wood, period dress, feather pen, old paper. Warm flickering candlelight contrasts with cool blue-gray rainy daylight. Emotional rhythm: WRITE → THINK → WINDOW GLANCE → WRITE → PAUSE → SECOND WINDOW GLANCE → FINISH → LEAN BACK → LOOK OUT WINDOW. 0–3s: Close-up of feather pen slowly writing beside a candle. The pen briefly hesitates above the paper. Camera gently reveals her thoughtful face. 3–6s: Writing slows. She pauses, subtly tightens her lips, then makes a brief FIRST glance toward the rainy window. She returns to the page and continues writing. 6–9s: The pen pauses again. Her free hand lightly touches the letter. She makes a SECOND, longer window glance. Her eyes soften and lose focus slightly, as if remembering someone far away. Candlelight warms one side of her face while cool window light shapes the other. 9–12s: She returns to the page and slowly writes a final line. Her expression softens from emotional weight into tenderness and longing. 12–15s: She stops writing, lowers the pen, and gently LEANS BACK INTO THE CHAIR. She remains seated. Her shoulders relax as she looks toward the rainy window with a soft, distant gaze and the faintest private smile. End there. Performance: Natural micro-expressions: small writing hesitations, hovering pen, trembling lashes, subtle breathing, slight mouth tension, softened eyes, brief loss of focus. No crying or melodrama. Camera: One continuous smooth cinematic shot. Start close on pen, paper and candle, drift naturally to hands and face, capture both window glances, then end on her seated profile looking toward the window. No abrupt cuts or flashy movement. Audio: No dialogue or voiceover. Soft pen on paper, rain, candle ambience, fabric movement, light wind against glass, subtle chair creak. Minimal melancholic period piano or strings, warming slightly near the end. STRICT: She stays seated for the entire scene. Never stands or walks to the window. Two window glances while writing. Writing is repeatedly interrupted by thought. End leaning back in the chair looking toward the window. No subtitles, captions or on-screen text.
15-second realistic cinematic American Western scene at golden sunset. One continuous unbroken shot, no cuts, no montage, no overlays, no time jumps. Two lovers ride slowly on horseback through open Western country on a dusty trail. Dry grass, canyon rocks, warm sunset light, drifting dust, and a quiet romantic atmosphere. The scene feels like a high-quality modern Western film: intimate, natural, and emotionally alive. The woman rides slightly ahead at first, relaxed and confident. One hand lightly controls the reins, her hair moves in the warm wind, and she enjoys the peaceful ride. The man follows behind, then gradually catches up and rides beside her naturally. Their horses move realistically at a slow walking pace with gentle body sway, proper rein control, and smooth spacing. Emotional progression: peaceful ride → quiet affection → playful conversation → shared comfort → subtle concern → alert attention. 0–3s: Camera begins low beside the woman's horse, tracking the slow walk. Show hoof movement, dust, reins, and her relaxed posture. She looks over the landscape with a calm expression. 3–6s: The man slowly catches up beside her. He smiles warmly. She notices him and gives a small restrained smile. Their connection feels familiar and effortless. 6–9s: They talk naturally while riding side by side. Man: "You always ride ahead when something’s on your mind." Woman: "And you always catch up when I want a quiet minute." She adjusts the reins and briefly touches her horse’s neck. Their horses continue walking calmly. 9–12s: They share a quiet affectionate moment. Then the man notices something far away to the left. His smile fades slightly. His eyes focus on the distance and he gestures subtly. Man: "Hey... look over there." 12–15s: The woman follows his gaze. Her expression changes slowly from warmth to focus. She gathers the reins slightly. Both horses remain calm but become more attentive, ears forward, heads slightly raised. Wind moves her hair, dust glows in the sunset, and the camera holds both riders and the distant horizon in the same frame. End with the feeling that something may happen beyond the horizon. Performance: Natural micro-expressions only: small smiles, subtle eye contact, brief glances away, changing breathing, slight tension in the face, and realistic emotional transitions. No exaggerated acting. Camera: Smooth cinematic tracking shot. Start close on horse, reins, dust and rider. Slowly widen to include both lovers. Stay intimate and fluid. Follow their shared attention shift toward the distant horizon. No camera cuts. Lighting: Golden sunset backlight, warm highlights, long shadows, realistic dust haze. Audio: Natural hoofbeats, saddle leather, reins, wind, dry grass, open-country ambience. Minimal Western guitar or low strings. Music becomes slightly tense near the end. Rules: * Both riders remain on horseback for the entire scene. * Horses move slowly and realistically. * Keep romance subtle and believable. * No forced danger, only a hint of something approaching. * English dialogue only. * No subtitles, captions, or text.
15-second realistic cinematic Western adventure, one continuous shot, no cuts. A female explorer stands at a dusty desert excavation site, holding a smartphone at arm’s length on a live video call. Behind her is a gigantic half-buried ancient spherical artifact. Harsh golden sunlight, dry rocks, excavation equipment, drifting sand, realistic skin, sweat and dust. Serious high-budget adventure-film style, not flashy sci-fi. Emotion: excitement → curiosity → unease → fear. 0–4s: She excitedly frames herself and the enormous artifact in the video call, looking between the phone and discovery. Light wind moves her hair and hat. "Tell me you can see this. This could be the find of my life." 4–8s: She turns slightly to show the artifact and gestures toward it with her free hand. Her smile fades as she studies its strange surface. Wind grows stronger and hair blows across her face. "No... wait. This isn’t just ancient." 8–12s: She brushes hair away while keeping the phone raised. Dust increases. Her eyes narrow as excitement becomes concern. "The surface looks engineered... almost biological." 12–15s: A strong gust blows her hat completely off and out of frame. She startles and tries to catch it with her free hand but misses. The phone wobbles. Hair whips across her face as she looks back at the artifact with alarm. "We may be standing way too close." Camera: smooth continuous cinematic tracking, keeping her face, phone and artifact visible. Audio: natural dialogue, desert wind gradually strengthening, sand, fabric, distant excavation sounds, restrained suspense score. STRICT: one hand holds the phone throughout; realistic live-call behavior; wind builds gradually; hat stays on until the final gust, then flies completely away; artifact remains stationary; natural acting; no cuts, subtitles, captions or text. 16:9
8K photorealistic vertical 6-second funny urban chase motion clip, ultra-low fisheye wide-angle shot close to the pavement. The overall movement and camera speed are increased by 25%. Continuous single long take with no cuts. Vintage 35mm motion picture film color grading, subtle film grain, soft vignetting, warm cinematic lighting tones, delicate contrast, movie-level depth of field atmosphere, theatrical American street comedy visual texture. Bright daytime city commercial street flanked by towering skyscrapers, blue sky dotted with white clouds, fluttering white pigeons in the air, passersby stopping to watch along the street. A huge fluffy long-haired gray tabby cat darts forward rapidly with a fresh fish clamped in its mouth. A curly-haired lady in a black leather jacket chases closely behind. Intense longitudinal motion blur stretches across the pavement throughout the sequence, crisp hard shadows cast by sunlight. Ultra-detailed rendering of fur, fish scales and leather jacket textures. The cat sprints forward, paws pushing hard against the ground, firmly holding the fish in its jaws with a mischievous, triumphant glint in its eyes. Its long fur and whiskers stream backward in the wind. The woman leans forward chasing hurriedly, looking anxious, and speaks softly: Stop! She stretches her right hand forward trying to catch the cat, her curly hair and leather jacket fluttering wildly with the running motion. The camera glides swiftly forward close to the ground with subtle bumpy vibration; motion blur gradually builds up on the pavement.[00:00-00:02] The cat flattens its ears backward, lowers its head to speed up and dodge. The woman closes the gap slightly, panting faintly, and says helplessly in a quiet tone: Drop the fish! Her fingertips are nearly touching the cat’s fluffy coat, her expression a mix of anxiousness and absurd amusement. The camera keeps accelerating forward, with stronger bouncing jitters and deeper motion blur.[00:02-00:04] The cat nimbly sidesteps to evade the grab, lifts its head with a pleased, victorious look and dashes far away still clutching the fish. The woman staggers slightly after missing her lunge, her outstretched hand falls empty. She sighs in resignation and mutters: Seriously… wearing a weary yet amused look before breaking into a run again. The camera swerves slightly following the cat and rushes ahead at full speed, with maximum high-speed motion blur. A slight pullback shot wraps up the scene.[00:04-00:06] Pigeons flap their wings and scatter continuously, pedestrians turn their heads to watch, buildings on both sides streak rapidly backward. The body movements of the cat and woman are smooth and natural without stiff AI distortion. Matching audio: Snappy, lighthearted playful percussion soundtrack accelerated accordingly; layered with rapid footsteps, wind noise and pigeon wing flutters. Voice lines are calm and restrained, never loud or harsh.
15s, 9:16, realistic cinematic Western, one continuous shot. She walks slowly toward the camera as it smoothly tracks backward. Her coat and hair move naturally in the warm desert wind, boots kicking up fine golden dust. At 4s, she hears a distant sound. Her eyes shift first, then her confident expression slowly becomes serious. She keeps walking for two more steps, then stops. At 8s, a stronger gust pushes dust across the street. Without looking away, her hand slowly moves toward the revolver at her hip, fingers resting on the grip but never drawing it. The camera gradually pushes closer. Her eyes narrow slightly as she stares at someone beyond the camera. At 12s, she quietly says: "You should've kept riding." She remains perfectly controlled as wind moves her coat and dust drifts through the golden sunset. Hold on her intense stare until the end. AUDIO: realistic boots on dirt, leather creaks, dry wind, distant wooden signs, faint horse sounds, restrained Western tension score. Natural body movement, subtle micro-expressions, realistic wind and dust physics, restrained acting. No cuts, no subtitles, no text.
15s, 9:16, photorealistic psychological horror, grounded high-budget film realism. One continuous shot. The woman walks cautiously down the abandoned hospital corridor, gripping the flashlight with both hands. Camera smoothly tracks backward. Her footsteps are slow, breathing controlled, flashlight beam moving naturally across peeling walls and dark doorways. She nervously whispers to herself: "Okay... just keep moving." Around 4s, a metallic object suddenly falls somewhere deep in the corridor. She freezes instantly. Her eyes move first. Her grip tightens and her breathing becomes shallow. She raises the flashlight toward the darkness. "Hello?" Silence. A distant fluorescent light behind her suddenly flickers on with an electrical buzz. Then another light closer to her flickers. She slowly looks over her shoulder. Her expression shifts naturally from confusion to genuine fear. "Is someone there?" Another light flickers on closer. She takes a small step backward, keeping the flashlight raised. A faint scraping sound suddenly comes from the dark doorway beside her. Her eyes snap toward it. After a tense pause, she whispers: "Don't come any closer." Hold on her frightened, controlled expression. Nothing is revealed. AUDIO: realistic footsteps, nervous breathing, fabric movement, distant metallic impact, fluorescent buzzing, electrical flicker, building creaks, faint scraping from the doorway. Her voice is quiet, breathy and naturally frightened. Minimal low atmospheric tension, almost no music. CAMERA: realistic handheld-stabilized tracking, subtle natural movement, no cuts or sudden zooms. REALISM: physically accurate flashlight movement and shadows, natural walking, breathing and micro-expressions. Fear builds gradually, never theatrical. No monster visible, no supernatural effects, no jump scare, no screaming, no subtitles, no text.
What you can build with grok-imagine-video-1.5
Prompt driven generation
Submit a concise prompt and get production-ready video output through one TTAPI job.
Core controls only
Keep integration simple with prompt, model, output settings, and callback handling.
Async result flow
Use polling or webhook callbacks so long-running generations do not block your UI.
Ready for product UI
Return grok-imagine-video-1.5 results that can be displayed, stored, or passed into downstream workflows.
Per-action, in quota
Usage is metered in quota by model, action, and output settings.
grok-imagine-video-1.5 operations and request fields
Use the endpoint below with your TTAPI key and the request fields shown.
/grok/generationsSubmit a non-official Grok video task with video_length, resolution_name, reference images, and an optional callback.
Official reference/grok/fetchRetrieve a non-official Grok video result.
Official reference/grok/extensionsExtend a non-official Grok video.
Official referenceHeaders
TT-API-KEYstringrequiredYour TTAPI API key.
Content-TypestringrequiredUse application/json.
Body
promptstringrequiredGeneration prompt. Base and 1.5 accept up to 4,096 characters; 1.5 Fast supports longer prompts.
modelenumoptionalDefaults to grok-imagine-video-1.5-fast. Also supports grok-imagine-video and grok-imagine-video-1.5.
aspect_ratioenumoptionalDefaults to 16:9. Supported: 2:3, 3:2, 1:1, 9:16, and 16:9.
video_lengthstringoptionalDefaults to 10. Fast supports 6–30 seconds; base and 1.5 support 1–15 seconds. Base supports up to 15 seconds with one image and up to 10 with multiple images.
resolution_nameenumoptionalDefaults to 720p. 480p and 720p work across all models; 1080p is supported only by grok-imagine-video-1.5.
refer_imagesstring[]conditionalReference-image URLs. Fast accepts up to 7; base accepts multiple; 1.5 requires exactly one image.
voice_idenumoptionalVoice role ID for the generated video audio. Options include carina, zagan, helix, orion, luna, iris, altair, and zenith.
hook_urlstringoptionalCallback URL notified when the job completes or fails; otherwise retrieve the result through Fetch.
From key to first result
Copy the resolved endpoint, authentication headers, and request body for this model.
curl --request POST \
--url https://api.ttapi.io/grok/generations \
--header 'TT-API-KEY: $TTAPI_KEY' \
--header 'Content-Type: application/json' \
--data '{"model":"grok-imagine-video-1.5","prompt":"a young man closes his eyes on a crowded nightclub dance floor","aspect_ratio":"2:3","video_length":"6","resolution_name":"1080p","refer_images":["https://example.com/first-frame.jpg"],"voice_id":"luna"}'