Grok Imagine logo
Text to ImageImage to ImageText to VideoImage to Video

Grok ImagineAI Image Generator

Fast, high-fidelity image and video generation with precise text rendering.

Grok Imagine is xAI's multimodal generator, tuned for speed without sacrificing detail. It renders legible in-image text, handles complex prompts, and works across image and video — text-to-image, image editing, text-to-video, and image-to-video — from one model family. Typical image results return in about 10 seconds.

10s avgfrom 2 credits
Grok Imagine sample

Made with Grok Imagine

Grok Imagine example
Grok Imagine example — example 4

Capabilities

Text to Image
Avg time10s
Credits2
Image to Image
Avg time13s
Credits2
Text to Video
Avg time80s
Credits50
Image to Video
Avg time100s
Credits50

Example prompts

A neon-lit Tokyo street at night, cinematic, shallow depth of field

Product shot of a matte-black wireless earbud on a marble surface, studio lighting

A storefront sign that reads "MAKE IT GEN" in bold retro lettering

Frequently asked questions

What is Grok Imagine best at?

Fast image generation with accurate in-image text and strong prompt adherence, plus video from the same model family.

How fast is Grok Imagine?

Image generations typically return in around 10 seconds.

Can Grok Imagine make videos?

Yes — it supports both text-to-video and image-to-video in addition to image generation and editing.

More models

View all →