Generate and edit images and short video
Turn a prompt into a visual or a short clip without leaving your workspace — DALL·E, GPT Image, FLUX, Stable Diffusion, Imagen, Kling, Wan, Hunyuan and more, metered against your plan.

15image models
14video models
4images per request
60 svideo via chained 5 s clips
8image providers
What it does
01
Image generation & editing
Create visuals for decks, proposals and posts, then edit them in place with models that support it. Up to four variations per request.
02
Short-form video
Text-to-video, image-to-video and video-to-video editing. Base clips are 5 seconds; chain up to 8 clips into one 60-second video.
03
Inside agents and workflows
Image and video agents generate media on execution; the media-generation workflow step does the same on a schedule. Everything lands in your Files vault.
How it works
01Describe what you need — an image or a clip — and pick a model.
02Upload a reference image for editing or image-to-video, or start from text.
03Generate, review and regenerate; long videos are stitched from 5-second clips.
04Drop the result into a chat, a workflow or your Files.
Models out of the box
Image
- DALL·E 3, GPT Image 1, Grok 2 Image
- FLUX & FLUX.1 schnell, Stable Image Core, SD 3.5 Large, SDXL
- Imagen 3, Gemini 2.5 Flash Image, CogView 3/4
Video
- Kling 1.6 & 2.1, Wan 2.1 & 2.2, Hunyuan Video, LTX Video
- MiniMax Hailuo & Director, Veo 3.0 Fast, Seedance 1.0 Lite
- Text-to-video, image-to-video and video-to-video
Voice AI (separate module)
- Spoken assistant and translator in 24 languages
- Speech-to-text on our own servers
- Text-to-speech in 6 voices, metered by the second
Built for
Real-world scenariosClient-facing proposal visuals
Product shots and social drafts
Quick explainer clips
Scheduled media in a workflow