Blog·Guides··8 min read

Prompting video models in 2026: what actually works

The MediaMint Team

Video is the most expensive thing you can generate — a premium 8-second clip costs as much as hundreds of draft images — so bad prompting hurts more here than anywhere else. This is the internal playbook we use, written for the models on MediaMint today.

Start with the shot, not the story

Video models generate one shot well. Prompts that describe a scene the way a director would brief a camera operator — subject, action, camera, light — reliably beat prompts that describe a narrative arc. Compare:

  • Weak: “A story about a lighthouse keeper who discovers something in the storm”
  • Strong: “Handheld medium shot: a rain-soaked lighthouse keeper leans into the wind on a catwalk, storm waves exploding against the rocks below, one flickering beam sweeping past the lens”

If you need a sequence, generate shots separately and cut them together — that’s what the canvas is for. Boards keep every take next to its prompt, so assembling a sequence is drag and drop, not archaeology.

Choose the model for the job, not the leaderboard

Each of our video models has a real specialty:

  • Veo 3.1 — the premium pick: highest fidelity, native synchronized audio, up to 4K. Use it for finals, not exploration; at 200 credits per second it deserves a locked-down prompt.
  • Kling 2.5 Turbo — the workhorse. Smooth, cinematic motion at 35 credits per second, ideal for finding the shot before you commit.
  • Kling O3 — when you need control: it accepts a start frame and an end frame, so you can pin exactly where the shot begins and lands. Storyboard transitions live here.
  • Seedance 2.0 — dynamic action and believable physics, with native audio. Sports, crashes, crowds, anything kinetic.
  • Hailuo 02 Pro — strong 1080p motion realism at a mid price, particularly good with human movement.

Or skip the decision: leave the model on Auto and the assistant routes each prompt using exactly these heuristics.

Iterate cheap, finish expensive

The single biggest credit-saver: never explore on a premium model. Draft the shot on Kling 2.5 Turbo at 5 seconds. When composition, motion, and timing are right, re-run the final prompt on Veo 3.1. A typical exploration of five drafts plus one premium final costs about a third of exploring directly on Veo — for an identical end product.

Audio-native changes the prompt

Veo 3.1 and Seedance generate sound with the picture, and they listen when you describe it: “the only sounds are boots on gravel and distant thunder” gives you a sound design, not just a picture. If your model doesn’t do native audio, generate the clip silent and add foley on the canvas afterwards — ControlFoley and ThinkSound both take a video input and return a synced track for a few credits.

Image-to-video is the secret weapon

The strongest workflow on the canvas isn’t text-to-video at all. Generate a still you love — Flux 2 Pro for photoreal frames, at a fraction of video cost — then feed it to Kling or Veo as the start frame. You get precise art direction on the cheap medium and spend video credits only on motion. This is also how you keep a character or product consistent across shots: same source still, different motion prompts.

The checklist

  • One shot per prompt; sequences are assembled on the board.
  • Name the camera: “static wide”, “slow dolly-in”, “handheld tracking”.
  • Name the light: “golden hour backlight”, “harsh fluorescent office”.
  • Pick duration deliberately — 5s drafts, longer only when the shot earns it.
  • 16:9 for screens, 9:16 for social. Deciding later means paying twice.
  • Draft cheap, upgrade the survivor.

Every example in this post runs on the free tier’s daily credits — open a board and try the image-to-video loop; it’s the fastest way to feel the difference.

Try it on the canvas

100 free credits a day. Every model. No card.

Start creating