Verification: 234cbc2215f1fb96
Pricing

Start frame

Optional

End frame

Optional

Multi-scene mode
720p
8 seconds
9:16

AI video generator

Generate cinematic videos in just minutes

Select the motion effect
1

Select the motion effect

Decide how your image will move

Add image
2

Add image

Upload or generate an image to begin your animation

3

Get video

Click generate to produce your final animated video!

Super Promotion

90% OFF

Create stunning AI photos & videos with essential tools

Unlock the Basic Plan for just $1

Auto-renewal is active. Cancel anytime. 90% off applies to the first billing cycle.

By choosing your age and continuing you agree to our Terms of Use and Privacy Policy
Please review before continuing

Veo 3.1 AI Video Generator

The old version of this page tried to win with scale: benchmark drama, speculative pricing, broad claims about every Veo feature Google has ever shown, and a lot of language that did not map cleanly to the route people actually open on Cleep. That is the wrong shape for a useful route page. A strong page for /generate/video/veo31-video should start from two grounded layers: what Google officially documents for Veo 3.1 today, and what the current Cleep route actually exposes.

On Google's side, the current Vertex AI Veo 3.1 documentation is specific. The core 3.1 Generate and 3.1 Fast Generate capability tables describe support for text-to-video, image-to-video, prompt rewriting, and videos from the first and last frames, with 16:9 and 9:16, 720 and 1080, 24 FPS, English prompt language, and a maximum of 4 videos returned per request. Google also documents Veo's prompt strategy very clearly: the official Veo prompt guide keeps returning to framing, camera motion, style, lighting, character detail, location, action, dialogue, and sound.

On the Cleep side, this route is narrower and more practical than the full Veo family story. The current veo31-video configuration exposes text-to-video, image-to-video, first-last-to-video, and a reference-to-video path, plus multishot, up to three image inputs, a visible clip length of 8 seconds, visible resolutions of 720p and 1080p, both 16:9 and 9:16, and an estimated turnaround of roughly 3 to 5 minutes. That makes the honest reading much simpler: Veo 3.1 on Cleep is best treated as a premium short-form planning route for tightly directed 8-second clips, not as a catch-all summary of every Veo workflow Google documents across every surface.

Quick Answer

Start with Veo 3.1 when the real task is not "make any AI video," but "plan a polished 8-second shot with strong direction." On Cleep, this route is especially useful when you want both vertical and horizontal output, need to move between text, image, reference-led, or first-last-frame workflows, or want a more premium planning surface than a simpler short-form route gives you.

The most honest mental model is this: Cleep's Veo 3.1 page behaves more like a fast-generation Veo lane with extra planning inputs than a full product matrix. It is very good at deliberate short clips. It is not the right page to promise automatic long-form editing, universal multilingual prompting, or every preview feature in Google's wider Veo stack.

Where Veo 3.1 is actually strong on Cleep

The strongest thing about this route is not one isolated spec. It is the combination of premium short-form quality and a more planning-heavy input surface. A lot of nearby routes give you one entrance: maybe text, maybe one image, maybe a first-last transition. Veo 3.1 on Cleep gives you several ways in without turning the page into a full editor.

The fixed 8-second shape matters more than it first seems. It forces better shot design. Instead of trying to cram a whole narrative into one generation, you can plan a clean reveal, a product moment, a fashion beat, a stylized motion plate, or a short branded transition with a clear beginning, middle, and finish. That is often exactly what premium short-form social or campaign work needs.

The route also sits in a practical production lane: 720p or 1080p, 16:9 or 9:16, estimated completion in 3 to 5 minutes, and enough image inputs to move beyond "text only." That makes it more useful for teams who already know the look they want and need the route to hold onto a brief, a still, a beginning and ending frame, or a simple multi-shot plan.

Veo 3.1 capability board showing 8-second premium clips, 720p and 1080p output, 16 by 9 and 9 by 16 aspect ratios, plus text, image, reference, first-last, and multishot workflows
Veo 3.1 is strongest on Cleep when you read it as a premium short-form route with several planning inputs: text, stills, reference-led direction, first-last transitions, and multishot structure, all aimed at deliberate 8-second clips.

The 8-second limit is a feature, not just a constraint

It pushes the route toward clean shot design: one reveal, one move, one transition, one emotional beat, one ad moment. That discipline often produces better premium short-form work than a vague "make me a longer video" brief.

Veo 3.1 covers both social and landscape output

Because the visible route supports both 9:16 and 16:9, it can serve vertical campaign cutdowns and standard landscape shots without forcing you to move to a completely different route family.

This route is planning-heavy in a useful way

Text, still image, first-last frame, reference-style input, and multishot all matter for different briefs. That gives Veo 3.1 more production shape than a simpler "text in, clip out" page.

Prompt behavior is part of the product, not an afterthought

Google's Veo documentation treats prompt rewriting and prompt detail as core behavior. On this route, writing better prompts is not optional polish. It is part of getting the route to act like a premium tool.

What Google officially documents, and what that means on this route

A strong Veo page should separate family-level facts from route-level facts. Google documents the Veo 3.1 family through Vertex AI and DeepMind. Cleep then exposes a narrower, very practical route on top of that. That distinction matters because it keeps the page honest when Google's broader docs mention capabilities that may be preview-only or not surfaced as a visible flow on this exact route.

Area Official source What it means here
Core 3.1 capabilities The Vertex AI Veo 3.1 page lists text-to-video, image-to-video, prompt rewriting, and first/last-frame generation for the main 3.1 Generate and Fast Generate models. The Cleep route lines up well with that core shape: text, still-image motion, and first-last planning are all central to how this page should be used.
Reference image to video Google's broader Veo 3.1 preview documentation describes reference image to video under preview-style capability tables rather than the plain GA list. Cleep does expose a reference-style path, but the safest reading is to treat it as a narrower workflow that should be tested carefully for your exact brief, especially when the reference carries most of the creative load.
Video extension Google documents extend-video under preview capability tables, not as part of the basic Generate / Fast Generate list. This Cleep route should not be described as an extend-video surface. There is no visible extend-video flow here today.
Prompt language Google's Veo 3.1 capability table lists English as the prompt language. Even if the article is localized, the route should still be treated as an English-prompt-first workflow for strongest results.
Clip length Google lists 4, 6, or 8 seconds for Veo 3.1, with reference-image-to-video documented more narrowly. Cleep currently exposes 8 seconds on this route, which reinforces the idea that this page is about high-control short-form clips, not broad duration flexibility.
Resolution Google documents 720 and 1080 for the main 3.1 Generate / Fast Generate path. The current route matches that visible production lane with 720p and 1080p, so there is no reason for the page to promise wider resolution coverage.
Aspect ratios Google documents 16:9 and 9:16 for the main route families. Cleep exposes both, which makes this route practical for both standard landscape shots and vertical social work.
Frame rate The official Veo 3.1 docs list 24 FPS. The route should be discussed as a cinematic short-form generation path, not as a broad frame-rate playground.
Response count Google's capability table lists a maximum of 4 videos returned per request. That fits the idea of directed iteration: generate a small batch, judge which shot is closest, then refine.
Prompt rewriting The official prompt rewriter documentation says Veo 3 and 3.1 do not let you turn prompt rewriting off, and the rewritten prompt is returned only when the original prompt is fewer than 30 words long. Prompt behavior is part of the route. You should not treat Veo 3.1 like a raw, unchanged prompt parser. It actively helps shape the input.
Prompt ingredients The official Veo prompt guide keeps stressing framing, camera motion, style, lighting, character detail, location, action, dialogue, and sound. That guidance is directly useful on Cleep. Better prompts are one of the biggest quality levers this route gives you.

How to choose the right Veo workflow on Cleep

Veo 3.1 is not one single interaction style. The most useful question is not "is Veo good?" but "how much of the shot do I already know?" If you know the mood but not the frame, use text-to-video. If you already own the first frame or key look, use image-to-video. If you know the opening and ending image, first-last can be much more practical than trying to describe the whole arc in one long prompt. And if your brief already has multiple beats, the route's multishot surface can help you think in production logic instead of one vague monolithic request.

Because Google documents Veo 3.1 as an English-prompt model family, the example prompts below intentionally stay in English. That is not accidental. It reflects the model's documented prompt language rather than some generic SEO choice.

Veo 3.1 mode board comparing text-to-video, image-to-video, first-last transitions, reference-led setup, and multishot planning
Pick the Veo 3.1 mode based on what is already fixed. Use text when only the idea is stable, stills when the look is known, first-last when the shot arc is known, and multishot when one beat is not enough.

Text-to-video

Best when the scene concept is clear, but the visual starting frame is still open.

Low-angle fashion promo shot, model steps out of a yellow taxi into wet neon light, camera tracks backward, subtle lens bloom, premium street-luxury tone, distant traffic hiss and muffled city ambience.

Image-to-video

Best when the first frame already solves styling, identity, or product composition.

Animate the uploaded skincare product still with a slow push-in, soft glass reflections, gentle vapor drift, elegant golden side light, luxury beauty ad pacing.

First-last-to-video

Best when the opening and ending states are both known and the job is really the bridge between them.

Start exactly from the first uploaded frame and end on the last uploaded frame, keep the shoe centered, create a fluid studio transition with precise lighting continuity and a clean premium campaign feel.

Reference-led planning

Best when you need the route to protect a look, mood board, or visual identity more tightly than text alone can do.

Use the reference images to preserve the heroine's styling and palette, then generate a cinematic 8-second sci-fi corridor shot with a careful handheld camera drift and restrained tension.

Multishot setup

Best when the brief already has beats instead of one isolated visual moment.

Shot 1: close product detail in darkness. Shot 2: reveal on glossy plinth with fast side light. Shot 3: hero pack with brand color reflections and a clean final hold.
Input path Best when What to lock clearly
Text-to-video You are still exploring the shot and mainly know the tone, setting, and motion energy. Framing, one key action, camera move, lighting, and visual style.
Image-to-video The first frame already matters and the route should animate rather than re-invent the scene. What should stay fixed, what may move, and how subtle or expressive the motion should feel.
First-last-to-video You know the opening and ending visual states and want the route to design the bridge. Continuity, pacing, visual identity, and what cannot drift between start and finish.
Reference-led path A look, brand world, or character identity matters enough that text alone is not enough. The protected visual cues: face, styling, silhouette, palette, art direction, or product setup.
Multishot The brief has multiple beats and the clip should feel planned rather than improvised. The order of beats, what changes between them, and the emotional or visual payoff at the end.

Prompting advice Google actually backs

One of the easiest ways to waste Veo 3.1 is to use premium video software with low-information prompts. Google's own prompt materials are unusually concrete here. The Veo prompt guide does not tell users to "be creative" in the abstract. It pushes them to specify framing, motion, style, lighting, character detail, location, action, dialogue, and sound. The separate Vertex AI prompt-rewriter docs make the same underlying point from another angle: Veo 3 and 3.1 actively rewrite prompts, and you cannot turn that behavior off.

The practical consequence is simple. If the route is important enough to use Veo 3.1, the prompt should read like shot direction, not like a lazy keyword bag. The better you map camera, motion, subject, environment, and tone, the more the route behaves like a premium planning tool instead of a generic clip generator.

  • Start with the shot, not just the subject: say whether the camera is close, wide, static, pushing in, tracking, or handheld.
  • Name the visual style early: if the result should feel like film noir, stop-motion, glossy commercial work, anime, or editorial luxury, say that directly.
  • Write lighting and atmosphere explicitly: Google's own guidance keeps returning to light, texture, and mood because they strongly shape output.
  • Use richer character detail when humans matter: specific identity cues beat generic labels.
  • Describe action in sequence when the shot is dense: the more complicated the motion, the more you should direct the order of events.
  • Use dialogue or sound cues on purpose: the official Veo materials talk about dialogue and sound, but this Cleep route should still be tested on your exact workflow rather than oversold as a separate visible audio-control surface.
  • Keep the prompt in English: that is what Google's Veo 3.1 documentation currently lists as the prompt language.
  • Remember that prompt rewriting is always on for Veo 3 and 3.1: if you want to inspect the rewritten prompt via API response, the original prompt must be fewer than 30 words.

When Veo 3.1 is the right route, and when nearby routes are easier

Veo 3.1 should not be sold as the default answer for every short video request. It is strongest when the brief benefits from its premium short-form planning surface. If the task is more about flexible short duration choices, simpler Hailuo-style iteration, or edit-first camera control, another route may fit better.

Stay on Veo 3.1

when you want a carefully directed 8-second premium clip, need both vertical and landscape output, and care about moving between text, image, first-last, reference-led, or multishot planning inside one route.

Compare with Hailuo Standard

when you want a broader short-form Hailuo workflow with visible 6- and 10-second options and a simpler first-last planning story.

Compare with Hailuo Pro

when the job is a narrower premium short clip and you do not need Veo's wider planning surface.

Compare with Seedance Pro

when you want a different creative personality for text-led generation and are testing which model family gives the more useful aesthetic for the campaign.

Compare with Kling Motion Control

when explicit motion or edit control matters more than Veo's premium directed-shot feel.

Hand the winning clip into your edit stack

when generation is no longer the hard part and the next task is subtitles, music, pacing, VO, cutdowns, or timeline assembly.

Veo 3.1 workflow board showing shot planning, English prompt design, mode selection, 8-second generation, take review, and post handoff
The cleanest Veo 3.1 workflow is usually: decide what is already fixed, choose the right input path, write the prompt like shot direction, generate a small batch, keep the strongest take, then move audio and edit polish downstream.
  • Think in one high-value shot: this route is strongest when one scene, one reveal, one transition, or one premium ad beat carries the whole clip.
  • Do not promise yourself long-form inside an 8-second route: if the brief is really multi-scene storytelling, generate beats and assemble them later.
  • Use 9:16 only when the shot truly wants vertical framing: do not just crop the idea mentally and hope the route fixes composition for you.
  • Reference-led work deserves extra testing: Google's own Veo docs split plain GA generation from preview-style reference-image flows, so validate the exact path you need.
  • Keep route-level honesty: there is no visible extend-video flow here today, and the page should not pretend otherwise.

What we verified for this page

This rewrite is grounded in primary Google material and the live route configuration inside Cleep. The fact layer comes from the official Vertex AI Veo 3.1 capability tables, the official prompt-rewriter documentation, the official Veo prompt guide, and the current veo31-video route configuration that exposes 8-second duration, 720p and 1080p, 16:9 and 9:16, text/image/first-last/reference paths, multishot, and up to three image inputs. Unsupported benchmark numbers, speculative pricing claims, and broad future-facing promises were removed on purpose.

Frequently asked questions about Veo 3.1 on Cleep

What is Veo 3.1 on Cleep?

On Cleep, Veo 3.1 is best treated as a premium short-form video route for carefully planned 8-second clips. It combines Veo-family generation quality with several practical entry points such as text, still image, first-last planning, reference-led input, and multishot.

Which workflows does the current route expose?

The current route exposes text-to-video, image-to-video, first-last-to-video, and a reference-style path, plus multishot and up to three image inputs.

What clip length does Cleep currently expose for Veo 3.1?

The current visible route is locked to 8 seconds, even though Google's broader Veo 3.1 documentation discusses 4-, 6-, and 8-second support at the family level.

What resolutions and aspect ratios are visible on this route?

The route currently exposes 720p and 1080p, plus 16:9 and 9:16, which makes it useful for both landscape and vertical campaign work.

Should I prompt Veo 3.1 in English?

Yes. Google's Veo 3.1 capability tables currently list English as the prompt language, so the safest workflow is to write the generation prompt in English even when the page itself is localized.

Can I turn prompt rewriting off for Veo 3.1?

No. Google's official Vertex AI docs say prompt rewriting cannot be disabled for Veo 3 and 3.1 models.

Does this route support reference-style generation?

Yes, the current Cleep route exposes a reference-style path. At the same time, Google's own Veo 3.1 documentation treats reference-image-to-video more narrowly than the basic GA feature list, so it is smart to test this path carefully for mission-critical briefs.

What should my Veo 3.1 prompts focus on?

Follow Google's own prompt guidance: specify framing, camera motion, style, lighting, character details, location, action, and any important dialogue or sound cues. Treat the prompt like shot direction, not loose keywords.

Does this page expose every Veo feature Google documents elsewhere?

No. The broader Veo family story is bigger than this route. For example, this page should not be described as a visible extend-video surface, and some preview-style capabilities belong to narrower workflows than the main Generate / Fast Generate tables.

When should I compare another route instead of staying on Veo 3.1?

Compare another route when you need longer short-form duration options, simpler iteration, more explicit motion-control or edit-first behavior, or a workflow that is not centered on a premium 8-second directed shot.