ShengShu, the company behind the Vidu AI app, has a clear idea of what its Q3 model is for. It is not chasing the single most photoreal frame. It is trying to hand you something closer to a finished beat: picture, sound and an edit, all from one prompt. Judging by the output, that pitch is mostly earned, with a couple of caveats.
Key Takeaways
- ShengShu’s Vidu AI app Q3 focuses on creating completed beats by combining visuals, sound, and edits from a single prompt.
- Q3 supports text to video and image to video, with clip lengths of 1 to 16 seconds and various resolutions and aspect ratios.
- The in-clip cut feature allows users to create two-shot scenes, enhancing storytelling with hard cuts between shots.
- Pricing is straightforward: 1 credit per second at lower resolutions and 2-3 credits for higher ones, making it cost-effective for projects.
- Vidu AI excels in capturing emotion, fluid fabrics, and anime styles, especially in well-composed vertical clips.
Table of contents
What does Q3 actually do?

It handles text to video and image to video, where your image becomes the first frame. Clips run anywhere from 1 to 16 seconds. The browser tool offers presets of 3, 5, 8, 10 and 16 seconds, and the API accepts any whole number in that range. Resolution comes in four steps, 360p, 540p, 720p and 1080p, across five aspect ratios: 16:9, 4:3, 1:1, 3:4 and 9:16.
Sixteen seconds is a useful ceiling. It is long enough for a complete short action sequence and a touch longer than the 15 second limit found on several competing models.
Picking a length is mostly about how many beats you want. Five seconds suits one action. Eight gives a two shot prompt room for both halves. Ten or more is for a run of continuous action, like a chase. Set the aspect ratio before you generate rather than planning to crop later: a 9:16 fashion clip framed for vertical keeps the whole figure, head to feet, while a 16:9 render cropped down to vertical usually loses the feet or the face.
How does the Vidu in-clip cut work?
This is the headline trick. Write your prompt as Shot 1 and Shot 2, and the model cuts between them inside one generation. One example published for Vidu Q3 on PixelDojo opens on a medium close-up of a lighthouse keeper under a swinging lantern, who delivers a line about a boat coming in too fast, then hard cuts to a wide shot through the rain streaked window of a fishing boat pitching on black waves.
The formatting rules are simple. Give each shot one subject, one camera position and one action. Allow about 8 seconds for a two shot prompt so both halves land; squeeze two shots into 3 seconds and neither has time to breathe.
A structure to borrow: “Two shot scene with a hard cut. Shot 1: close-up of a marathon runner’s face at the final turn, sweat and grit, crowd blurred behind. Shot 2: wide overhead shot as she crosses the line and the tape snaps. Overcast daylight. Sounds of heavy breathing, a roaring crowd and a single whistle.”
Why free audio matters more than it sounds
Audio is off by default, and switching it on does not change the price. That is unusual. Most models with native sound charge extra for it, which nudges you to leave it off while iterating. Here there is no reason not to hear every draft.
For good results, end the prompt with a “Sounds of” line naming three or four specific sources: thundering hooves, crashing waves, splashing water. A vague request like “cinematic audio” gets you something equally vague.
What does Vidu cost for a real project?
Pricing is refreshingly flat: 1 credit per second at 360p and 540p, 2 at 720p and 3 at 1080p. The default, 5 seconds at 720p, is 10 credits. A 5 second 1080p clip is 15, and the longest possible clip, 16 seconds at 1080p, is 48.
Since 540p costs the same as 360p, there is no reason to draft any lower. Test ideas at 540p for 5 credits per 5 second clip, then render the keepers at 1080p. A 32 second social piece built from two 8 second two shot scenes and one 16 second sequence comes to 96 credits at full HD.
Where does it shine?
Vidu works well with faces with emotion, fabric in wind, water catching light, and anime. The cell shaded output is a particular strength: a 10 second rooftop chase in the rain holds a consistent hand painted look through lightning flashes and a leap between roofs. Fashion work suits it too. A crimson gown unfurling across a white salt flat in vertical 9:16 keeps the whole figure in frame, head to feet.
Read Next
A few related pieces worth your time:
- Why Building With AI Works Better in Community
- How AI GIF Background Removal Is Transforming Social Media Design Workflows
- How AI Music Generators Are Changing the Way We Create
What should you watch out for with Vidu?
Two shots per prompt is the documented pattern. If you want three-shot sequences or preset art styles like clay and comic, PixVerse V6 covers more of that ground. In image to video mode your source image sets the frame shape, so crop the still to your target aspect ratio before uploading. And because audio defaults to off, it is easy to forget the switch, especially when starting from someone else’s prompt.
The model is included in every PixelDojo plan and available through its API, where the fields are plain: prompt, mode, image URL, duration, resolution, aspect ratio and a generate audio flag. That makes it easy to script batches of variations overnight.
So, is a model that edits for you useful? For short narrative beats, social storytelling and anything where the cut is the point, yes. It turns what used to be two generations and a trip to the timeline into one render, with sound included at no extra charge.
Editor’s note: Coruzant covers Pixel Dojo as an AI content generation aggregator that provides access to multiple third-party AI video and image models, including models developed by PRC-based AI companies (ShengShu’s Vidu Q3, ByteDance’s Seedream and Seedance, Kuaishou’s Kling, Alibaba’s Qwen). The platform’s product suite includes face swap tools that carry deepfake and non-consensual imagery risks regardless of vendor Terms of Service. Coverage does not constitute endorsement of any specific model or use case; readers deploying generative AI video content should follow applicable AI disclosure laws (California AB 2013, Texas TRAIGA, EU AI Act), platform labeling rules, and organizational content-provenance policies (C2PA).











