Shengshu Technology's Vidu Q4 surges into the top three in image-to-video: previous generation ranked only 19th
Beating AI News Flash: Shengshu Technology has released its new-generation flagship video model Vidu Q4 Preview, now available via web and API. The new model currently supports image-to-video and reference-to-video generation, but does not yet support text-only video generation.
On Artificial Analysis's image-to-video (with audio) blind test leaderboard, the new model ranks 3rd with a score of 1,179, trailing only MiniMax H3 Max and MiniMax H3. The previous-generation Vidu Q3 Pro scored 1,056, ranking 19th. The two generations differ by 123 points.
Vidu Q4 Preview mainly improves facial expressions, body movements, camera transitions, and visual effects such as explosions. It can generate visuals and audio simultaneously, with a single video lasting up to 16 seconds and supporting up to 4K resolution. Users can also upload up to 15 reference images and 3 audio references to specify character appearance, scenes, and voices, keeping characters as consistent as possible across different shots.