Skip to content

FLUX 3 can generate a 20-second audio-video, Audi has used it to train factory robots

Jul 24, 19:14

According to Dynamic Beating monitoring, video generation AI company Black Forest Labs has released the multimodal model FLUX 3. The previous two generations were mainly used for image generation, while FLUX 3 is the first to incorporate images, videos, and audio into the same model. It can generate and edit images and create videos with sound up to 20 seconds long.

FLUX 3 supports text-to-video, image-to-video, video rewriting, and continuation. During scene generation, it can synchronously add dialogue, ambient sound, and action sound effects. It also supports multilingual dialogue and concatenation of multiple segments.

Black Forest Labs has also collaborated with mimic robotics to develop the robot model FLUX-mimic based on FLUX 3. It uses the movement and physical laws learned by the video model to predict the next step of a robotic arm. Audi has tested and deployed it in factories to perform tasks such as part sorting, component insertion, and handling of flexible materials like cables and seals.

Depending on the task's complexity, FLUX-mimic requires a minimum of only 30 minutes of robot demonstration data, significantly less than the previous solutions that usually required over 30 hours. The video version is currently open for early applications, the image version will be open in the coming weeks, and the robot version will be initially available to partners. The open-weight version FLUX 3 Dev is planned for release later this year.

View source