MiniMax Preview Generation Model: A single model for both live and post-production preview based on the H3 architecture
According to Dynasty Beating monitoring, the MiniMax H3 team revealed in a Reddit AMA that they are developing a dedicated image model and plan to open-source the weights. This model will combine textual image generation and general image editing into a single framework and is currently in the post-training optimization phase.
This model shares the same architectural philosophy as H3. It will utilize H3's VAE Encoder and introduce a separately designed VAE Decoder for image generation. The envisioned workflow by MiniMax is to first use the image model to generate the initial frame, which will then be passed to H3 to continue generating the video.
The team also mentioned that H3 itself has shown potential for both image generation and editing. Previously, they only trained the model to predict the final frame based on the "initial frame + text description," without specific image editing training. However, the model still demonstrated strong zero-shot capabilities in various image editing evaluations. This serves as a key rationale for MiniMax to extend this architecture to an image model.