There’s a new AI video gen tool hitting the blogosphere, being tested by users for the first time, and challenging competing models for popularity among creatives.

Chinese company Minimax has revealed their “Minimax H3” model, which can generate video with sound for up to 15 seconds, with approximately 2K resolution.

Here are some key components of the design:

  • Contextual Omni Representation enables H3 to integrate text, images, audio, video, and other types of input into a unified contextual understanding model, for reasoning.
  • H3-VA is MiniMax’s vision-language model that can understand and reason across images and text, for tasks like visual question answering, document analysis, and image-grounded conversations.
  • The H3-Omni Transformer is MiniMax’s multimodal transformer architecture that processes and integrates text, images, and audio. This provides what experts call “contextual understanding,” with reasoning and generation across different modalities or input types.
  • In-Context Regeneration is a technique that allows the model to revisit and refine parts of its output using the surrounding context, improving coherence, factual accuracy, consistency, and overall response quality without retraining the model.

So all of that offers a robust ability to utilize context in putting together short videos.

Here are more things to know about H3’s development and release.

Open Source, Open Weights

In a way, H3 from Minimax continues a Chinese trend of using open source design, offering the public community the code behind the service. Here’s part of the announcement directly from the company’s July 31 press release:

“Closed-source models have long dominated video generation, with slower iteration and a less open ecosystem than fields like large language models. To support the open-source community, accelerate compatibility with a broader range of AI hardware, and make it easier for users to build their own customized versions, we plan to open up the model weights in the coming days, subject to applicable laws and regulations. Hardware compatibility has been a key consideration since the earliest stages of H3’s design.”

That shows some of the rationale behind the decision to open-source this promising new video tool.

Outsourcing Context

Although there seems to be no formal acknowledgement of this, reports from the open-source community indicate that H3 may use Qwen3-VL-32B as a text encoder/component, rather than being a model derived from Qwen itself.

“It’s just a DiT like most other image/video models,” writes Redditor Nextil of the new H3 tool. “The multimodal ‘understanding’ comes from the text encoder, which is just Qwen3-VL-32B, a 10 month old model. There have been many open models released capable of computer use.”

Territorial Restrictions

The third caveat here is that Minimax H3 is actually not available in some pretty big markets, namely:

  • The United Kingdom
  • The European Union
  • The United States
  • South Korea

Lest you think that this is just more of the sort of protectionism you see practiced by China and the U.S. in their knock-down, drag-out war for AI dominance, the decision actually has to do more with regulations.

With a little digging, I found this explanation of Minimax’s release strategy on Hugging Face:

“MiniMax-H3 was built with the goal of global availability. The current territory scope is not about excluding specific countries or regions, but about recognizing that video generation models are facing a more complex and rapidly evolving regulatory environment compared with text or code models.”

Pointing out that regions such as the EU, UK, South Korea, and the US are currently developing or enforcing AI-related regulations that may have specific implications for generative video models, the writer enumerates a few of them, which I am including verbatim:

  • The EU AI Act has started enforcement, while practical requirements for models capable of generating video and likeness-related content are still evolving.
  • Similar regulatory uncertainties exist in the UK and South Korea regarding AI-generated content and video generation.
  • In the US, AI regulation remains a rapidly changing landscape, and MiniMax is also involved in ongoing copyright-related legal proceedings specifically concerning generative video AI.

The author continues to elaborate:

“For open-weight models, once the weights are released, developers can deploy and modify them independently. This creates different compliance challenges compared with hosted services.”

That’s a little bit about what’s behind the regional moratorium. Experts point out that Minimax can still control various aspects of H3 use, for safety, since every request goes through the company’s own servers.

Iterative Video

I also wanted to include the Minimax press release’s explanation of how H3 works by making two passes:

“For H3’s 2K output, instead of using a conventional dedicated super-resolution module, we have the H3 base model regenerate its own low-resolution output in-context,” spokespersons write. “This brings two advantages: first, the regeneration process maximally reuses the generative capability already built into the H3 base model; second, the in-context approach lets it draw on the original multimodal context again to produce high-resolution output, recovering details that traditional super-resolution can only ‘guess’ at and often can’t restore, like small text and fine detail.”

Cool.

So now you know more about a generative AI platform to rival Google’s Veo or Synthesia and others. Stay tuned for more.

Share.
Exit mobile version