ComfyUI v0.36.0 roughly halves MiniMax H3 VAE time in Comfy's tests

Comfy reports that ComfyUI v0.36.0 cuts a 1344x768, 129-frame MiniMax H3 VAE round trip from 24.3 to 12.7 seconds on an RTX 5090, but the figures are…

ComfyUI v0.36.0 roughly halves MiniMax H3 VAE time in Comfy's tests

Comfy says it has roughly halved the time MiniMax H3 video workflows spend in the VAE. According to a post on the Comfy blog published September 22, 2026, ComfyUI v0.36.0 encodes MiniMax H3 video up to about 2.2x faster and decodes it about 1.4–2.7x faster. In Comfy's own measurements, a 1344x768, 129-frame encode-and-decode round trip fell from 24.3 to 12.7 seconds. These are self-reported, single-GPU figures for the VAE stage only. For teams doing AI video production in ComfyUI, the practical question is how much of their pipeline that stage actually represents.

Three changes, three different requirements

Comfy describes three separate optimisations. Each one has a different adoption cost.

ChangeWhat it does (per Comfy)Claimed gainWhat you need
Fused encoder kernelMerges per-frame normalisation, activation and edge padding into a single pass that writes directly into the layout the next convolution expects~1.5x encoder, identical outputv0.36.0 or later; on by default
Custom convolution honouring fp16 accumulationComfy says PyTorch sends convolutions to cuDNN, which offers no way to request fp16 accumulation; the new custom convolution honours the flag and folds the bias and skip connection into the convolutionEncoder to ~2.2x in totalLaunch with --fast fp16_accumulation
Faster int8 decoderUses 8-bit weights with normalisation, activation and skip connection folded into the matrix multiplications; attention also runs in 8-bit1.4x over the previous int8 VAEThe int8 MiniMax-H3 VAE file, which Comfy calls a drop-in replacement. Comfy's setup steps also tie --fast fp16_accumulation to a faster decoder; the post does not say whether the 1.4x int8 figure depends on the flag

Before this change, Comfy says, the fp16 accumulation flag applied only to matrix multiplications, so the encoder got no benefit from it. Comfy says the flag helps on consumer cards where fp16 accumulation runs at double rate. The post does not list which cards those are.

The post does not say exactly which combination of settings produced the 24.3-to-12.7-second round trip or the top 2.7x decode figure. Treat the mapping between configurations and headline numbers as unconfirmed until it is checked with Comfy.

What the quality figures measure

Comfy says the fused encoder kernel is lossless: it computes the same values in a different order. Encoding still runs in fp16 even when the int8 VAE file is loaded. For the paths that do not produce identical output, Comfy reports these figures:

  • The int8 decoder matches the standard decoder at 67.7 dB PSNR.
  • The faster encoder matches at about 68 dB.
  • The VAE's own reconstruction of real footage scores roughly 38 dB.

The two sets of numbers use different references. The ~68 dB figures compare the new paths against the standard VAE's output. The 38 dB figure compares the VAE against source footage. Comfy says what "the faster int8 paths" add is "roughly 30x smaller loss again" than the VAE's inherent loss, and that "you will most likely not see a difference by eye." The post does not make clear whether that 30x comparison also covers the faster encoder, which runs in fp16 rather than int8. The post does not publish visual comparisons or perceptual metrics. Motion Engineering Today has not tested the output, and the sources reviewed for this article report no independent testing.

What the figures do not cover

  • One card, one resolution. All measurements were taken on an RTX 5090 at 1344x768 and 129 frames. Comfy describes them as same-day A/B runs against the pre-change build, taken through the VAE Encode and VAE Decode nodes. Comfy says the fused kernel's reduced memory traffic "helps any NVIDIA GPU," but the post reports no results from other cards.
  • No end-to-end timing. The post does not report total generation time or sampling time. It therefore does not establish what share of a full MiniMax H3 run the VAE accounts for. The savings will matter most where VAE work is a large share of the total, such as encode-heavy video-to-video or repeated re-decoding. That is our analysis, not a measured result.
  • No non-NVIDIA data. The post discusses only NVIDIA hardware and cuDNN.

How to adopt it

Comfy's instructions:

  1. Update ComfyUI to v0.36.0 or later. The fused encoder kernel needs nothing else.
  2. Start ComfyUI with --fast fp16_accumulation for the faster encoder and decoder.
  3. Load the int8 MiniMax-H3 VAE file for the fastest decode.

Comfy says matching workflows are available from the blog post and in ComfyUI's template library. Our recommendation: before switching production pipelines, teams should time a full run on their own hardware with and without the flag and the int8 file.

Frequently asked questions

How much faster is the MiniMax H3 VAE in ComfyUI v0.36.0?

According to a Comfy blog post published September 22, 2026, ComfyUI v0.36.0 encodes MiniMax H3 video up to about 2.2x faster and decodes it about 1.4–2.7x faster. In Comfy's tests, a 1344x768, 129-frame encode-and-decode round trip fell from 24.3 to 12.7 seconds. These are self-reported figures from a single RTX 5090 and cover only the VAE stage. Comfy does not say which settings produced the headline numbers, and it reports no end-to-end generation timing or results from other GPUs.

How do I turn on the faster MiniMax H3 VAE in ComfyUI?

Comfy gives three steps. First, update to ComfyUI v0.36.0 or later. The fused encoder kernel is on by default and needs nothing else. Second, launch ComfyUI with --fast fp16_accumulation for the faster encoder and decoder. Third, load the int8 MiniMax-H3 VAE file for the fastest decode. Comfy says matching workflows are on its blog and in ComfyUI's template library. Before switching production pipelines, it is worth timing a full run on your own hardware with and without the flag and the int8 file.

Does the faster ComfyUI MiniMax H3 VAE reduce video quality?

Comfy says the fused encoder kernel is lossless and produces identical output. For the paths that change output, Comfy reports these PSNR figures against the standard VAE: 67.7 dB for the int8 decoder and about 68 dB for the faster encoder. For comparison, the VAE's own reconstruction of real footage scores roughly 38 dB. Comfy says you will most likely not see a difference by eye. However, the post publishes no visual comparisons or perceptual metrics, and no independent testing has been reported.

AI-assisted, reviewed by JJ Rooney on 6 October 2026.