diffusion_decoder differs from the current ltx-2.5-video-vae and lacks decoder.type_emb

#19
by christopher5106 - opened

Hi! While porting the SDR-To-HDR IC-LoRA (with seam keyframes) to diffusers, we compared this repo's diffusion_decoder with the original Lightricks/LTX-2.5 vae/ltx-2.5-video-vae-bf16.safetensors. The two don't match:

  • Missing tag: the original file has the keyframe tag decoder.type_emb (BF16, [128]), and this repo's decoder does not.
  • Different weights: none of the 161 decoder tensors that can be matched by name are equal. The relative differences range from 2–6% on conv_in/conv_out to 10–40% on w_down, context_proj and upsamples.*. The latent statistics are identical.

On a decoder-only check with identical inputs and noise against Lightricks' own decoder (49 frames, 480x736), this repo's decoder matches at 44.4 dB for a plain decode and 35.2 dB with keyframes. The original VAE converted to diffusers matches at 62.2 dB and 68.6 dB.

The diffusers converter is updated to carry decoder.type_emb and set decoder_keyframe_type_embedding=True in https://github.com/huggingface/diffusers/pull/14975. Would you consider re-converting diffusion_decoder from the current original VAE with it? Happy to share the comparison script.

Hi @christopher5106
The diffusion vae decoder weights were updated.
Thank you

Thanks @art-alex ! The new diffusion_pytorch_model.safetensors is byte-identical to our conversion of ltx-2.5-video-vae-bf16.safetensors (sha256 bb3801cc…2ee1), so plain decoding now matches the reference (62 dB in our check, up from 44 dB).

One small follow-up for later: the file now carries decoder.type_emb, but config.json doesn't set "decoder_keyframe_type_embedding": true, so diffusers drops the tag as an unexpected key. That flag only matters once diffusers supports keyframe-aware decoding, which is being discussed in huggingface/diffusers#14981, so there's nothing to do until then. Closing the discussion from our side is fine.

LTX.io org

We are working on adding the keyframe-aware decoding support for Diffusers library. Once implemented I will set the decoder_keyframe_type_embedding in the config.
Once merged I will update in this thread.
Thank you

Thanks, great to hear! In case it's useful as a reference: huggingface/diffusers#14975 (closed) has a working port of the keyframe-aware decode (joint attention with the two nearest planes, type_emb, converter change), matching your decoder at 68.6 dB with keyframes on a decoder-only check. Until it lands in diffusers, the same decode ships as an interim custom class in https://huggingface.co/scenario-labs/ltx25-sdr-to-hdr-modular (SDR-To-HDR with seam keyframes, bitwise-identical to that port on the real weights). Happy to share the parity script or help test your implementation.

Sign up or log in to comment