---
format: "aidr-story-markdown/v1"
id: "02a178ff700f1bbeddd541b658c59d9a3a93650fb7012061246822292a8163da"
canonical_url: "https://aidr.today/02a178ff?lang=en"
title: "PixelUMM Removes VAEs and ViTs From Multimodal Video Models"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-02T07:38:13.000Z"
category: "Open Source"
topics: ["open-source"]
source_urls: ["https://huggingnews.com/ai/pixelumm-removes-vaes-and-vits-from-multimodal-video-models-6ce06acd","https://x.com/CongWei1230/status/2105808995910816159","https://marketbrief.now/ai/pixelumm-removes-vaes-and-vits-from-multimodal-video-models-6ce06acd"]
summary: "The code and trained weights for a new image and video system are now public for developers to use in multimodal tasks. Developed by CongWei1230, the model known as PixelUMM performs generation and understanding directly in pixel space, which eliminates the need for Vision Transformers (ViTs) and Variational Autoencoders (VAEs) typically used to encode visual data in AI models. The project was released on October 1 and utilizes an encoder-free architecture to bypass standard compression steps found in most generative AI. By removing the reliance on latent spaces, the model seeks to unify how images and videos are processed and produced within a single framework."
---

# PixelUMM Removes VAEs and ViTs From Multimodal Video Models

> [Open the canonical story](<https://aidr.today/02a178ff?lang=en>)

**Published:** 2026-10-02T07:38:13.000Z
**Category:** Open Source
**Topics:** open\-source

## Summary

The code and trained weights for a new image and video system are now public for developers to use in multimodal tasks\. Developed by CongWei1230, the model known as PixelUMM performs generation and understanding directly in pixel space, which eliminates the need for Vision Transformers \(ViTs\) and Variational Autoencoders \(VAEs\) typically used to encode visual data in AI models\. The project was released on October 1 and utilizes an encoder\-free architecture to bypass standard compression steps found in most generative AI\. By removing the reliance on latent spaces, the model seeks to unify how images and videos are processed and produced within a single framework\.

## Sources

- [Story source](<https://huggingnews.com/ai/pixelumm-removes-vaes-and-vits-from-multimodal-video-models-6ce06acd>)
- [Story source](<https://x.com/CongWei1230/status/2105808995910816159>)
- [Story source](<https://marketbrief.now/ai/pixelumm-removes-vaes-and-vits-from-multimodal-video-models-6ce06acd>)

