← Back to live feed1 story
Friday, Oct 2, 2026
The code and trained weights for a new image and video system are now public for developers to use in multimodal tasks. Developed by CongWei1230, the model known as PixelUMM performs generation and understanding directly in pixel space, which eliminates the need for Vision Transformers (ViTs) and Variational Autoencoders (VAEs) typically used to encode visual data in AI models.
The project was released on October 1 and utilizes an encoder-free architecture to bypass standard compression steps found in most generative AI. By removing the reliance on latent spaces, the model seeks to unify how images and videos are processed and produced within a single framework.
Sign in to suggest edits
Key sources
- SOURCE@congwei1230“Introducing 𝗣𝗶𝘅𝗲𝗹𝗨𝗠𝗠: an encoder-free unified multimodal model for image and video understanding and generation, directly in pixel space.”x.com
- SOURCEmarketbrief.now