Training Text-to-Image Models Without a VAE

3 pointsposted 10 hours ago
by schopra909

6 Comments

schopra909

10 hours ago

Hi HN, author here!

For context, we're a 2-person lab training generative video models. Goal is a new set of controllable, animation tools (you can read more about that here https://www.linum.ai/about if you're curious).

The biggest bottleneck for our last text-to-video model in terms of training and inference cost is attention. Video models are incredibly token dense (e.g. 110K tokens for a several second clip). If we can condense that context window more aggressively, we can train bigger models for a lot less $$ and offer them to prosumers at reasonable price points (unlike the big models today like Seedance, which cost an arm and a leg to run).

Traditionally, image and video models have two disjoint components: VAE (Variational Autoencoder) and Diffusion Transformer (DiT). They're trained separately, and empirically VAEs seems to struggle to get past 16x16 token reduction.

Here, we're switching to pixel-space, throwing away the VAE, and achieving 32x32 token reduction (4x smaller context windows) while learning a better overall model in a fraction of the training samples.

The central thesis is "simpler is better". If we can put the compression problem into the more powerful Diffusion Transformer (DiT) would should be able to learn a "latent space" optimized for generation and get better compression without hurting generation quality.

I'll be checking this post off and on the next couple of hours, so feel free to drop questions below. And I'll try to answer them to the best of my ability.

P.S. The model checkpoints from this blog are Apache 2.0, so feel to try playing with it yourself on a GPU!

E-Reverance

8 hours ago

I know it goes a tiny a bit against the spirit of what y'all are doing, but applying a few layers of pixel-wise local attention (so 1x1 "patch", with 3x3 or 5x5 attention window, basically treating it as a dynamic conv) has worked way better than both linear and conv unpatching in my recent experiments.

Diagram for reference https://x.com/1rreverant/status/2107546198093730287 (In my most recent recent experiment I actually removed the MLP and just used a linear project on the pixel's hidden states)

schopra909

4 hours ago

That sounds like an interesting idea!

Can you confirm I'm understanding correctly?

1) Linear unpatchify as usual to go from hidden states to pixel space

2) Attention within a local window (e.g. 3x3, 5x5) to "blend" pixel space data and come up with a better image (as an alternative to MLP or Convolution)

And follow up questions:

1) How do you handle boundaries between your "attention windows"? Do you move the window just like a convolution does or are the "attention windows" all mutually exclusive from one another?

2) How much faster/slower is this operation vs. a linear layer + MLP?

E-Reverance

4 hours ago

1) yes, but with more than channels than 3 (and no correspondence to color, its "pixel space" in the sense of position, not value)

1D grid example for clarity (obviously meant to be done in 2D though)

So embed -> [N dim, N dim, N dim ... , N dim] instead of embed -> [RGB, RGB, RGB ... , RGB]

2) / [followup 1)] Stride of 1, so we place a window at each pixel (so lots of overlapping)

Also I wouldn't phrase it as "come up with a better image", the point is to give less spatial decoding pressure to the patch tokens so that they can almost completely focus on feature learning instead. There is no reason to have the model learn spatial decoding when the structure prior of images is comically strong (especially compared to text), its a waste of training time and parameters

[followup 2)] I haven't measured but it was passable is all I can say (my experiments setup are horrendous right now lol)

(Also don't mind the phrasing, I just wanted to be 100% clear)

schopra909

4 hours ago

I appreciate the clarifications!

If we have extra compute lying around in the coming weeks, we’ll try this out and report back. It’s a good idea :)

E-Reverance

4 hours ago

I loved to hear! Just ping me on twitter (same account I sent the link with)