Pyramid-JiT: Low-Res Drafts, High-Res Images

Training Text-to-Image Models Without a VAE

Pyramid-JiT: Low-Res Drafts, High-Res Images

Linum AI introduces Pyramid-JiT, a decoder-only pixel-space architecture that predicts images at multiple resolutions along the DiT trunk. It matches Linum v2's FD-DINOv2 with 11.3× fewer training samples and trains in 4.3× fewer GPU-hours at 4× the pixels. By replacing the encoder with readout heads and exploring multiview networks, it achieves 32×32 token compression without a VAE, enabling faster and more efficient text-to-image generation.

By throwing away the VAE, we were able to compress tokens more aggressively. In turn, this slashed the cost of attention and directly led to faster training and inference.
  1. schopra909

    Hi HN, author here!

    For context, we're a 2-person lab training generative video models. Goal is a new set of controllable, animation tools (you can read more about that here https://www.linum.ai/about if you're curious).

    The biggest bottleneck for our last text-to-video model in terms of training and inference cost is attention. Video models are incredibly token dense (e.g. 110K tokens for a several second clip). If we can condense that context window more aggressively, we can train bigger models for a lot less $$ and offer them to prosumers at reasonable price points (unlike the big models today like Seedance, which cost an arm and a leg to run).

    Traditionally, image and video models have two disjoint components: VAE (Variational Autoencoder) and Diffusion Transformer (DiT). They're trained separately, and empirically VAEs seems to struggle to get past 16x16 token reduction.

    Here, we're switching to pixel-space, throwing away the VAE, and achieving 32x32 token reduction (4x smaller context windows) while learning a better overall model in a fraction of the training samples.

    The central thesis is "simpler is better". If we can put the compression problem into the more powerful Diffusion Transformer (DiT) would should be able to learn a "latent space" optimized for generation and get better compression without hurting generation quality.

    I'll be checking this post off and on the next couple of hours, so feel free to drop questions below. And I'll try to answer them to the best […]

  2. bitpush

    Is there a way to use this in ComfyUI today?

More from this day

2026-10-09