FLUX 3 x mimic: The Next Generation of Video-Action Models

FLUX 3 x mimic: The Next Generation of Video-Action Models

We combined our FLUX 3 multimodal foundation model with mimic robotics' expertise to create FLUX-mimic, a new video-action model for real-world automation. By treating actions, video, and audio as views of a single physical reality, we proved that one backbone can master both content creation and robot control. This approach allows robots to learn complex manipulation tasks with minimal data, scaling our world model from the lab to actual production lines at Audi.

If one model does both, it was never really only a content creation model. It is a model of how the world behaves, and content creation is one thing one can do with it.
  1. GiffertonThe3rd

    The video at around 3.30min, where the robot arm took 3 attempts to reseat the window trim, was quite unnerving - I have not seen such resolving before. Is it new or am I way out of the loop?

  2. vessenes

    Really interesting. Upshot: a well trained multimodal video generation model has a world representation model trained inside it. They’ve done some work lifting this world model out and deploying it to robots, where it seems to work well.

    On the one hand, this isn’t a new idea, and the quality video models certainly have understanding of materials, light, the world (at least in an Occam’s razor sense of understanding). I’m not aware of a video lab that’s turned itself into a robot lab yet, though; perhaps this would be a first, or a new sort of obvious-in-retrospect business path: train video model, sell video generation, scale, use scale to train robot things: profit.

    I found their hands very interesting - looks like a bunch of stuff hidden in gloves - Xiami’s Robotics-1 foundation model just released videos of users training on some pretty standard looking grippers; to the point that there are demo videos of people putting on gripper type gloves to make video to train that model.

    The BFL model looks like it doesn’t need that at all. Given the difficulty of the hardware side, I’ll be curious to see what they do with this.

  3. flufluflufluffy

    Ok the phrasing here, it’s, it’s just -

    > However, compared to more specialized approaches for representation learning they produce less disentangled representations, which puts a ceiling on their usefulness for tasks that require world understanding.

    Only an LLM would use a less disentangled representation of the concept “more entangled” when trying to explain to people in the real world why less disentangled representations are not as useful for modeling the real world.

  4. altern8

    This is awesome.

    What's sad to me is that we have all this awesome technology, but movies are worse than ever.

    I usually watch movies from decades ago just to find something decent and it's amazing how with goofy-looking puppets the storytelling was 1000 times better.

    This is probably an unrelated rant, sorry

  5. rlupi

    It's nice to see partnerships between European startups.

More from this day

2026-07-24