Blog · · The Rubiic team

Code, not pixels: why Rubiic writes your video in Remotion

Generated video is pixels with nothing underneath. Rubiic writes each scene as Remotion code, inside a harness built for it. Why, and where pixels still win.

Two ways to make a video with AI

Ask an AI for a video today and, underneath, one of two things happens.

The first is generation. A video model turns your prompt into frames, and what comes back is the clip itself: pixels, sampled one frame after another.

The second is authorship. A language model writes a program, and the program draws the video. You get the program as well as the pixels, and a program can be read, changed and run again.

Rubiic is the second kind. Every scene in a Rubiic video is a Remotion composition — React code — written for your script and rendered in the cloud. This post is about why we built it that way, and about the less obvious half of the answer: why “an AI that writes Remotion” is not enough on its own.

Generated pixels

Video models are extraordinary at what they were built for. They produce motion, texture and light that would take a crew and a day to film. For atmosphere — a street at dusk, moving water, a face turning toward the camera — nothing else comes close.

But look at what they hand back. A generated clip is the finished surface with no editable document underneath. There is no layer that says “this is the title”, no chart that knows its numbers, no timeline that marks where the second sentence begins. It is all pixels, and pixels have to be asked for again.

That makes a particular kind of video hard to pin down:

  • On-screen text. Letters are part of the image, so exact wording and spelling are something you check for rather than something you set.
  • Precise diagrams. An arrow that leaves the receptor rather than entering it, a bar at 42 rather than 40, a flowchart whose boxes say what the narration says — these are facts, and a sampled frame has nowhere to keep a fact.
  • Your brand. A colour you specify as a hex value comes back as an impression of that colour.
  • One-line edits. “Change the second title” means generating again, and the new clip can differ in everything else too.

None of that is a failing of video models. It follows from what they produce. But an explainer, a lesson or a product video is made mostly of the things on that list.

Video as code

Remotion is an open-source framework for making video with React. A scene is a component, and each frame is whatever that component renders for a given frame number. Here is a title that fades in over its first second:

import {AbsoluteFill, interpolate, useCurrentFrame} from "remotion";

export const Title = ({text}: {text: string}) => {
  const frame = useCurrentFrame();
  // Fade in over the first second (30 frames).
  const opacity = interpolate(frame, [0, 30], [0, 1], {extrapolateRight: "clamp"});
  return (
    <AbsoluteFill style={{justifyContent: "center", alignItems: "center"}}>
      <div style={{opacity, fontSize: 96, fontWeight: 600}}>{text}</div>
    </AbsoluteFill>
  );
};

The consequences are the reverse of the list above. Text is set by a browser in a real font, so the words on screen are the words in the code: crisp, and fixed by editing a line. A chart is drawn from its numbers. A colour is a value in the code, not an impression of one. And a change is an edit to a file followed by a render: “make the second title shorter” touches one line, and the rest of the file is the same file it was.

Code also composes. A scene can hold an image, an animated diagram and a caption at once, each placed to the pixel and timed to the frame, because they are all just parts of one component.

And it makes the work legible. When the agent changes a scene, it changes code that someone could read, in a new version you can put beside the last one — so “what did it just do?” always has an answer.

A language model with a code editor is not a video studio

So why not ask a general-purpose assistant to write the Remotion? You can, and it will get surprisingly far. It will also run into problems that have nothing to do with how well it writes React. Most of Rubiic is the answer to those problems.

It can reach for anything

An assistant will happily import a charting library that is not installed. In Rubiic, scene code may import React and Remotion — and images you have accepted — and nothing else, and that is checked when the scene is compiled. A scene that reaches for something missing fails there, with a message the agent can act on, instead of halfway through a render.

It cannot see what it made

Code that compiles can still put white text on a white background. Rubiic previews each scene in a sandboxed page, and the agent can take screenshots of its own frames in a headless browser, read any runtime errors, and fix what it finds. For a complex composition it can also ask a separate visual reviewer to inspect the frames. Both are checks the agent chooses to run, not a certificate: still frames do not show motion, and nothing in it knows whether your diagram is scientifically right.

Versions drift

Remotion's bundler and its renderer have to agree exactly. When they do not, a render does not fail — it comes out black. Rubiic pins Remotion to one exact version and refuses to render when the cloud function or the build environment was made with another. A refusal you can read is better than a black video you have paid for.

Narration has to fit

A voiceover is a length of audio; a scene is a number of frames. Rubiic aligns the recorded narration to the approved script from measured timestamps, so the agent knows when each word lands and can time the visuals to it. If the narration would run past the end of its scene, the scene is refused with the length it would need — the last sentence is never quietly cut off.

What you hear should be what you get

Music sits under the voice, lower while someone is speaking. Rubiic mixes the two with a single mixer that the preview in your browser and the render in the cloud share, so the balance you approved is the balance you download.

A minute of video is a lot of frames

Rubiic renders on AWS Lambda, splitting a longer scene across up to fifty workers. A render job is identified by the scene, its version and the render settings, so asking for the same render twice picks up the same job rather than paying for a second one.

The file has to be right, too

After a render, Rubiic can decode the whole MP4 and check its duration, dimensions, frame count, audio signal and narration timing against what was planned. It is an automated check, not a promise about taste. Each example on our site lists the checks its file actually passed.

What that means when you use it

  • Every version is kept. Each revision is a new version, and you can compare it with the one before, so an edit that makes things worse costs nothing to walk back.
  • Your brand is reusable. Rubiic proposes a palette, fonts and logos from pages it has read; you review them, and every scene takes them as an input. Save the look as a style and the next project starts from it. It is guidance, not enforcement — a font it does not have is swapped for one it does — so look at the result.
  • You stay in the loop. In Review mode, the default, you approve the narration script before it is recorded, and the music and the export as they come.
  • One project, many formats. Any size from 320×240 to 3840×2160, landscape or vertical, plus still images, GIFs of up to ten seconds, carousels of up to twelve pages and SRT or VTT captions, all from the same scenes.
  • The MP4 is yours. You own what you make, and you download a file that plays anywhere.

Where pixels still win

This is not an argument that generated video is a mistake. If you need a photoreal person, a real-looking place or cinematic camera movement, a video model is the right tool, and Rubiic does not generate video clips. Where a picture helps inside a scene, Rubiic can generate an image grounded in a web search and place it in the composition. The image is pixels; the scene around it — the words, the timing, the diagram it sits beside — is still code.

That is the split we would suggest to anyone: generate what has to look real, and write what has to be right.

See it for yourself

The quickest way to judge is to watch one. The examples page shows finished videos alongside the brief that produced each. The first, a 60-second explainer about using Rubiic from Claude Code, was made by us with Rubiic, and the checks its file passed are listed under it. The use-case pages go through what to bring for research explainers, documents, lessons and product videos. And if you already work in Claude, Rubiic is available there too, as a connector.

See what it makes for you

Rubiic is invite-only while it scales. Join the waitlist and we will tell you when a seat opens.