AI Media Studio: One Prompt Box for Images and Video

I wanted a place where I could type a sentence and get a picture back. Then I wanted the picture to move.

That is the entire origin story of AI Media Studio. It is a web app that takes a plain text prompt and hands back an image or a short video clip, using Google’s Gemini image models and the Veo video model. No node graphs. No twelve-slider control panel. A prompt box, a few sensible knobs, and a gallery.

This post is a field note on what it does and how it is built.

What it does

The core loop is boring on purpose: write a prompt, pick a model, pick a shape, hit generate.

Images run on three Gemini image models, which I’ve nicknamed after the “Nano Banana” family:

  • Nano Banana (gemini-3.1-flash-lite-image): fast and lightweight. Good for sketching ideas when you don’t yet know what you want.
  • Nano Banana 2 (gemini-3.1-flash-image): the balanced default. More detail, still quick.
  • Nano Banana Pro (gemini-3-pro-image): the slow, high-fidelity one for intricate concepts that the lighter models smear.

Video runs on Veo (veo-3.1-lite-generate-preview). Video generation is not instant, so the app polls for status and shows you where the job is instead of leaving you staring at a spinner that may or may not be alive.

Both paths share the same controls:

  • Aspect ratio: square 1:1, widescreen 16:9, portrait 9:16 for phones, plus 4:3 and 3:4.
  • Toggles for prompt enhancement, safety filtering, and resolution.

Nothing exotic. The point is that the same handful of choices covers almost everything I actually want to make.

The part that made it feel like a product

A blank prompt box is intimidating in a way that people underestimate. So the app ships with a curated prompt library, sorted by style: photorealism, digital art, cyberpunk, architecture, cinematic. There is also a shuffle button that drops a random prompt into the box.

Neither of those features is technically interesting. Both of them are the difference between “I opened it and closed it” and “I opened it and made six things.”

The other piece is the gallery. Generated images and videos land in a tabbed grid. Click anything and a lightbox opens with the full-resolution result, zoom, and the metadata that matters later: the prompt, the model, the aspect ratio, and when it was made. From there you can download the PNG, JPEG, or MP4 in one click, or copy the prompt to riff on it.

That metadata turned out to be the feature I use most. Half of generative media work is remembering what you typed the time it worked.

How it’s built

Nothing here is clever, which is how I like it.

  • Next.js 15 with the App Router, on React 19.
  • Tailwind CSS for styling, with Radix UI headless components underneath the menus, dialogs, and toggles.
  • @google/genai for the model calls. These live in server-side API routes: one for image generation, one for polling video status, and one that streams downloads back to the browser so the model API key never touches the client.
  • Framer Motion for route transitions and the little bits of movement that make a media app feel less like a form.
  • Lucide for icons.

Dark and light themes are both first-class. The header sticks to the top so the model and ratio controls stay in reach while you scroll a long gallery.

What I learned

Model choice should be a dropdown, not a decision. Early on I tried to explain the three image models in the UI. Nobody reads that. Now the names carry the hint (lite, default, pro) and people figure it out by generating the same prompt on two of them.

Polling is a UX feature. For video, the difference between “generating…” and “generating, step 3 of 5, about 40 seconds left” is the difference between trust and a closed tab.

Keep the key on the server. Obvious, but worth saying. Every call to Gemini or Veo goes through an API route. The browser never sees credentials, and I can rate-limit and log in one place.

Metadata is the product. The lightbox that shows prompt, model, ratio, and timestamp cost an afternoon and gets used more than anything else.

Try it

It’s live at image-video-factory.ai.studio. Type something absurd, pick 16:9, and see what Nano Banana 2 thinks a “cyberpunk noodle stall at dawn, wet asphalt, neon reflections” looks like. Then ask Veo to make it move.

If you make something good, copy the prompt. Future you will want it.