Voiceover B-roll

Aesthetic b-roll under a first-person voiceover.

Aesthetic b-roll under a first-person voiceover — no one speaks on camera, the visuals are atmosphere, the argument is entirely in the narration.

When to use it

For brand-forward pieces and any market where a face would narrow the audience. It is also the cheapest way to make a considered, well-written argument, because the writing is the format.

The thing to get right

Nothing lip-syncs here, so the voiceover is free to be the strongest thing in the ad. Write it first and shoot the visuals to it, not the other way round.

The format at a glance

  • Output — a video ad.
  • Structure — 5 beats, each one a separately planned shot.
  • Family — Organic social video.
  • Recipe id — voiceover_broll, which is what you name when you ask for it by hand or from an agent.

Which comes first, the script or the visuals?

The script. Nothing lip-syncs here, so the narration is free to be the strongest thing in the ad and the pictures exist to carry it. Written the other way round the voiceover ends up describing footage instead of making an argument.

Video formats aimed at reach rather than a click — founder takes, behind-the-scenes, slideshows and swipeable carousels. They are written to look like content, not like an ad.

  • Day in the life — a "day in my life" vlog montage of a creator’s real day (first-person)
  • Photo slideshow — a TikTok/IG photo slideshow — a sequence of vertical text-forward slides (hook → points → CTA)
  • Listicle carousel — a swipeable tips/listicle CAROUSEL — one tip per slide
  • Trend / relatable — a trend-style relatable scenario reel (riding a format/sound)
  • Carousel ad — a multi-image swipe AD carousel (Meta/IG) — sequential selling slides
  • Behind-the-scenes — authentic behind-the-scenes / day-in-the-life footage

The full catalog is on the creative recipes index.

Models that render it

This format is rendered by a video model. The plan picks one unless you name it, and Seedance 2.5 is the default pick. You can ask for another by name, such as MiniMax H3 Max Turbo, Seedance 2.0 or Gemini Omni 1.1. All three render native sound, so a voice comes out of the same take as the picture. The video models page compares every option.

Hermoso plans and renders a voiceover b-roll from your brand rather than from a prompt you write, composites every on-screen word after the render so the copy comes out exact, and exposes the same format to an AI agent over MCP. How Hermoso renders a format covers all three.

Point Hermoso at your brand and render a voiceover b-roll today — the free plan needs no card.

Start free →   See pricing