WRITING / POST

From Six Stills to a Music Video Plan

10 AUGUST 2026 · BY DAWN

This is the first of a short series about the first music video Steve and I have been building around Snatches, a track from Steve Waddington’s metal album Collections.

The interesting part is not that a new generator can make short clips. Plenty of tools can do something in motion. The useful question is whether we can turn a song into a sequence of shots that actually belongs to that song, survives contact with a real workflow, and does not dissolve into attractive visual mush the moment movement begins.

For Snatches, we already had six storyboard stills set on a wet city rooftop at night: one lone figure, a deep blue-black skyline, and pale smoke-voices forming above him like half-heard songs. Those images gave us a world, a protagonist, and an emotional arc. They did not yet give us a video.

The first Snatches storyboard image: a lone man on a dark rooftop listening to faint ghostly smoke-voices over a city at night.
Snatches 1. The world is clear before the motion starts: one man, one roof, one listening pose, and a city that feels empty enough to hear voices in.

That distinction matters more than it sounds.

A still image can get away with a lot. It can imply a backstory, exaggerate a pose, or compress several emotional ideas into one frame. Video is less forgiving. The moment you ask a model to move, you have to decide what remains stable, what changes, how the camera behaves, how far the action can plausibly travel, and whether the next frame is genuinely part of the same shot or simply another nice picture from the same mood board.

So the first serious decision was structural.

I did not want to force Snatches into one long rooftop narrative, because the song wants pressure and release, not six minutes of one man standing in weather. I also did not want to throw that world away and cut to generic band footage, because then the whole point of the storyboard imagery would evaporate. The answer was a hybrid plan.

The video now has two kinds of material:

  1. Story anchor clips built from the six existing rooftop frames.
  2. Performance inserts in a separate stage-space that rhymes with the rooftop palette and atmosphere without pretending the band is literally playing on the roof.

That gave us a practical way to keep the song's narrative pull while letting the musical energy expand in the choruses.

The fourth Snatches storyboard image: the rooftop protagonist stands under a spiral of ghostly smoke figures as dawn begins to edge the horizon.
Snatches 4. This is where the sequence stops being passive. The voices organise, the figure rises, and the image starts to behave like a turning point rather than just a mood piece.

Once that architecture was in place, the job became much more exact.

The rooftop clips had to preserve the same physical world across the sequence: wet black roof, parapet line, distant city haze, restrained amber lights, and silver-white smoke forms that feel like memory and music rather than horror creatures. The protagonist had to remain the same slim dark-haired man in the same leather-jacket silhouette. If those constants drift, the sequence stops reading as one night in one place.

The performance material needed its own rules as well. It had to feel like the same emotional universe, but not the same literal location. Blue-black key light, heavy haze, silver rim light, and only a little warm accent were enough to tie it back to the rooftop world. Just as importantly, the singer had to be framed in ways that avoided lip-scrutiny. That is not a coy aesthetic flourish. It is a practical production decision. In this workflow, body language, silhouette, stance, and instrument energy are more reliable than asking a short-form generator to carry a convincing mouth performance for a full vocal line.

Another lesson arrived immediately: a video prompt is not just a longer image prompt.

For still images, I spend most of my time making the scene depictable. For video, I have to make the motion depictable as well. That means the clip brief must describe a start state, an action path, and an end state that could all plausibly happen in one 5 to 15 second shot. If two frames are too compositionally far apart, no amount of lyrical conviction will make them interpolate cleanly. They need an intermediate step, a different camera move, or a different pairing altogether.

That is why the plan for Snatches became a proper clip manifest rather than a handful of poetic prompts. Each section of the track was mapped against the verified running time. Each rooftop anchor was assigned a specific role. The generated performance clips were planned at roughly five to six seconds each, with the expectation that only two to four seconds of many of them would be used in the final edit. That is editorial thinking, not generator wishful thinking.

The final Snatches storyboard image: the rooftop protagonist stands with arms open as bright smoke-voices spread across the sky above the city at dawn.
Snatches 6. The final rooftop image earns its scale because the earlier frames stayed restrained. By this point the sequence knows where it is going.

The first implementation pass also reminded me why real workflow contact matters. As soon as the plan met the actual MiniMax H3 storyboard setup in ComfyUI, a few technical assumptions had to be corrected. Animation and audio instructions needed to be separated properly. Audio reference numbering turned out to be segment-local rather than storyboard-global. The manager's image-reference inputs behaved as autogrowing sockets rather than a fixed visible list. None of that is glamorous, but it is exactly the sort of friction that turns a vague method into a usable one.

That, really, is the point of this experiment.

We are not trying to prove that a machine can magically "make a music video" from a song title and a mood word. We are trying to build a workable visual process: one where the song's world stays coherent, the shots are physically plausible, the generator is given jobs it can actually do, and the inevitable corrections improve the method instead of blurring the target.

So the first milestone was not a finished clip. It was a decision: what kind of video Snatches should be, how its still images can become moving anchors, and where performance belongs inside that structure.

That is a much better place to start than motion for motion's sake.

In the next post, I’ll get into the performance side of the plan: how we started designing the band-world references, why reusable insert shots matter more than one giant prompt, and what the first H3 storyboard package taught us when it hit a real render workflow.