How I Made My Book Trailers with AI
An honest, step-by-step account for other independent authors — the tools I used, the decisions they couldn't make, and why the difference matters.
I didn't know book trailers were a thing.
A reader on Threads left a comment asking where he could find the trailer for my novel. I didn't have one — and until that moment, I had never thought to make one. That single comment is the reason these two trailers exist.
I couldn't hire a film crew, a composer, and an editing house to make them. So I tried something else, and I want to show exactly how it worked — the tools, the process, and, more importantly, the judgment the tools could never supply.
Before anything else, one thing matters more than any tool below: the novels themselves are written entirely by me — word by word, over roughly five years, out of a life only I lived. AI played no part in the writing. Everything that follows is about the trailers — the short promotional videos — and nothing else.
It also helps to say where I'm coming from: alongside writing, I'm a visual artist. I paint, and years of thinking about composition, light, and restraint gave me an eye for this footage — knowing which frame was wrong, which was too much, and which quiet image would carry the feeling. The trailers weren't my first time directing an image; they were my first time doing it with these tools.
That same eye made the book covers. I designed both covers myself, in Photoshop and Canva — and the imagery on them isn't stock decoration. It draws on my own life: the places, objects, and textures behind the story. The instinct was the same one that shaped the trailers — begin from something personal and true, then build outward. The tools change; the source doesn't.
The tools, at a glance
ChatGPT — storyboard development and iteration (translating scenes from the novels into shot instructions). (paid subscription)
Google Flow — generating the individual cinematic clips. (paid)
Suno — composing the original music. (paid)
InShot — final editing, pacing, and sound. (paid)
An honest word on cost: none of this was free. All four are paid tools or subscriptions, so budget for that before you start — it's cheaper than a film crew by a wide margin, but it isn't nothing.
Scale: I generated roughly 20–25 four-second clips per trailer and kept only a fraction. The finished trailers run about 34 and 48 seconds — so most of what the tools produced never made the cut.
It started with the books, not a video tool
The process began not with a video-generation tool, but with the books themselves.
Before creating any footage, I had to decide what each trailer was meant to communicate. What the Gate First Knew is a story of departure, migration, discovery, and the search for belonging. What the Gate Remembers enters a later stage of life, when children have grown, parents are aging, familiar houses become quieter, and memory begins to carry a different weight.
I did not want the trailers to summarize the plots. I wanted them to recreate the emotional atmosphere of the novels.
Finding the visual language
The first creative decision was to keep the characters largely in silhouette.
This was partly practical, because AI-generated faces can change between scenes. More importantly, it was an artistic choice. I did not want a generated face to replace the Abir, Harleen, Arin, Mira, or other characters that readers had already imagined for themselves.
By showing people from behind, through doorways, against windows, or in partial shadow, the characters could remain recognizable without becoming visually fixed. Their emotions had to be communicated through posture, movement, distance, and gesture rather than facial expressions.
That limitation became part of the trailers' identity.
The recurring images came from the novels themselves: the green gate, tea, rain, architectural drawings, letters, notebooks, airports, trains, empty chairs, family tables, and houses that seem to retain the presence of those who once lived within them.
I also knew what these places were supposed to feel like, because I have lived in them. Kolkata, Boston, Berlin — the light, the streets, the interiors weren't abstractions to me. That's how I could tell when a generated scene was wrong: not just badly composed, but false to a place I remembered. I kept directing the footage back toward the versions I had actually seen and lived.
Developing the storyboards
I used ChatGPT as a creative development and iteration tool.
The process involved identifying the central emotional question for each trailer, testing different scene orders, and translating literary themes into visual moments. A paragraph in a novel cannot simply be copied into a video prompt. It must first be reduced to an image, an action, or sometimes a sound.
Loneliness after children leave home could become a dining table still set for four, with only two people sitting there. The passage of time could become an old clock whose pendulum slowly stops. Loss could be communicated through a quiet room, an empty chair, or family members gathered around a bed without any face being clearly shown.
Each scene was written as a detailed production instruction: location, the age and background of the characters, lighting, camera position, camera movement, architecture, weather, emotional tone — and the elements that should not appear.
The exclusions were often as important as the instructions. I repeatedly specified no visible faces, no exaggerated acting, no text generated within the scene, no fantasy effects, no melodrama, and no visual detail that would make the footage feel detached from the world of the novels.
Generating the footage
I used Google Flow to generate the individual cinematic clips.
Because the clips were created separately, consistency required constant attention. A gate could change shape. A house could become too ornate. A Bengali home could begin to resemble a palace or an abandoned mansion. A character could appear older, younger, taller, or dressed differently from one generation to the next.
Many clips were not used.
Some were technically impressive but emotionally wrong. Others looked attractive on their own but did not belong in the larger narrative. Occasionally one small movement inside an otherwise imperfect generation — a hand reaching toward a notebook, rain beneath a gate, a person pausing after a phone call — was strong enough to keep.
The process was generate, review, reject, rewrite, and generate again. AI accelerated the creation of possibilities. It did not remove the need to evaluate them.
Creating the music
I used Suno to develop original music for the trailers.
The two books needed different musical identities. The first trailer needed movement, departure, and discovery. The second needed maturity, tension, loss, memory, and return.
As with the visuals, the first result was rarely the right one. I experimented with mood, instrumentation, pacing, and emotional progression. The music had to support the images without overwhelming the literary restraint of the stories. I then shaped the tracks around the evolving edit. The music was not background decoration; it became part of the storytelling architecture. What helped was , my wife is a singer and has has great ears for music. I asked her for feedback throughout the process.
Editing the final trailers
I assembled the footage in InShot. This was where the generated material became an actual trailer.
I selected the clips, shortened them, changed their order, adjusted transitions, aligned visual changes with the music, and decided when a scene should be allowed to breathe and when it should disappear almost immediately.
Sound did a surprising amount of the work. A phone vibration, a gate creaking, rain against a window, suitcase wheels across an airport floor, a water drop falling into a metal bowl, a clock suddenly going silent — any of these could create more tension than another visual scene.
The final trailer was not a chronological collection of generated clips. It was built through pacing, contrast, repetition, and omission. Some of the most attractive footage was cut because it slowed the narrative. Other scenes were reused briefly because their recurrence established a motif. The editorial question never changed: does this moment make the viewer want to keep watching?
What the tools did — and what I did
The tools made it possible to explore ideas that would otherwise have required actors, locations, cinematographers, musicians, editors, and a far larger budget.
But the tools did not decide:
what emotional story the trailer should tell;
which moments from the novels mattered;
why the characters should remain in silhouette;
which generations were represented;
how Kolkata, Boston, Berlin, and the other locations should feel;
which generated scenes should be rejected;
where the suspense should begin;
when the music should rise or recede;
or when the trailer was finally complete
Those remained creative decisions.
For an independent writer, the distinction matters. The technology lowered the barrier to production, but authorship stayed in the direction, the selection, the interpretation, and the final edit. And it goes without saying — though I'll say it anyway — that the same is true, many times over, of the novels the trailers were made for. Those were mine long before any of these tools entered the process.
The tools generated material. I made the trailer.
Watch the trailers
What the Gate First Knew — Official Book Trailer:https://youtu.be/khnPyN8-nuU
What the Gate Remembers — Official Book Trailer:https://youtu.be/H57nq3W8kKg
I am also dropping the embedded videos at the bottom , or you can explore on other pages of my website.
For other indie authors: five things I'd pass on
Keep faces out. Silhouette solves the consistency problem and, better, leaves your readers' imagined characters intact.
Exclusions matter as much as instructions. Half of every prompt was what not to generate — no melodrama, no on-screen text, no fantasy effects.
Reject generously. Most clips won't make it. Judge each one against the whole trailer, not on its own.
Let sound carry the tension. A creak, a vibration, a sudden silence often does more than another image.
Use the eye you already have. If you draw, paint, photograph, or design — even your book covers — that judgment transfers directly. It's what tells you which generated frame is wrong when the tool can't.
Stay the author. The tools generate options; the story, the taste, and the final cut have to stay yours. That's the part no tool replaces — and the part worth being transparent about.
A closing thought
I've spent a career watching technologies arrive, unsettle people, and then quietly become ordinary — the things we once argued about become the tools we stop noticing. I think we're at one of those thresholds again.
I don't see these tools as a replacement for the writer, the painter, or the composer. I see them as companions — instruments that shorten the distance between an idea and its expression, so a person with something to say can say it more fully, and more people can say anything at all. That's what they did for me: they let one author, at a laptop, direct something that once required a crew.
I understand the unease — every real change in how we make things has met it, and some of that caution is worth taking seriously. But my honest belief, having lived through more than one of these shifts, is that this will settle the way the others did: not as the end of human creativity, but as a new way of expressing it. The work still has to be yours. The tools only widen who gets to make it.
The novels — What the Gate First Knew and What the Gate Remembers — are part of The Gate of Belonging*, a five-book literary saga. You can read selected lines from the books here.*