• Skip to primary navigation
  • Skip to main content
  • Skip to primary sidebar
  • Skip to footer

Social Media Examiner

Your Guide to the Marketing Jungle

  • 🔥 Free Newsletter
  • 🎙️ Podcasts
    • Social Media Marketing Podcast
    • AI Explored Podcast
    • Our YouTube Channel
  • 🌟 AI Society
  • 🗓️ Marketing Conference
  • 🤖 AI Conference
  • 😊 NoteGo App
  • 👋 About Us
    • Marketing Events
  • Search
  • Social Media Marketing WorldImprove your strategy & find your next big marketing ideasDISCOVER WHAT YOU'VE BEEN MISSING

    How to Create Pro-Quality AI Videos With Gemini Omni

    by Michael Stelzner / September 22, 2026

    Want to create professional-looking video content without filming yourself or hiring a production crew? Wondering how to use AI-generated video hooks to stop the scroll and boost engagement on social media?

    In this article, you'll discover how to use Google's Gemini Omni to create an AI avatar, produce scroll-stopping video hooks, edit existing footage, and assemble longer videos from 10-second clips.

    This article was co-created by Eve Whitaker and Michael Stelzner. For more about Eve, scroll to the end of this article.

    What Is Gemini Omni and How to Access It

    Google's Gemini Omni, officially called OmniFlash, is the successor to Veo3, Google's first AI video generator. Omni is built on Google's world models, which were trained on YouTube's massive video library. That training gives it an understanding of movement, environments, people, and speech inflections that sets it apart from other AI video tools.

    Omni is available in several places, making AI video creation accessible at different levels.

    The simplest way to access it is via the Gemini app, available on both desktop and mobile. Inside the Gemini chat interface, tapping the plus button reveals a “Video” option that runs Omni behind the scenes.

    For more advanced features, particularly video editing and modification, Google Labs offers a fuller set of tools. Omni is also available through third-party aggregator platforms such as Open Art and Higgs Field, which bundle multiple AI models into a single interface.

    Eve Whitaker recommends that beginners start with the Gemini app and a $20/month subscription to explore what Omni can do. The ultra tier runs about $200/month and includes access to all of Google's tools, but most users won't need that right away. Credits are consumed per generation, with higher resolution and longer clips using more. Starting at the lowest tier and buying additional credits as needed keeps costs manageable.

    One limitation to be aware of is that the Gemini app imposes generation limits. Early users could generate roughly five or six videos before being throttled, though the limit resets within an hour or two. Google Labs and aggregator platforms offer more generous allowances.

    Pro Tip: For anyone serious about AI video production, Eve suggests pairing a direct Google Labs subscription with a third-party aggregator. The aggregator provides access to multiple AI models from a single dashboard, while Google Labs unlocks Omni's full editing capabilities. The Gemini app alone doesn't support video editing or modification.

    #1: How to Create an AI Avatar With Gemini Omni

    A clone, like those created by HeyGen, accepts a long script and generates a talking-head video in one pass. An avatar is different. Eve calls it an “AI twin”: a digital representation of a real person that can be placed into any scene, environment, or scenario through text prompts. The key constraint is that Omni generates 10-second clips of the avatar, which then need to be edited together for longer content.

    Setting Up the Avatar

    Each avatar is tied to the user's account. Unlike Sora, where creators could share their avatars for others to use in their own videos, Omni restricts avatar access to the account that created it. No one else can call up or use another person's avatar. Clients who want their own avatars need to create them on their own accounts. And, while there's no option to maintain multiple avatars simultaneously, your avatar can be deleted and recreated at will.

    The process takes about five minutes. Open the Gemini app on a phone, tap the plus button, and scroll down to “avatar.” The app walks through a Face ID-style capture sequence: look left, look right, look up, look down. It then displays nonsensical sentences to read aloud so it can capture your voice. The deliberately strange text is designed to capture natural speech patterns without the speaker overthinking delivery.

    Once the avatar is created and named, it syncs across platforms. This means that a user who builds their avatar on their phone can also call it up in Google Labs on their desktop by typing the @ sign followed by the avatar's name.

    Tips for Creating a High-Quality Avatar

    The recording environment makes a significant difference in avatar quality, and the AI won't flag problems during capture. If the lighting is poor, the camera has a smudge, or the background is noisy, Omni will simply fill in the gaps with generated content, which may not look or sound like the real person.

    Go Deeper Than an Article Can Take You

    Social Media Marketing World

    Join thousands of small-business marketers at Social Media Marketing World — three days of AI, social media, and marketing strategy taught by world-class experts. April 1–3, 2027 in Anaheim.

    Every session is pitch-free and built around tactics you can use yourself, no big team required. Your All-Access ticket includes an extra day of workshops, full access to AI Business World, and recordings of everything.

    GET MY ALL-ACCESS TICKET — SAVE $800

    Lighting: Natural light works best. Stand in front of a window or to the side of one. Filming with a window behind creates a silhouette that degrades the capture. Eve recorded hers in her living room with natural light and reports excellent results.

    Audio: Minimize background noise. The AI needs a clean voice capture to reproduce speech accurately. If it can't hear something clearly, it will fabricate it.

    Clothing: Whatever is worn during the recording becomes the avatar's default outfit. A red shirt and necklace worn during setup will appear in every generation unless the prompt specifies something different. Eve suggests choosing an outfit that works across multiple scenarios, since prompting a change is easy, but the default will recur.

    Hats and glasses: Avoid wearing them during setup. The AI has no data on what's underneath a hat, so removing one via a prompt produces unpredictable results. The better approach is to record without a hat, then upload a photo of the desired hat as a reference image and prompt the avatar to wear it.

    Voice tone: Speak naturally. If the recording is animated or exaggerated, the avatar will default to that energy level in every generation, making it harder to prompt calmer or more conversational deliveries later. Recording in a natural speaking voice preserves flexibility, since animation can always be added through prompts.

    Background: The recording background matters less than it might seem. When the avatar is prompted into a specific environment, such as “put me on a beach in Mexico,” Omni knocks out the original background entirely and replaces it with the prompted scene. A clean, well-lit recording is still important for accurately capturing the face and voice, but the physical backdrop doesn't need to be polished.

    #2: The Four-Element Prompting Formula for AI Video

    Effective AI video prompting follows what Eve Whitaker calls the four-element prompting formula: subject, action, environment, and camera. Omni responds well to natural language, meaning the days of formatting prompts in JSON or coded syntax are over.

    Subject: Who or what appears in the video. When using an avatar, this means calling it in with the @ sign and the avatar's name.

    Action: Describes what the subject is doing: dancing, walking, talking to the camera, parachuting onto a lawn.

    Environment: Sets the scene: a street, a beach in Mexico, a kitchen countertop, a park full of dogs.

    Camera: Tells Omni where the camera sits relative to the subject and how it moves.

    A prompt using all four elements might read:

    @Eve Whitaker is dancing in the street.

    It is raining tennis balls.

    A bunch of dogs are walking around her.

    She walks up to the camera and says, 'Did that get your attention?'

    Eve advises starting simple and iterating. If the first generation places the subject too far from the camera, the next prompt adds “she's standing close to the camera” or “she walks up to the camera.” Each generation reveals how Omni handles different types of instructions.

    Omni also supports iterative refinement within a single concept. After generating a clip, the next prompt can specify:

    Keep everything else exactly the same, but change the shirt from blue to red.
    

    Omni will reproduce the scene with only that one element adjusted. This makes it possible to test variations without re-prompting the entire concept from scratch.

    Tips to Handle Scripts and Dialogue

    Including a script in the prompt is important. Without one, Omni will generate its own dialog, which sometimes produces interesting accidents but more often produces unusable audio.

    The most common prompting mistake is giving Omni more than 10 seconds of dialog. Rather than cutting the speaker off mid-sentence, the AI compresses all the dialog into the time frame, producing rushed, garbled speech toward the end.

    Eve's workflow for scripting starts with an LLM. She tells ChatGPT, Claude, or another tool something like:

    I want to make a 30-second commercial in Omni, and it can only generate 10-second clips.

    Help me break down the script.

    She describes what she wants, often using voice input with rough, unpolished language, then asks the LLM to suggest clarifying questions. After refining the idea collaboratively, she asks the LLM to turn the final concept into a prompt formatted for Omni.

    AI Business Society

    Tired of Guessing Which AI Moves Matter?

    New AI tools and strategies launch every week — and figuring out which ones matter takes hours you don't have.

    The AI Business Society filters the noise for you, with expert-led training you can put to work the same day — plus a community of marketers sharing what's actually working.

    I'M READY FOR REAL AI RESULTS

    Tip on Camera Direction for Non-Filmmakers

    Camera placement shapes how the viewer feels about the subject.

    An extreme wide shot establishes the full environment and context. A medium shot is conversational. A close-up creates intimacy and is ideal for delivering a message directly. An extreme close-up generates discomfort or tension, which can be used deliberately to grab attention.

    A “hero shot,” where the camera is positioned below the subject looking upward, makes the subject appear larger than life. A camera positioned above the subject does the opposite, making them appear smaller.

    Eve's practical advice for learning camera language is to start watching television and movies, paying attention to what the camera does during emotionally resonant moments. When a scene triggers an emotion, noticing whether the camera is tight on a face, pulling back to reveal a setting, or shifting angles builds an instinct that transfers directly to prompting. Once those patterns are noticed, they become impossible to unsee.

    #3: How to Create Scroll-Stopping Video Hooks

    The primary use case Eve recommends for Gemini Omni is creating attention-grabbing hooks for short-form video across TikTok, Instagram, YouTube Shorts, and other platforms.

    The concept is straightforward: film a standard talking-head or direct-to-camera video the traditional way, then use Omni to generate a visually striking 2-3-second hook to open the video.

    Examples of hooks Eve has created include her avatar getting Nickelodeon slime dumped on its head, getting cake shoved in its face while wearing an evening gown, standing in a park with tennis balls raining from the sky, riding a buffalo in front of a house listing, parachuting onto a lawn next to a for-sale sign, and walking across a kitchen countertop as a six-inch avatar saying, “The only thing small about this kitchen is me.”

    The speed of creation is part of the appeal. Eve created the buffalo-riding hook while waiting in line for Coffee: she opened the Gemini app on her phone, called in her avatar, prompted it to ride a buffalo in front of a house listing using a Zillow photo as a reference image, and had the finished clip by the time she picked up her order. That mobile-first workflow makes it possible to produce polished hooks anywhere, without a desktop or editing software.

    How to Brainstorm Hook Ideas

    Eve uses an LLM to generate AI video hook concepts rather than brainstorming them alone.

    The process starts by describing the topic and asking for creative visual hooks. The initial suggestions from any LLM tend to be predictable, which is where Eve's “Captain Obvious” technique comes in. The “Captain Obvious” technique is a method for steering an LLM away from its default, predictable suggestions toward more creative and distinctive ideas. The term comes from television production, where producers would pitch story ideas, and the team would flag the obvious ones.

    When an LLM suggests a predictable visual, such as a person sitting at a computer with too many tabs open to represent “AI overwhelm,” Eve tells it:

    That's Captain Obvious.

    Think outside the box.

    Use metaphors or analogies.

    This redirect consistently produces more creative and distinctive concepts.

    How to Match Hooks to Existing, Traditional Video

    When pairing an AI-generated hook with a traditionally filmed talking video, visual continuity matters. If the avatar appears in a different outfit than the talking video, the transition feels jarring. Eve's solution is to take a screenshot from the video and upload it as a reference image, then tell Omni to dress the avatar in the same outfit.

    The hook clips and talking video are then edited together in CapCut, Instagram's Edits app, or a similar tool.

    Pro Tip: Eve has tested as many as seven different hooks for the same talking video using Instagram's trial reels feature, then put promotion behind whichever hook performed best.

    #4: How to Edit Existing Video With Gemini Omni

    Gemini Omni can modify existing video using natural language prompts, making edits that once required a professional editor accessible to anyone. These capabilities, available through Google Labs and aggregator platforms, go well beyond generation. Real footage, whether shot on a phone or generated by AI, can be uploaded and transformed with a text description of the desired change.

    The modifications include changing environments in real footage. For example, a summer park scene showing feet walking along a path and a camera panning up to leafy trees can be converted to winter with a single prompt. Omni keeps the shoes, the pants, and the movement identical while stripping leaves from trees and adding snow. The result looks realistic.

    Motion graphics can also be added to real video by describing them in a prompt. A cooking video can show animated text slates labeling ingredients as they appear on screen, and the result looks like a motion graphics artist produced it.

    Other editing capabilities include changing a subject's outfit, removing objects from a scene, converting a person into a cartoon version of themselves, and changing the background entirely. For example, a talking-head video recorded at a desk can be re-homed on a beach, in a news studio with on-screen text, or any other environment.

    Pro Tip: Video uploaded for editing must be cropped to 10 seconds or less. Omni provides a built-in cropping tool, so longer footage can be trimmed to the relevant segment before applying modifications.

    #5: How to Assemble 10-Second Clips Into Longer Videos

    The 10-second generation limit doesn't prevent the creation of longer AI video content. It requires pre-production planning and a third-party editor.

    For a 90-second video, the approach is to generate nine 10-second clips and edit them together. The key is planning what the last frame of each clip looks like and how it transitions into the first frame of the next clip. Eve suggests using an LLM for this planning, giving it the full concept and specifying constraints:

    I'm making a 90-second video using nine 10-second clips.

    I need them to cut together smoothly.

    I don't want any jump cuts.

    The LLM will suggest transitional strategies and help plan each segment. For example, Eve recently created an avatar as Gilligan on an island, saying he'd been stranded for 62 years. The LLM suggested inserting a montage of island life between clips, produced across two additional 10-second generations.

    Alternatively, Google Labs offers a feature that lets the next prompt build on the clip just generated, which, in theory, helps with continuity. In practice, Eve finds the results inconsistent and prefers to generate clips independently, then edit them together in CapCut.

    Pro Tip: Eve suggests using 720p resolution for social media content to conserve credits. Reserve 1080p for content that will be displayed on large screens at conferences or on television. Videos can also be upscaled after generation using third-party tools, so starting at 720p and upscaling later is a cost-effective workflow. Omni supports both vertical (9:16) and horizontal (16:9) aspect ratios, specified in the prompt.

    Eve Whitaker, an AI video strategist and founder of VEI Creative, helps small businesses create professional-quality video content using AI tools. Learn more about the Video Spark membership community on Skool. Follow them on Facebook, Instagram, and LinkedIn.

    Other Notes From This Episode

    • Connect with Michael Stelzner @Stelzner on Facebook and @Mike_Stelzner on X.
    • Watch this interview and other exclusive content from Social Media Examiner on YouTube.

    Listen to the Podcast Now

    This article is sourced from the AI Explored podcast. Listen or subscribe below.

    Where to subscribe: Apple Podcasts | Spotify | YouTube Music | YouTube | Amazon Music | RSS

    ✋🏽 If you enjoyed this episode of the AI Explored podcast, please head over to Apple Podcasts, leave a rating, write a review, and subscribe.


    Stay Up-to-Date: Get New Marketing Articles Delivered to You!

    Don't miss out on upcoming social media marketing insights and strategies! Sign up to receive notifications when we publish new articles on Social Media Examiner. Our expertly crafted content will help you stay ahead of the curve and drive results for your business. Click the link below to sign up now and receive our annual report!

     
    Find Your AI Level

    Where Do You Actually Stand With AI?

    Most marketers use AI, but they don't know what's holding them back from real results.

    Our 3-minute quiz pinpoints the one gap keeping you stuck — and gives you the exact next step to fix it.

    GET MY RESULTS

    Tags: AI Explored podcast

    About the authorMichael Stelzner

    Michael Stelzner is the founder of Social Media Examiner and Social Media Marketing World—the industry's largest conference. He's also the founder of the AI Business Society and the AI Business World conference. Michael hosts the Social Media Marketing Podcast and the AI Explored podcast, and is the author of the books Launch and Writing White Papers.
    Other posts by Michael Stelzner »

    Get Social Media Examiner’s Future Articles in Your Inbox!

    Get our latest articles delivered to your email inbox and get the FREE Social Media Marketing Industry Report (44 pages, 50+ charts)!

    Industry Report Cover

    Worth Exploring:

    Facebook

    Marketing Help Explore More →

    Instagram

    Marketing Help Explore More →

    YouTube

    Marketing Help Explore More →

    Linkedin

    Marketing Help Explore More →

    AI

    Next Frontier Explore More →

    Social Media Marketing Industry Report

    Get Free Report →

    Social Marketing Trends

    The data you've been missing!

    Need a new plan? Discover how marketers plan to change their social activities in the 18th annual Social Media Marketing Industry Report. It reveals what marketers have planned for their social activities, content marketing, and more! Get this free report now and never miss another great article from us. Join more than 385,000 marketers!

    Simply click the button below to get the free report:

    Footer

    Your Guide to the Marketing Jungle
    Copyright © 2026 Social Media Examiner®
    All Rights Reserved. Terms of Use | Privacy Policy.

    Helpful Links

    • About us
    • Our content via email
    • Our podcasts
    • Our YouTube channel
    • Our live show
    • Our social media marketing industry report
    • Our AI marketing industry report
    • Sponsorship opportunities
    • RSS