You will be able to line up the first frame, on-screen text and spoken words so the hook works with or without sound.
Watch the next few videos in your feed with the sound off, the way you might on a crowded train. Some still make complete sense: you can see what is happening and the words on screen tell you why it matters. Others turn into a person talking silently with a vague caption, and you move on without knowing what they were saying.
The hook you wrote in lesson 3.2 has to survive both situations. With sound on, the viewer hears your first line. With sound off, they only see the first frame and read whatever is written on it. Ideally all three, the picture, the text and the spoken line, say the same thing, so the hook works however someone is watching.
Platforms have built captioning tools and text overlays into their apps for good reason: a lot of viewing happens with the sound off, on public transport, in offices, late at night beside a sleeping partner. Nobody outside the platforms can tell you exactly how much, and it differs by platform and audience, so plan as if a large share of your viewers will never hear you.
That means the on-screen text has to carry the hook alone. If your spoken line is "The reason your blouse gaps isn't your size", and the text on screen says "New arrivals", a muted viewer has no reason to stay. Put the hook itself on screen, in the first frame, in words a viewer can read in a moment.
The on-screen hook does not have to be word for word what you say. It is often shorter. Farah's spoken line was "The reason your blouse gaps at the buttons isn't your size, it's the cut." Her on-screen text was "It's not your size." Short, clear, and it means the same thing.
The first frame is the picture a viewer sees before they have decided anything. If it shows you settling into the shot, adjusting the phone or taking a breath before speaking, you have used the deciding moment on nothing.
Start with something already happening. Hands already piping icing. The blouse already gapping as Farah sits down. Priya already holding up a worksheet with one answer circled in red. Something moving or unexpected gives the eye a reason to stay while the brain catches up.
A practical habit is to start recording a second or two before the action, then trim the start in editing so the video opens mid-movement. Lesson 5.3, Edit for pace: cut the start, the pauses and the end, shows how.
On-screen text has two practical jobs. It has to be readable at a glance, and it has to be visible.
Readable means short. A viewer deciding whether to stay will not read a paragraph. A few words, in a plain, large font with strong contrast against the background, works far better than a full sentence in a decorative script. If you would need to pause the video to read it, it is too long.
Visible means placed where the app will not cover it. Every short-video app puts its own things on top of your video: the like, comment and share buttons down one side, your account name and post caption along the bottom, and sometimes a search bar or menu at the top. Text placed in those areas is partly hidden. Keep your hook text in the middle area of the screen, away from the edges. Most editing apps and the platforms' own editors show guide lines or a preview so you can check, and the exact layout changes between apps and updates, so look at a preview on your own phone before you post.
The most common mistake at this stage is a hook where each part says something different. The picture shows a finished cake. The text says "Party season is here". The spoken line is "So many of you have asked about delivery." Each piece might be fine on its own. Together, the viewer cannot tell what the video is about, so they understand none of it.
When the three line up, they reinforce each other. Mei Ling's video on transporting cakes opened with a cake box on the back seat of a car, the text "Will it melt in the car?" and her voice saying "Here's how to get a cream cake home in a hot car without it melting." Muted or not, a parent knew in a moment exactly what the video would give them.
A quick test: describe each of the three parts in a few words. If the three descriptions are about the same thing, the hook works. If they are about two or three different things, choose one and change the others to match.
You do not need to film anything yet. For your strongest hook from lesson 3.2, sketch the first frame on paper or describe it in a line: what is in the shot, what is already moving. Write the on-screen text as it will appear. Write the first spoken line.
Then read the three side by side and check that they say one thing. That is your activity: sketch the first frame, the on-screen text and the first spoken line for your best hook, and check that all three say the same thing.
Sketch the first frame, on-screen text and first spoken line for your best hook and check that all three say the same thing.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).