Answer in brief
Vertical is a different composition, not a horizontal frame with the sides trimmed off. A working guide to the safe areas the interface takes, the opening second, where captions can and cannot sit, and why a horizontal archive rarely survives the reframe.
The short answer: what to settle before the camera rolls: A vertical frame stacks subjects above one another, so…
Vertical video comes down to four decisions taken in advance: where the subject sits inside a narrow frame, which strips of that frame the platform interface will cover, what the viewer sees in the opening second, and which line the captions occupy. Everything else follows from those four answers.
A fifth decision is where the material comes from. Footage shot for a horizontal frame can be moved into vertical, but that is a rebuild of the composition rather than a trim of the edges, and some shots do not survive the move.
For a business the conclusion is plain. Vertical belongs in the shooting brief, not only in the editing brief. Fixes made at the edit stage cost less than a reshoot, yet no edit creates pixels that were never recorded in the first place.
The rest of this article works through it in order: composition, safe areas, the opening second, captions, salvaging horizontal footage, capture, export, and how the work itself is packaged on the /services/video-production page.
Vertical is a composition, not a crop
A horizontal frame lays things out beside each other: a person on the left, the product on the right, the room behind them carrying the context. A vertical frame stacks them instead, and "next to" turns into "above and below".
That changes how headroom behaves, how movement reads and where attention travels. The eye runs down a narrow column rather than across a wide band, which leaves the middle third of the frame carrying the weight of the shot.
The first consequence is shot size. What worked as a mid shot in horizontal usually has to come closer in vertical, or the subject shrinks towards the centre while the top and bottom of the frame fill with nothing in particular.
The second consequence is the background. A narrow frame gives it far less area, so it reads as a block of colour rather than as a scene. The location detail the shoot was booked for frequently does not fit at all.
Safe areas: the strips the interface takes
The platform draws its own furniture over your video: a status strip along the top, the account name, the description and system captions along the bottom, a column of action buttons down one side. Those regions belong to the app, not to your composition.
The exact height of each strip differs between platforms, between app versions and between handsets, so no single pixel figure is correct everywhere. Check those values against each platform's current specification instead of carrying them around in your head.
The working method is to keep the top and bottom bands free of meaning. No text there, no logo, no faces: leave wall, floor or the edge of a table, which is to say anything you would not mind handing over to an interface you do not control.
The side column deserves the same treatment wherever a platform puts buttons there. The fix belongs to the shoot — you move the composition by choosing the camera position, rather than sliding the whole picture sideways in the edit.
The opening second: what has to be on screen immediately
A vertical feed scrolls, and the decision to watch or move on is taken on the first frames. So the subject of the clip has to be visible from the start: the product, a face, an action, something recognisable without a sentence of setup.
A hard editing rule follows from that. Idents, slow push-ins, the run-up before the point and the logo on black all move out of the opening. The logo does not disappear anywhere; it simply stops being the first thing anyone sees.
Work on the assumption that the opening frame may be reused as the cover image in a profile grid. That makes it a photograph as well as a frame: a clear subject, a calm background, no motion blur and no word caught halfway through.
The first line of text is a separate job. One short phrase in the middle third does more than a full paragraph, because nobody reads two sentences while deciding whether to stop or keep scrolling.
Where captions can sit, and where they cannot: Vertical is a different composition, not a horizontal…
In a vertical frame captions live in the middle third, low but above the platform's own bottom strip. Any lower and they slide under the description and the system text; any higher and they land across a face or across the product.
That constraint is useful back at the shoot. Once the crew knows the lower third is reserved for text, the subject is framed slightly above centre and the captions stop competing with the picture for the same pixels.
Lines have to stay short. Two lines of a few words each can be read while the shot is moving, four cannot, and a line break dropped into the middle of a phrase makes the viewer re-read and lose the sentence.
Burned-in captions are a decision of their own. They are visible always and on every platform, but they cannot be switched off and cannot be translated without rebuilding the edit, so a multilingual campaign needs a separate export per language.
Legibility: contrast, type size and the background underneath
WCAG 2.2 asks for captions on prerecorded audio and sets a minimum text contrast of 4.5:1 at normal size and 3:1 for large text. Those contrast rules are written for text on a page rather than for lettering inside a video frame, yet the ratio stays a sensible target to design burned-in captions against.
The difficulty in vertical video is that the background under the text changes on every frame. White lettering disappears against a bright wall, so the caption needs a plate, an outline or a soft shadow: anything that holds the contrast steady while the picture moves.
Check type size on a phone at arm's length rather than on a monitor. A size that looks generous on the editing timeline sits right at the edge of readability on a six-inch screen held in one hand.
Weight counts for as much as size. A thin face with weak contrast falls apart when the platform re-encodes the file, because fine strokes lose more to compression than solid shapes do.
Why horizontal footage rarely survives the reframe
The first reason is arithmetic. A 1920×1080 frame cropped to 9:16 while keeping its full height leaves a strip 1080 × 9 ÷ 16 wide, roughly 608 pixels. Getting to 1080×1920 then means enlarging that strip.
Enlargement invents no detail. Softness, grain and compression artefacts grow along with the image, and on a phone that shows on faces, on signage and on the fine edges of a product.
The second reason is composition. In a horizontal shot the subject often sits near one edge with something that matters near the other. A vertical frame is forced to choose one of them, and a scene built on their relationship falls apart.
The third reason is movement. A pan that leads the eye across a wide frame pushes the subject straight out of a narrow one. You are left either tracking the motion by animating the crop or cutting the shot earlier than it was designed to end.
What can still be rescued from a horizontal archive
Close-ups survive the move. A face, a pair of hands, a product on a table — anything already occupying the centre of the frame carries over into vertical with barely any loss, provided the original was recorded with resolution to spare.
A 3840×2160 source cropped the same way leaves a 1215×2160 strip, and that fits into 1080×1920 with no enlargement at all. There is the real argument for shooting at a higher resolution: headroom for reframing rather than quality for its own sake.
Locked-off shots help as well. With no pan and no handheld drift, the crop can be moved deliberately inside the frame, and a slow travel across a still image reads as a device rather than as damage control.
What does not survive is the wide establishing shot, the group scene and anything built for width simply because width was available. Reshooting those as a short vertical scene is more honest than stretching them and apologising for the result afterwards.
Shooting for vertical: what to agree in advance
Write the format into the shooting brief: device, orientation, resolution, frame rate and how much room to leave around the subject. One paragraph agreed before anyone travels to the location saves a working day in the edit.
The habit worth building is a separate vertical take instead of a plan to reframe later. Same words, same scene, a different camera position and a tighter composition; the take costs minutes and the difference on screen is obvious.
When there is one camera and a second take is impossible, centre the subject and leave air above and below. Not for elegance, but so that a 9:16 rectangle has somewhere to land inside the recorded frame.
Record clean audio even for a clip meant to be watched without sound. That track is what the captions are transcribed from, and it is what the version somebody plays with the sound on will depend on.
Export: container, codec and versions per placement
A container and a codec are not the same thing. As the MDN documentation explains, the container — MP4, for example — holds the streams, while the codec determines how the picture and the sound inside them are compressed. Platforms have requirements about both.
MP4 carrying H.264 video and AAC audio is a pairing accepted broadly across social placements, and one that needs no explaining on the client side. Unusual combinations are kept for the archive, not for publishing.
Export the frame at 1080×1920 and resist the urge to hand over a much larger master. The platform re-encodes the file regardless, so extra weight does not turn into extra quality once it has been through their processing.
There are usually several versions: with burned-in captions and without, with music and with a clean track, sometimes with the text placed differently for different interfaces. In the Single Reel package, exports for every vertical placement are part of the work.
What it costs and what each package covers
Single Reel costs $70, is delivered in 1-2 working days and carries 1 round of revisions. It includes captions, a licensed track and exports for every vertical placement. Shooting, scripting and actors are not included.
Reels Pack costs $90 for up to 5 edits, delivered in 1-2 working days from footage delivery, with 2 rounds on the edit. Filming is not part of it: you supply the footage.
Reels Pack Express costs $150 for 1 working day, a 24-hour turnaround from footage, and 1 round of revisions. Shooting is not included, and neither is licensing beyond the supplied track.
Campaign Edit costs $180 for a hero cut plus short derivatives, in 3-5 days, with 2 rounds on the cut and 1 on the grade; licensed music is not included, as the licence is purchased by you. Monthly Content Engine costs $290/mo on a monthly cycle with 30 days notice to stop, the volume agreed at the start of each cycle, and filming days and paid stock excluded.
What to decide before you hand over the footage
Four answers remove much of the risk before any work starts: the orientation of the source files, their resolution, whether a vertical take exists, and which language the captions are in. Those four decide whether the job is an edit or a rescue.
The fifth answer is the list of placements. Different platforms occupy the edges of the frame in their own way, and knowing the list in advance means the text is positioned once instead of being moved after the first post goes live.
The sixth is music. A track sets the rhythm the cuts are built on, so it is chosen at the start rather than at the end: swapping music into a finished clip means rebuilding the edit, not replacing a file.
Taken together, those answers already amount to a short brief for the service page at /services/video-production, and one question is left — which package matches the volume you actually need.
Practical checklist
- Write orientation, resolution and frame rate into the shooting brief before anyone travels to the location.
- Shoot a separate vertical take instead of planning to reframe a horizontal shot later.
- Keep the top and bottom bands of the frame free of text, logo and faces.
- Put the subject of the clip into the opening second and move the ident out of the start.
- Check caption size and contrast on a phone at arm's length rather than on a monitor.
- Collect the list of placements before the edit and order a separate export for each one.
Questions and answers
Can you simply crop a horizontal video into vertical?
You can crop it, but that rebuilds the composition rather than adjusting the edges. A 1920×1080 frame cropped to 9:16 leaves a strip roughly 608 pixels wide, which is then enlarged to 1080×1920. Close-ups carry over acceptably; wide shots and scenes with two subjects at opposite edges do not.
What does a vertical clip cost and what is included?
The Single Reel package is $70, delivered in 1-2 working days with 1 round of revisions. It includes captions, a licensed track and exports for every vertical placement. Shooting, scripting and actors are not included. Reels Pack is $90 for up to 5 edits in 1-2 working days from footage delivery, with 2 rounds on the edit.
Where should captions sit so the interface does not cover them?
In the middle third of the frame, low but above the platform's own bottom strip. Any lower and they slide under the description and system text; any higher and they land on a face or the product. The exact strip heights differ between platforms and app versions, so check the current specification.
What has to be in the opening second?
The subject of the clip: a product, a face or an action recognisable without a sentence of setup. Idents, slow push-ins and a logo on black move further into the running time. Build the first frame as though it were also a cover image: a clear subject, a calm background, no motion blur.
What format should the finished clip be delivered in?
MP4 carrying H.264 video and AAC audio is accepted broadly across social placements. As the MDN documentation explains, a container and a codec are separate things: the container holds the streams, the codec determines the compression. Export at 1080×1920; a heavier master gains nothing because the platform re-encodes the file.

