A wide recording cut to a phone-shaped clip keeps a strip down the middle and loses the rest. That is not a flaw to fix - it is the format. The only question is which strip you keep, and most bad clips are bad because nobody chose.
Take the most common recording there is: two people either side of a table, framed nicely in widescreen. Crop the middle out of that and you get the gap between them - two half-faces at the edges and an empty table in the centre. The conversation is intact, the picture is useless.
The fix is not a better centre. It is to follow whoever is speaking, and to move when the speaker changes. That way each moment is framed on the person making it, and a cut between speakers looks deliberate rather than like a camera nobody was operating.
SnapDub's default does this: it finds the person talking and holds the frame on them, handing off when someone else takes over.
It also knows when not to move. Point it at a screen recording, gameplay or a slide and there is no speaker to follow - so it locks to a steady centre frame and stays there, instead of drifting around hunting for a face that was never in the shot. A wandering frame on static footage looks broken in a way a still one never does, and that steadiness is a decision, not a fallback.
A centre crop is a bet that the subject will stay where they were when you started.
Different footage wants different answers, so all four are one dropdown apart.
The default, and right for almost anything with people in it - interviews, podcasts, panels, a talk.
No tracking at all. Best when the subject genuinely is centred and still - a single presenter to camera, a fixed product shot.
The whole wide frame, shrunk to fit, with a soft blur filling above and below. For footage where losing the sides would lose the point - gameplay, a screen share, a chart.
The same, without the blur. Cleaner for anything where the blur would read as a mistake - text, diagrams, a slide deck.
Plenty of footage has nobody talking in it, and it still has a subject.
Footage you edited yourself does not hold still - it changes what it is about every time the shot changes.
A crop chosen once for a whole clip has to be a compromise between every shot in it, and a compromise between two good answers is usually a bad one.
So the question gets asked again at every cut: this shot, who is it about, where are they. The frame arrives already where it should be, the way it does when somebody is operating the camera and knows what is coming.
Compilations are made of clips that were vertical to begin with, so the recording you upload is a wide frame with a tall strip of actual footage in the middle and filler either side - black, grey, or a blurred copy of the picture. Crop that the ordinary way and you can end up with a clip of the filler.
SnapDub looks for the strip and stays inside it. And it looks per shot rather than once, because a compilation alternates: an ordinary wide clip, then a vertical one with bars, then another wide one. Deciding once for the whole video is how the bars get into half of it.
Where the strip is already the shape of the output - a vertical clip inside a widescreen frame usually is - there is nothing left to choose. The frame sits on the strip and stays there, which is the crop that footage was always going to get.
A camera in the corner, and the thing you are reacting to everywhere else.
This is the case where a better crop is not the answer at all. Whichever position you pick, a tall window is either on the camera or on the screen - they are at opposite ends of a wide frame, and one window cannot be at both ends.
So SnapDub stops cropping and starts arranging: it finds the camera box, gives it the top of the frame, and gives the rest to the screen. Your reaction stays readable, the thing you are reacting to stays watchable, and nothing is squeezed into a letterbox to make room.
The screen's half does not chase anything either. A screen has no subject that walks about - the interesting thing might be a kill feed in a corner or a map in another - so that panel takes the widest view it can of everything the camera box is not covering, and holds it.
Drop in a video. It takes a frame, shows you the slice that survives, and puts the app's own buttons back on top so you can see what they cover. Switch between TikTok, Reels and Shorts to see where each one sits.
Drop in a video and it takes a frame itself. A screenshot works too.
It is read on your machine - nothing is sent to us.Reading a frame…
That is a centre crop, which is what happens when there is nobody to follow. With speaker tracking the box moves to whoever is talking instead of sitting in the middle.
You are not editing for a rectangle. You are editing for a rectangle with buttons on it.
Every short-video app puts its own furniture over the video: the account name and caption along the bottom, a column of buttons down one side, a progress bar at the top. None of it is your content and all of it is on top of your content.
The switch above the preview draws each of the three in turn on your own frame. What you will notice is how little they differ: the side rail moves a little, one app starts its text slightly higher, two of them take a strip off the top. Frame for the busiest of them and the same clip is safe everywhere.
The exact sizes differ between apps and between phones, so pixel measurements go stale fast. What holds everywhere is simpler: the bottom quarter and one vertical edge are busy. Anything that has to be read - a face, a product, burned-in text - belongs in the middle band.
The habit worth building takes ten seconds: before posting, look at the clip and ask what would sit under a caption two lines long. If the answer is somebody's mouth, move the frame up a little.
Yes, and that is usually the right call - a vertical frame does not have room for two people side by side without shrinking both of them to nothing. The frame follows whoever is talking, so each moment is framed on the person making it.
It holds a steady centre frame. There is no speaker to follow in that footage, and a frame that drifts around looking for one is the thing that makes a gameplay clip look amateur - so it stays locked and still. When the sides of the picture carry the point, the fit modes keep the whole frame and fill the space above and below instead.
Yes. Every clip opens in the editor with the crop as its own control - change the mode, or zoom in and out to sit the subject where you want it - and rebuilding one clip does not touch the rest of the batch.
It is the safe answer and rarely the best one. A blurred-edge clip uses about half the screen for the actual picture, which is a lot to give away on a phone. Use it where the wide frame carries information, and crop properly where it does not.