ClipEmber burns captions into the pixels of the video itself — no SRT file to upload separately. Transcription returns word-level timestamps, which is what makes word-by-word highlighting possible; nine styles, adjustable size and position, 90+ languages.
Word-level timestamps, nine caption styles, uppercase pop or subtle sentence case — burned into the video, no SRT juggling.
Open the app to run it — 5 free videos every month, no card required.
The transcriber produces word-level timestamps — that's what makes karaoke highlighting possible.
Nine templates from BOLD POP to subtle sentence case; adjust size, position and words per line.
Captions are rendered into the pixels — they show up on every platform, no separate subtitle file.
Run on 2026-08-17 · source: Tutorial SCREENCAST-O-MATIC 2020 — Cómo grabar la pantalla by EducaTIC · Creative Commons Attribution (CC BY)
This one is not a fair timing: we had six runs queued at once and this was last in line, so crop planning and rendering spent most of that time waiting, not working. Our quiet-queue runs on this site finish in 4 to 5 minutes.
We ran this one in Spanish on purpose, because "multilingual" is easy to claim and easy to check. The captions came back in Spanish, and so did the clip titles the model wrote for them — "¿El límite de grabación gratuita te corta? Usa este sencillo truco." Nothing was translated into English along the way, which is the behaviour you want: captions keep the language people are actually speaking.
Request 155fbffe-890e-4137-abe2-2d237704ba0c — every number above comes from that run, not from an average.
90+ languages via multilingual speech recognition; captions keep the original language and casing rules you choose.
Yes — position and size are sliders in the editor, per clip.
Karaoke word-highlighting is the default; static-line templates are available too (Subtle, Inkbox).