Skip to content

Auto

Combine fragmented subtitles from Whisper and YouTube into clean, readable cues. Stop the rapid-fire flashing of one-word subtitle blocks.

Drag & drop your subtitle file here

Drag and drop subtitle files here or click to browse

SRT, VTT, ASS, SSA, LRC, SBV, TTML, SCC, STL, SMI, TXT, JSON
Browser-Based Processing · Files Stay on Device
Core Processing Architecture

What the Auto-Merger Does

Engineered for video editors and localization teams. Helps enforce common subtitle timing rules and formatting.

Detects adjacent short cues separated by tiny pauses (configurable, default under 400ms).
Merges them into a single cue only if the combined text fits within your character limit.
Respects a maximum duration so merged cues never become too long to read.
Sets the merged cue's end time to the last cue's end, preserving timing accuracy.
Perfect for cleaning up Whisper, YouTube auto-captions, and other speech-to-text fragmentation.
Adjustable thresholds: max characters per line, pause threshold, and max merged duration.
Works with SRT, VTT, ASS, SSA, LRC, SBV, TTML, SCC, STL, and SMI formats.
Non-destructive: preview the merge count and result before downloading.
Live Preview & Execution Pipeline

Auto Live Execution Demo

dirty.srt
Raw Input
00:01.200♪ [Upbeat Synth Music] ♪
00:04.500Hello... world!!
00:08.100Duplicate line
00:08.100Duplicate line
cleaned.srt
Cleaned
00:04.500Hello, world!
00:08.100Duplicate line
00:11.400Verified syntax check

In-Memory Processing

File parsing, mutations, and time shifts for standard tools run in browser memory without sending subtitle files to our servers.

Millisecond Precision

Millisecond-accurate timecode parsing and arithmetic prevent synchronization drift across desktop and web players.

Fault-Tolerant Engine

Automatic multi-encoding detection and syntax sanitization gracefully handle malformed cues and unexpected linebreaks.

Production Workflows

Who Needs Auto-Merging?

Real-world scenarios across video post-production, course transcription, and global localization.

SCENARIO #1

Whisper & AI Transcription

AI transcription tools fragment speech into tiny one-word cues that flash rapidly. Merge them into natural sentences automatically.

SCENARIO #2

YouTube Auto-Captions

YouTube's automatic subtitles break text into very short blocks. Combine them for a smooth reading experience before publishing.

SCENARIO #3

Post-Production Editing

Before editing or translating, merge fragmented cues so you work with complete sentences instead of fragments.

SCENARIO #4

Accessibility & Reading Comfort

Rapid single-word subtitles are hard to read. Merging improves comprehension for all viewers.

Execution Pipeline

How It Works

Streamlined end-to-end processing pipeline engineered for speed and precision.

STEP 01

Upload & Import

Upload your subtitle file (especially effective on Whisper or YouTube output).

PHASE 01
Next
STEP 02

Set the merge thresholds

max characters per line, pause threshold, and max merged duration.

PHASE 02
Next
STEP 03

Configuration

The merger scans for adjacent cues with short gaps and combines them greedily.

PHASE 03
Next
STEP 04

Processing

Each merge only happens if the combined text fits the character limit and duration cap.

PHASE 04
Next
STEP 05

Verify & Export

Download the cleaned subtitle file with natural, readable cue blocks.

PHASE 05
Next
Engine Capabilities

Why Choose AllSubConverter?

Client-Side Privacy

Standard tools run in your browser so standard files are not uploaded to our servers.

Fast Local Parsing

Optimized processing engine parses and converts files directly in your browser.

Generous Batches

Up to 50 MB per file, 200 files or 1 GB total per batch.

Universal Compatibility

Runs in modern browsers across Windows, macOS, Linux, iOS, and Android.

Frequently Asked Questions

AI transcription tools like Whisper and YouTube auto-captions generate very short cues, often one or two words each, with tiny gaps between them. This causes rapid flashing on screen and poor readability.

No. Cues separated by pauses longer than your threshold are never merged. This preserves natural breaks where the speaker actually paused.

Input: SRT, VTT, ASS, SSA, LRC, SBV, TTML, SCC, STL, SMI. Output: SRT.

It merges adjacent cues only when three conditions are met: the gap between them is shorter than your pause threshold (default 400ms), the combined text fits your character limit (default 35), and the merged duration stays under the maximum (default 5 seconds).

Only minimally. The merged cue keeps the first cue's start time and takes the last merged cue's end time. The on-screen timing stays accurate to the original speech.

Yes, completely free with no sign-up or limits. It runs entirely in your browser.

Ready to convert your subtitles?

Combine fragmented subtitles from Whisper and YouTube into clean, readable cues. Stop the rapid-fire flashing of one-word subtitle blocks.