Descript
Underlord Engine v4.2 Vendor: Descript Inc.Institutional AI audio and video workstation that translates waveform timelines into malleable text documents. Features real-time neural speech clean-up, automated multi-speaker diarization, and gaze vector reconstruction.
Word Error Rate (WER) archive-recorded on multi-accent conversational datasets.
Signal-to-Noise Ratio (SNR) enhancement with room echo cancellation.
Sub-second generative audio patch latency in timeline text buffer.
Editing speed multiplier over conventional timeline NLE interfaces.
Dual-Stage Studio Sound Reconstruction
Evaluation metrics demonstrating spectrogram cleansing before and after Descript Studio Sound dual-pass denoising on an untreated 85dB noise floor environment.
// Stream Diarization & Gaze Tracker [14:22:01.092] INFO: Ingest multi-track chunk: 48kHz PCM [14:22:01.104] DIAR: Speaker_01 ("Sarah") confidence 0.998 [14:22:01.121] DIAR: Speaker_02 ("Alex") confidence 0.992 [14:22:01.144] NLP: Filler words flagged: - "um" [00:04:12.100 -> 00:04:12.350] [REMOVED] - "you know" [00:04:15.820] [REMOVED] [14:22:01.189] GAZE: Angle theta -14.2 deg (off-axis) [14:22:01.210] GAZE: Tensor correction applied (0.00 deg) [14:22:01.245] EXPORT: Synced XML ready (NLE Bridge v1) [14:22:01.250] STATUS: Buffer verified deterministic
Multi-Modal Studio Pipeline
Studio Sound AI
Isolates speech with a single click, eliminating ambient mechanical hum, echo, and poor microphone acoustics.
AI Eye Contact
Corrects gaze trajectory seamlessly so presenters appear to look directly into the camera while reading teleprompters.
Filler Word Eraser
Identifies and excises "ums", "ahs", repeat words, and gap silences with smooth audio cross-fades and frame blends.
Multi-Cam Auto-Cut
Automatically switches camera perspectives to the active speaker based on speaker diarization without manually slicing clips.
Descript vs Industry Peers
| Evaluation Vector | Descript Underlord | Otter.ai | Adobe Premiere Pro | Synthesia |
|---|---|---|---|---|
| Primary Workload | Video & Podcast Creation | Meeting Notes & Summaries | Traditional NLE Production | Avatar Video Generation |
| Text-Based Waveform Editing | check_circle Native Native (Bidirectional) | Transcript Only (No Video Cut) | Text-Based Rough Cut (Basic) | Prompt / Script to Avatar |
| Neural Audio Clean-up | Studio Sound (Dual-Stage) | Standard AGC | Enhance Speech AI | Synthetic TTS Only |
| Generative Voice Patch (Overdub) | check_circle Yes (< 500ms patch) | No | No | Yes (Full Avatar Voice) |
| Automated Gaze Correction | check_circle Included (AI Eye Contact) | No | Third-party plug-in | Synthetic Face Rig |
| NLE Interoperability | FCPXML, Premiere XML, EDL | TXT, DOCX, SRT | Native Project Files | MP4 Video Render |
Transparent Pricing Index
Free
- check 1 hr transcription / month
- check 720p watermark-free export
- check Basic Studio Sound
- close AI Eye Contact not included
Hobbyist
- check 10 hrs transcription / mo
- check 4K video exports
- check Unlimited Studio Sound
- check Basic Underlord AI Actions
Pro
- check 30 hrs transcription / mo
- check Unlimited AI Eye Contact
- check Full Generative Voice Match
- check Filler word removal (all)
Enterprise
- check Unlimited hours pooled
- check Dedicated SSO & SOC2 Type II
- check Custom Voice Clone Governance
- check Tailored API SLA & Support
“Descript remains the gold standard for creator and podcast engineering, transforming complex multi-track timeline editing into as natural an action as editing a collaborative Google Doc.”
While professional cinema productions still demand traditional timeline precision inside DaVinci Resolve or Adobe Premiere, Descript drastically cuts turn-around time for interview, episodic, and corporate video pipelines by up to 80%.