Built a React interface for file-based transcription, playback synchronized to transcript segments, and export workflows.
AI/ML / PROJECT OVERVIEW
Audio & Video Transcription
An audio and video transcription workspace with speaker diarization, time-aligned playback, searchable transcripts, and multiple export formats. A React interface connects to Flask services running local speech and audio models.
01 / INSIDE THE PRODUCT
What it brings together.
Independent project
Integrated Flask services with Whisper, WhisperX, and pyannote.audio for local transcription and speaker diarization.
Built with the right tools.
ReactViteFlaskWhisperWhisperXPyTorchpyannote.audioWaveSurfer.js
Have a similar
challenge in mind?
Tell us what you need to build, connect, or improve.
Discuss this project ↗