Fastest Offline Voice-to-Text — Now with Multi-Speaker IDs
By Prompt Engineering
Key Concepts
- On-device Transcription: Voice-to-text processing that occurs entirely on the user's local machine, without sending data to external servers.
- Speaker Diarization: The process of identifying and separating different speakers within an audio recording.
- Large Language Model (LLM): An AI model used for natural language processing tasks, such as grammar correction and style enhancement.
- Real-time Transcription: The ability to convert spoken words into text as they are being spoken, with minimal delay.
- Local GPT: A previously developed on-device, private retrieval-augmented generation system.
- WBY: A voice assistant that utilizes local speech-to-text, LLM, and text-to-speech components.
- Retrieval Augmented Generation (RAG): A technique that combines information retrieval with text generation to produce more accurate and contextually relevant responses.
New Features and Enhancements
The world's fastest voice-to-text system for Mac OS has received significant updates, focusing on enhanced functionality and user privacy.
1. Local File Transcription
- Functionality: Users can now transcribe audio from local files directly within the application.
- Process: Select one or more files, initiate the transcription process, and receive instant results.
2. Speaker Diarization
- Core Feature: The system can now identify and differentiate multiple speakers within an audio recording, transcribing each speaker's contribution individually.
- Technical Process: This involves running multiple models, including clustering and speaker identification algorithms.
- Performance: While the diarization process itself takes longer than standard transcription due to its complexity, it is an offline process.
- Example: A podcast with three distinct speakers was demonstrated. The system identified four speakers, with the discrepancy attributed to background music in an earlier segment, highlighting the model's robustness even in challenging audio conditions.
- Customization: Users can modify speaker names (e.g., renaming identified speakers to "John," "Jack," "Matt") and export transcriptions with timestamps for each speaker.
- Application: This feature is particularly useful for building retrieval systems on top of transcriptions.
3. Enhanced Mode with LLM Integration
- Functionality: An "enhanced mode" leverages a Large Language Model (LLM) to correct grammatical errors and improve the style of transcriptions.
- User Control: Users can define how the LLM will handle grammatical corrections and stylistic adjustments.
- Performance: Despite LLM integration, the transcription remains remarkably fast.
- Intelligent LLM Loading: To address user feedback regarding memory usage, the LLM was previously offloaded when the app was inactive, causing a slight delay on restart. The new system intelligently loads the LLM during transcription, ensuring near-instantaneous performance even after periods of inactivity.
4. Improved Transcription Control
- Cancel Transcription: Users can now cancel an ongoing transcription by pressing the "escape" key.
- Options upon Escape:
- Cancel Transcription: Discards all transcribed text.
- Keep Going: Continues the transcription process without losing any data.
- Benefit: This provides a more user-friendly way to manage and correct accidental starts or unwanted transcriptions, preventing the need to manually stop and delete text.
Developer's Vision and Background
The developer emphasizes the importance of building local AI solutions for privacy and user control.
- Motivation: The project is driven by a belief that voice will be the primary interface of the future, reducing the need for typing.
- Previous Projects:
- Local GPT: A highly successful (22,000+ GitHub stars) on-device, private retrieval-augmented generation system.
- WBY: A voice assistant that integrates local speech-to-text, LLM, and text-to-speech for fully voice-driven interaction.
- Core Philosophy: The current voice-to-text system is presented as one of the fastest and most customizable on the market, running entirely locally.
Pricing and Availability
- Offer: A $10 discount is available, providing 12 months of updates without a subscription.
- Trial: A 3-day free trial allows users to test all features.
- Platform: The application is specifically for Mac OS.
- Update Schedule: The discussed updates are expected to be released by the end of the week (Friday or Saturday).
Conclusion and Call to Action
The developer is actively seeking user feedback to further improve the application. They encourage Mac OS users to test the system and provide their input. The core takeaway is the availability of a powerful, private, and fast on-device voice-to-text solution with advanced features like speaker diarization and LLM-powered enhancements.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

4 easy tips for writing better AI prompts | Kunalsinh Kathia | TEDxSaffrony Institute of Technology
TEDx Talks

URGENT Warning For Silver Holders: What Happens Next
GoldCore TV

'We got to get this dealt with in the next couple of weeks': Pelletier on Iran war and oil shortages
BNN Bloomberg

WALL STREET WATCH: SpaceX eyes HISTORIC Nasdaq debut
Fox Business Clips

Anthony O'Neal: From homeless at 19 to financially free
Yahoo Finance

Every Bond Market on Earth Is Breaking at Once. This Is 2008 x10
Peter Schiff

Officers who clashed with rioters on Jan. 6, 2021, sue to block "anti-weaponization fund” #shorts
CBS News