Cloud or Local Speech-to-Text? We Built the Same App Both Ways. Here's What We Learnt.

https://hackernoon.imgix.net/images/cZHlNDzCOTaC2f46Rngy5xZFnwI3-gfb3bhi.jpeg

One Electron app, two modes, one interface. What a two-day hackathon build taught us about on-device diarization, CoreML and DirectML, read against a production on-device engine.

At this year's internal hackathon, our esteemed colleagues Georgios Hadjiharalambous and Stuart Wood set out to build Inkwell, a dictation app centered around high-accuracy and speaker-aware note-taking, powered by Speechmatics speech-to-text.

Nothing too unusual there.

What makes Inkwell worth writing about is the decision they made early on: the app would run in two modes, cloud and fully local, behind one identical interface.

Flip a toggle, and the same Electron UI talks to either a realtime cloud endpoint or a speech recognition engine running entirely on the laptop in front of you, no network required after install.

That second mode is obviously the interesting one. "Local Whisper works" is a solved problem developers stopped being impressed by two years ago. What's still mostly undocumented...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more