You’re only partially correct about input speed. If you want to dictate an email then yes you need to think about each word you want to say and the order in which to say them. Coupled with an LLM that problem is diminished because you can just kind of have a conversation with the LLM and tell it to draft an email.
and how much of that conversation with an LLM is "No, what I want is..." because it assumed something; or just straight up hallucinated or the typo made it go off on a tangent?
As for whisper, I can find sources that are saying for American-English speakers in a not-noisy environment (aka the best case scenario,) the model has a word error rate between 2-8%. For reference, Dragon NaturallySpeaking had a WER of 3-5%. So I wouldn't say that Whisper has made any substantial improvements, and they're OpenAi. you can trust them if you want. I don't think that'll work out well in the long run, though.
which is funny because that's most established content on YT, whether you realize it or not.