I feel like another motivation is to keep the neural net parameters, AKA the most valuable part of the speech recognition algorithm, off-device. It'd have to be huge otherwise, and could be easily duplicated then used as a starting place for a competing product.