That’s not going to produce a good model. Test panels and dogfooding are both small scale and introduce biases based on who uses them. Until the gains stop coming primarily from supervised deep learning we won’t move away from the winner being whomever has the largest, most diverse dataset and compute resources.