Accuracy falls with declining recording quality and rising speaking rate, and falls most sharply on the shortest segments, where the least contextual evidence is available and where the remaining recognition error is concentrated. Error rates vary with speaker gender and age, but no more than for any other system measured on the same segments, and show no geographic concentration within or across languages — indicating that this variation reflects the composition of the benchmark population rather than the behaviour of the model.
Since the errors are substitutions, we examined what is substituted. In every language measured, the model emits words that nothing in the segment supports at around half again the rate of the strongest comparison system, and the consistency across languages is the notable part. Most are ordinary words of the language placed where nothing supports them rather than invented strings, and the word displaced is more often a common one than a rare one.
Within Hindi the error rate is uneven across the regions its speakers come from. It is highest in the Haryanvi-speaking districts, where the distance from the strongest comparison system is the largest of any state and extends across the state boundary into the neighbouring districts of the National Capital Region. Female speakers are transcribed more accurately than male in the large majority of districts, an advantage that every system evaluated displays, and that persists within every quartile of recording quality and every band of segment duration.
The model's headroom on this benchmark therefore lies less in output-layer correction than in recognition of short, context-poor conversational speech, with written-language identification in Urdu the one substantial exception.