arXiv:2609.03502v118 หน้า
A Thai voice from fifteen seconds of speech
An 82M model small enough to run on the device, trained entirely on speech a larger model generated, and the benchmark that says what it still gets wrong.
Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut
Three of the twelve teacher voices. None was recorded.
