Research

Thai has no shortage of text, speech, or documents. It has a shortage of labelled data. We build the supervision ourselves and publish what transfers.

arXiv:2609.03502v118 pages

A Thai voice from fifteen seconds of speech

An 82M model small enough to run on the device, trained entirely on speech a larger model generated, and the benchmark that says what it still gets wrong.

Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut

arXiv(opens in a new tab)PDF(opens in a new tab)

Three of the twelve teacher voices. None was recorded.

arXiv:2609.03595v120 pages

Thai OCR trained without a single real label

A 0.9B document model taught on 45,723 reconstructed pages, and the controlled experiments that say which part of a synthetic document actually transfers.

Kunat Pipatanakul

arXiv(opens in a new tab)PDF(opens in a new tab)

An annual-report page whose original text has been erased and replaced with rendered Thai in the same regions.
One page, reconstructed from a public source document.

Open artifacts

Models

Benchmarks and datasets

Code

Demos

Working on Thai language technology?

This work is a collaboration between Wayu Research and Paxa Labs, self-funded by Wayu Research, with the Typhoon team co-authoring the speech report and reviewing the document one.