AI research and products, built for Thai.

Paxa Labs is a product and research lab in Bangkok.

Thai users switch languages mid-sentence, write without spaces between words, depend on tone for meaning, and work with documents that mix scripts, tables, stamps, and handwriting.

We turn those everyday problems into speech, translation, and document products that teams can use in production.

01 / 03

The input is real Thai

People do not rewrite their lives into clean benchmark language before they use software.

ส่ง email หา HR หน่อย

เดี๋ยว join meeting แล้วส่ง file ให้

ใช้ AI summarize อันนี้ให้หน่อย

A customer can move from Thai to English, a product name, and a number in one breath. A useful model has to keep the context, register, and speaker intact.

Word boundaries depend on context.

ตากลม

Where one word ends and another begins can depend on context. Before a model can understand a Thai sentence, sometimes it first has to decide what the words even are.

Tone changes the word.

เสือ · เสื่อ · เสื้อ

Tiger. Mat. Shirt.

The consonants and vowels can look almost identical in romanization, but changing the tone changes the word.

The same complexity continues in business documents: Thai and English names, stacked marks, numbers, tables, stamps, and handwritten notes all compete for the same page.

These are not edge cases here. They are the input.

02 / 03

Each problem becomes a product

We build one focused product around each place a general-purpose pipeline breaks.

Thai speech where pronunciation, tone, rhythm, English words, numbers, abbreviations, and names are all part of the same problem.

AI TranslationAvailable now

Translation that understands context, register, terminology, loanwords, and how Thai people actually write.

Speech-to-TextResearch preview

Transcription for real Thai recordings, including sentences that change language halfway through, with word timestamps.

OCRAvailable now

Models built to read the documents Thai people and Thai businesses actually use.

Text-to-speech, speech-to-text, translation, and OCR are live. Each one started as a failure case in real Thai input that a general-purpose pipeline could not clear.

A product is the test: the research has to survive real input, real latency, and real users.

03 / 03

Research that ships

We start with failures from actual Thai speech, language, and documents. The roadmap begins from those failures.

A failure becomes a data example and an evaluation. Repeated patterns shape the model, training method, and serving stack. The improved system goes back into the product, where the next difficult input tells us what to investigate.

Language specialists, model researchers, and platform engineers work on the same loop. That keeps pronunciation, terminology, layouts, API behavior, and production reliability connected, with no handoff between disciplines.

The output of the lab is a better product.

Built in Bangkok. Tested on Thai reality.

Research earns its place when it makes the product more useful.

paxa-tts-flash-v1

The team

One lab in Bangkok, working in both of the languages it models.

The lab includes people who have done AI research inside a global technology company. It includes people who have worked on Typhoon, SCB 10X's open Thai language-model family, and people who have run AI products under very high traffic. The lab introduces itself by function. Every layer below is owned in house, from the phonetic front-end to the API you call.

Model research

Architectures and training methods, and the training runs that produce every model we ship.

Thai language and data

Where Thai breaks a general-purpose pipeline, written down as data, evaluation sets, and reading rules.

Platform and serving

AI products kept fast and stable under real production load.

The demo voices on this site are named after Thai desserts. The models are the public face of the lab.

Published research

The lab publishes its methods. Both technical reports are on arXiv, with the model weights, the benchmarks, and the evaluation code released beside them.

Every report and released artifact

Text-to-speech, speech-to-text, translation, and OCR are live APIs you can call today.

Read the API documentation