Open weights
Open base models where possible, hosted inference where it makes sense, with an eye on cost per token in a peso market.
Language models finetuned for the way Filipinos actually write: Tagalog, Cebuano and Bisaya, and the Taglish that mixes them. Built on open models, tested by people from the region.
Frontier models handle English fine and Tagalog badly. Bisaya, Cebuano and code-switched Taglish fall off a cliff. BudotsGPT and FilipinoGPT are our LoRA finetunes of open models for the languages and the register people use every day, in news, in tourism and online.
Open base models where possible, hosted inference where it makes sense, with an eye on cost per token in a peso market.
Filipino-context adaptation: instruction tuning with LoRA on local corpora.
Generated and human-audited datasets covering Cebuano, Tagalog and Taglish register, plus domain sets for news and tourism.
Retrieval and summarisation layers feeding Balita.ph and internal newsroom tools.
Benchmarks that test for local idiom, not just BLEU, and human raters from the region the text is for. A model that is fluent in Manila and lost in Cebu has not passed.