Measuring the Measure: Evidence and Limits
Text-free outcome counts, recomputation code and provenance limits for our Malay LLM safety case study, stating what the published evidence can and cannot establish.
Measurement-first work on AI safety and Malay language models. Oaica is a safety-focused AI company.
Text-free outcome counts, recomputation code and provenance limits for our Malay LLM safety case study, stating what the published evidence can and cannot establish.
One changed prediction can move a Malay LLM safety subtask by up to 9 points. A study of the SEA-HELM Malay safety track, an Indonesian cross-language probe, and a checklist for any language.
The Oaica 35B-A3B Malay model, its safety-tuned variant and two experimental specialists: what we measured, the trade we report, and what the copy may say.
Oaica's 35B-A3B Malay model runs on your own hardware, from small GPUs to datacenter servers, with our .oqm format and a statement of what we do not claim.
How far a Malay AI safety score can be trusted. One changed prediction can move a Malay LLM safety subtask by up to 9 points. What the SEA-HELM Malay safety benchmark can and cannot resolve, with an Indonesian cross-language probe, by Oaica.
Oaica, a safety-focused AI company: one changed answer can move a Malay AI safety sub-score by up to 9 points. On-premises Malay AI and safety audits.