Models › Malay Coder v1.0 260923
Experimental
Oaica 35B-A3B Malay Coder v1.0 260923
An experimental specialist: a 35B sparse mixture-of-experts model fine-tuned from the open-weight base Qwen3.6-35B-A3B, aimed at code generation from Malay prompts.
Specification
The base, the size and the container are shared by every model in the family; only the fine-tune differs.
| Field | Value |
|---|---|
| Base model | Qwen3.6-35B-A3B (open-weight) |
| Parameters | about 35 billion in total, about 3 billion active per token |
| Architecture | sparse mixture-of-experts (MoE) |
| Version | v1.0, release code 260923 (23 September 2026) |
| Weights | shipped in the .oqm quantised container |
| Serving | the .oqm engine can pin the model’s MoE layers in host CPU RAM and fill the available GPU memory automatically, offloading the remainder |
| Minimum hardware | laptop- and desktop-class GPUs at reduced context, up to datacenter servers |
| Context | set by the serving configuration; runtime and context memory require workload-specific budgeting |
| Language | Malay, with the base model’s own multilingual coverage |
Intended use
| Candidate use in a supervised pilot | Evidence still needed |
|---|---|
| Code generation from Malay prompts | Broader coding tests and review of generated code |
Measured results
MalayMMLU uses 484 scored questions, greedy generation with reasoning off and answer-letter extraction. HumanEval uses 164 coding problems. Figures describe each test separately; percentages across different tests are not interchangeable. Counts are included to make the denominators explicit. No paired significance claim is made.
| Measurement | Value |
|---|---|
| MalayMMLU correct / 484 | 351 (72.5%) |
| HumanEval passed / 164 | 77 (47.0%) |
| Safety composite (Saved Run 3) | 45.91 (39.40–51.54) |
| LiveCodeBench v6 | 28.6% on 175 problems — a different benchmark from HumanEval, and one that does not itself contradict it |
| Small Malay instruction coding set | 16 / 16 a small, easy 16-item set — recorded as a feasibility observation, not broad validation |
The composite is our reproduction of the SEA-HELM Malay balanced-accuracy formula — toxicity plus three cultural subtasks — over 1,572 scored items, greedy generation with reasoning off at batch size 8. The interval describes conditional item-sampling uncertainty, retaining shared prompt/response structure; it excludes generation reruns, training and model selection. No safety ranking follows from overlapping marginal intervals; a paired comparison would be needed.
Official SEA-HELM Malay panel
Not run on this model. The official pinned-harness panel covers the general and safety variants only; neither experimental specialist has a published row here. What the panel is and how it is measured is described on the family page.
Official SEA-HELM Indonesian panel
Not run on this model. The Indonesian panel, like the Malay one, covers the general and safety variants only.
Limits
These are descriptive point estimates, not population guarantees. The evidence still needed above has not been produced, so this is not a deployment recommendation. Long-context evaluation remains incomplete across the family.
Where it sits in the family
See all four models side by side ›
Run it on hardware you own
Install OAICA, point it at a model you control, and run Claude Code and other agents on it. Local models are never gated and do not expire.
What this page does and does not establish
- Use measured values with their protocol, denominator, run and limitations.
- Describe local deployment as a configuration to validate.
- Describe moderation and drafting as supervised pilot candidates.
We do not claim safety superiority, regulatory approval, verified absence of contamination, a universal capability ceiling, or replacement of human reviewers. Full protocol notes, counts, run hashes and the score-recomputation code are published on the research site.
See also: the whole family · install and usage docs · pricing.