Models
The Oaica 35B-A3B Malay family
Four models, built on one open-weight base and each measured on the same harness. Every number below carries its denominator, its run and its limits, because a model page that quotes a figure without those is quoting a different thing than the figure.
The four models
All four are fine-tuned from the open-weight base Qwen3.6-35B-A3B — about 35 billion parameters in total, about 3 billion active per token, in a sparse mixture-of-experts layout. The same weights are not tied to one accelerator class: the .oqm serving engine can pin the model’s mixture-of-experts layers in host CPU RAM and fill the available GPU memory automatically, so the family runs on laptop- and desktop-class GPUs at reduced context as well as on datacenter servers.
Oaica 35B-A3B Malay v1.0 260923
Malay drafting and assistance on local hardware.
MalayMMLU 392/484 · HumanEval 67/164 · safety 44.92
SafetyOaica 35B-A3B Malay Safety v1.0 260923
Malay screening with human review.
MalayMMLU 387/484 · HumanEval 52/164 · safety 47.31
ExperimentalOaica 35B-A3B Malay Coder v1.0 260923
Code generation from Malay prompts.
MalayMMLU 351/484 · HumanEval 77/164 · safety 45.91
ExperimentalOaica 35B-A3B Malay Researcher v1.0 260923
Malay instruction-following and drafting.
MalayMMLU 365/484 · HumanEval 129/164 · safety 45.24
Results, side by side
MalayMMLU is 484 scored questions, greedy generation with reasoning off and answer-letter extraction. HumanEval is 164 coding problems. Figures describe each test separately; percentages across different tests are not interchangeable. Counts are included to make the denominators explicit. No paired significance claim is made.
| Model | MalayMMLU correct / 484 | HumanEval passed / 164 | Safety composite |
|---|---|---|---|
| Malay v1.0 260923 | 392 (81.0%) | 67 (40.9%) | 44.92 (37.88–51.58) |
| Malay Safety v1.0 260923 | 387 (80.0%) | 52 (31.7%) | 47.31 (40.36–53.66) |
| Malay Coder v1.0 260923 (experimental) | 351 (72.5%) | 77 (47.0%) | 45.91 (39.40–51.54) |
| Malay Researcher v1.0 260923 (experimental) | 365 (75.4%) | 129 (78.7%) | 45.24 (38.60–51.24) |
| Stock Qwen3.6-35B-A3B (reference) | not interpretable under this protocol | 126 (76.8%) | 43.37 (36.92–48.95) |
The safety composite is our reproduction of the SEA-HELM Malay balanced-accuracy formula — toxicity plus three cultural subtasks — over 1,572 scored items, greedy generation with reasoning off at batch size 8. Each row is one saved run: Runs 1–4 are the four Oaica models, Run 5 the stock reference. The intervals describe conditional item-sampling uncertainty, retaining shared prompt/response structure; they exclude generation reruns, training and model selection. The 1,572 scored items are not 1,572 independent prompt clusters — the first two cultural subtasks share 71 prompts. The stock reference is not placed in that column because the protocol cannot score it: it does not write the option letter, answering by naming the choice’s text instead — a Benar/Palsu item is answered “Pernyataan tersebut adalah Benar.” — and scores 4 of 484, or 8 of 484 (1.7%) under the most tolerant extractor we have, against a 36.8% chance baseline. That sample is not uniformly four-option: 219 of its 484 items are two-option (Benar/Palsu), 231 four-option, 31 three-option and 3 five-option. All four of the stock’s extracted-correct answers come from two-option items; it scores zero on every three-, four- and five-option item. The extractor is not the problem: it reads a letter from all 484 of the general variant’s outputs. A stock figure produced under an instructed prompt, or with a longer generation budget, would be a different protocol and is not comparable to the four rows above. No safety ranking follows from overlapping marginal intervals; a paired comparison would be needed.
Official SEA-HELM Malay panel
A second, separate measurement: the same three arms scored on the pinned SEA-HELM harness (aisingapore/sea-helm at 9be3986d, v1.3.0) and aggregated with the harness’s own MultiRunAggregator — eight independent runs each, mean over runs with a 2,000-sample prompt bootstrap for the interval. These are leaderboard-comparable, and they are not on the same scale as the composite above; do not read the two tables as one series.
| Competency | Stock Qwen3.6-35B-A3B | Malay Safety v1.0 | Malay v1.0 |
|---|---|---|---|
| TOTAL | 69.43 (67.8–71.1) | 70.55 (68.6–72.5) | 69.84 (67.9–71.8) |
| Cultural | 56.67 | 76.90 | 76.41 |
| Safety | 37.63 | 45.47 | 43.88 |
| Multi-turn | 82.44 | 62.35 | 63.51 |
| Knowledge | 70.28 | 68.93 | 65.74 |
| Instruction-following | 81.92 | 80.70 | 80.88 |
Serving: llama.cpp llama-server on 1×A100-80GB, Q8_0 GGUF, -ngl 99 --jinja -c 65536 --parallel 8, fixed server-side sampling. Thinking off, the leaderboard-comparable no-think protocol. Judge for the multi-turn leg: openai/gpt-oss-120b, the pinned v1.3.0 default.
The TOTAL differences sit inside overlapping intervals — on this row the three arms are not separable at the benchmark level. What separates them is the per-competency trade: cultural and safety move clearly up for both fine-tuned arms, and multi-turn moves clearly down (82.4 against 62.4 and 63.5, well outside the intervals). The same direction appears in the Indonesian panel below. We report the trade rather than a total-score win.
Official SEA-HELM Indonesian panel
A third measurement, in a second language. The Indonesian SEA-HELM panel was run on the same pinned harness, the same eight-run protocol and the same serving configuration, with thinking off. No Indonesian text entered the fine-tuning mix, so this is a cross-language transfer probe, not a within-language result, and the base model was never tuned for the language. Figures here are the mean and sample standard deviation over the eight runs — a different statistic from the bootstrap intervals in the Malay table above, so the two are not one series.
| Competency | Stock Qwen3.6-35B-A3B | Malay Safety v1.0 | Malay v1.0 |
|---|---|---|---|
| TOTAL | 66.91 ± 6.83 | 70.00 ± 1.59 | 65.29 ± 3.59 |
| Cultural | 57.51 | 76.80 | 70.57 |
| Safety | 50.47 | 57.16 | 46.06 |
| Multi-turn | 84.95 | 63.33 | 61.66 |
| Knowledge | 67.92 | 72.50 | 66.96 |
| Instruction-following | 87.62 | 85.48 | 88.21 |
20 of the 21 id tasks: the gated syntax-criteria task could not be obtained, and none of the arms was scored on it. Same serving and judge as the Malay panel.
The three TOTALs span less than the run-to-run scatter of the stock arm, so this panel does not separate the arms overall; two of the stock runs collapse across the classification legs on the same seed positions and depress its raw mean, and excluding them leaves the ordering unchanged. The one signal that reproduces across both languages is multi-turn, where the stock reference scores about 85 against roughly 62 for both fine-tuned arms. On safety the order here reverses against Malay — the general arm sits below the stock reference. The fine-tune’s Malay safety gain does not carry into a language it was never tuned for. We report that rather than a total-score win.
How they are served
The weights ship in our own .oqm quantised container. Higher-throughput engine configurations are offered for API and cloud inference. Runtime and context memory still require workload-specific budgeting — the deployment is a configuration you validate against your own workload, not a spec sheet to trust.
Fully local operation requires external services and outbound integrations to be disabled.
Every model runs on hardware you own
Install OAICA, point it at a model you control, and run Claude Code and other agents on it. Local models are never gated and do not expire.
What we do and do not claim
- Use measured values with their protocol, denominator, run and limitations.
- Describe local deployment as a configuration to validate.
- Describe moderation and drafting as supervised pilot candidates.
We do not claim safety superiority, regulatory approval, verified absence of contamination, a universal capability ceiling, or replacement of human reviewers. Full protocol notes and the score-recomputation code are published on the research site.