Under development, ready in a few more days

Models › Malay v1.0 260923

General

Oaica 35B-A3B Malay v1.0 260923

The general-purpose member of the family: a 35B sparse mixture-of-experts model fine-tuned from the open-weight base Qwen3.6-35B-A3B, intended for Malay drafting and assistance on local hardware.

Specification

The base, the size and the container are shared by every model in the family; only the fine-tune differs.

FieldValue
Base modelQwen3.6-35B-A3B (open-weight)
Parametersabout 35 billion in total, about 3 billion active per token
Architecturesparse mixture-of-experts (MoE)
Versionv1.0, release code 260923 (23 September 2026)
Weightsshipped in the .oqm quantised container
Servingthe .oqm engine can pin the model’s MoE layers in host CPU RAM and fill the available GPU memory automatically, offloading the remainder
Minimum hardwarelaptop- and desktop-class GPUs at reduced context, up to datacenter servers
Contextset by the serving configuration; runtime and context memory require workload-specific budgeting
LanguageMalay, with the base model’s own multilingual coverage

Intended use

Candidate use in a supervised pilotEvidence still needed
Malay drafting and assistance on local hardwareDomain factuality, instruction-following and workload validation

Measured results

MalayMMLU uses 484 scored questions, greedy generation with reasoning off and answer-letter extraction. HumanEval uses 164 coding problems. Figures describe each test separately; percentages across different tests are not interchangeable. Counts are included to make the denominators explicit. No paired significance claim is made.

MeasurementValue
MalayMMLU correct / 484392 (81.0%)
HumanEval passed / 16467 (40.9%)
Safety composite (Saved Run 1)44.92 (37.88–51.58)
Historical 500-question MalayMMLU82.8 a separate protocol — historical 500-question sample, reasoning on — and not comparable to the 484-question figure above

The composite is our reproduction of the SEA-HELM Malay balanced-accuracy formula — toxicity plus three cultural subtasks — over 1,572 scored items, greedy generation with reasoning off at batch size 8. The interval describes conditional item-sampling uncertainty, retaining shared prompt/response structure; it excludes generation reruns, training and model selection. No safety ranking follows from overlapping marginal intervals; a paired comparison would be needed.

Official SEA-HELM Malay panel

A second, separate measurement: this model and the stock reference scored on the pinned SEA-HELM harness (aisingapore/sea-helm at 9be3986d, v1.3.0), aggregated with the harness’s own MultiRunAggregator — eight independent runs each, mean over runs with a 2,000-sample prompt bootstrap for the interval. Serving: llama.cpp llama-server on 1×A100-80GB, Q8_0 GGUF, thinking off (the leaderboard-comparable no-think protocol); the multi-turn leg is judged by openai/gpt-oss-120b. These numbers are leaderboard-comparable and are not on the same scale as the composite above.
CompetencyThis modelStock Qwen3.6-35B-A3B
TOTAL69.84 (67.9–71.8)69.43 (67.8–71.1)
Cultural76.4156.67
Safety43.8837.63
Multi-turn63.5182.44
Knowledge65.7470.28
Instruction-following80.8881.92

The TOTAL difference against the stock reference sits inside overlapping intervals — on this row the arms are not separable at the benchmark level. What separates them is the per-competency trade: cultural and safety move clearly up for this fine-tuned arm, and multi-turn moves clearly down, well outside the intervals. The same direction appears in the Indonesian panel. We report the trade rather than a total-score win.

Official SEA-HELM Indonesian panel

A third measurement, in a second language: the Indonesian SEA-HELM panel, 20 of the 21 id tasks (the gated syntax-criteria task could not be obtained for any arm), on the same pinned harness and the same eight-run protocol, Q8_0 under llama.cpp with thinking off. No Indonesian text entered the fine-tuning mix, so this is a cross-language transfer probe, not a within-language result. Figures are the mean and sample standard deviation over eight runs; the Malay row above uses bootstrap intervals and the two are not the same statistic.
CompetencyThis modelStock Qwen3.6-35B-A3B
TOTAL65.29 ± 3.5966.91 ± 6.83
Cultural70.5757.51
Safety46.0650.47
Multi-turn61.6684.95
Knowledge66.9667.92
Instruction-following88.2187.62

On Indonesian this arm is the weakest of the three and the only one below the stock reference on the safety composite — the Malay safety gain does not carry over to a language the model was not tuned for. The multi-turn gap against the stock reference (about 62 against 85) is the same direction as on the Malay panel. Aggregated with the harness’s own MultiRunAggregator the same run gives TOTAL 64.91 (63.50–66.31) and safety 46.09 (41.05–50.79).

Limits

These are descriptive point estimates, not population guarantees. The evidence still needed above has not been produced, so this is not a deployment recommendation. Long-context evaluation remains incomplete across the family.

Where it sits in the family

See all four models side by side ›

Run it on hardware you own

Install OAICA, point it at a model you control, and run Claude Code and other agents on it. Local models are never gated and do not expire.

Install OAICA

What this page does and does not establish

  • Use measured values with their protocol, denominator, run and limitations.
  • Describe local deployment as a configuration to validate.
  • Describe moderation and drafting as supervised pilot candidates.

We do not claim safety superiority, regulatory approval, verified absence of contamination, a universal capability ceiling, or replacement of human reviewers. Full protocol notes, counts, run hashes and the score-recomputation code are published on the research site.

See also: the whole family · install and usage docs · pricing.