Under development, ready in a few more days

Models › Malay Safety v1.0 260923

Safety

Oaica 35B-A3B Malay Safety v1.0 260923

The safety-tuned member of the family: a 35B sparse mixture-of-experts model fine-tuned from the open-weight base Qwen3.6-35B-A3B, intended as a Malay screening candidate with human review in the loop.

Specification

The base, the size and the container are shared by every model in the family; only the fine-tune differs.

FieldValue
Base modelQwen3.6-35B-A3B (open-weight)
Parametersabout 35 billion in total, about 3 billion active per token
Architecturesparse mixture-of-experts (MoE)
Versionv1.0, release code 260923 (23 September 2026)
Weightsshipped in the .oqm quantised container
Servingthe .oqm engine can pin the model’s MoE layers in host CPU RAM and fill the available GPU memory automatically, offloading the remainder
Minimum hardwarelaptop- and desktop-class GPUs at reduced context, up to datacenter servers
Contextset by the serving configuration; runtime and context memory require workload-specific budgeting
LanguageMalay, with the base model’s own multilingual coverage

Intended use

Candidate use in a supervised pilotEvidence still needed
Malay screening with human reviewHarmful-content misses, false alarms, subgroup and adversarial tests

Measured results

MalayMMLU uses 484 scored questions, greedy generation with reasoning off and answer-letter extraction. HumanEval uses 164 coding problems. Figures describe each test separately; percentages across different tests are not interchangeable. Counts are included to make the denominators explicit. No paired significance claim is made.

MeasurementValue
MalayMMLU correct / 484387 (80.0%)
HumanEval passed / 16452 (31.7%)
Safety composite (Saved Run 2)47.31 (40.36–53.66)
Historical 500-question MalayMMLU83.4 a separate protocol — historical 500-question sample, reasoning on — and not comparable to the 484-question figure above

The composite is our reproduction of the SEA-HELM Malay balanced-accuracy formula — toxicity plus three cultural subtasks — over 1,572 scored items, greedy generation with reasoning off at batch size 8. The interval describes conditional item-sampling uncertainty, retaining shared prompt/response structure; it excludes generation reruns, training and model selection. No safety ranking follows from overlapping marginal intervals; a paired comparison would be needed.

Official SEA-HELM Malay panel

A second, separate measurement: this model and the stock reference scored on the pinned SEA-HELM harness (aisingapore/sea-helm at 9be3986d, v1.3.0), aggregated with the harness’s own MultiRunAggregator — eight independent runs each, mean over runs with a 2,000-sample prompt bootstrap for the interval. Serving: llama.cpp llama-server on 1×A100-80GB, Q8_0 GGUF, thinking off (the leaderboard-comparable no-think protocol); the multi-turn leg is judged by openai/gpt-oss-120b. These numbers are leaderboard-comparable and are not on the same scale as the composite above.
CompetencyThis modelStock Qwen3.6-35B-A3B
TOTAL70.55 (68.6–72.5)69.43 (67.8–71.1)
Cultural76.9056.67
Safety45.4737.63
Multi-turn62.3582.44
Knowledge68.9370.28
Instruction-following80.7081.92

The TOTAL difference against the stock reference sits inside overlapping intervals — on this row the arms are not separable at the benchmark level. What separates them is the per-competency trade: cultural and safety move clearly up for this fine-tuned arm, and multi-turn moves clearly down, well outside the intervals. The same direction appears in the Indonesian panel. We report the trade rather than a total-score win.

Official SEA-HELM Indonesian panel

A third measurement, in a second language: the Indonesian SEA-HELM panel, 20 of the 21 id tasks (the gated syntax-criteria task could not be obtained for any arm), on the same pinned harness and the same eight-run protocol, Q8_0 under llama.cpp with thinking off. No Indonesian text entered the fine-tuning mix, so this is a cross-language transfer probe, not a within-language result. Figures are the mean and sample standard deviation over eight runs; the Malay row above uses bootstrap intervals and the two are not the same statistic.
CompetencyThis modelStock Qwen3.6-35B-A3B
TOTAL70.00 ± 1.5966.91 ± 6.83
Cultural76.8057.51
Safety57.1650.47
Multi-turn63.3384.95
Knowledge72.5067.92
Instruction-following85.4887.62

The strongest of the three arms on the Indonesian panel overall and on its safety composite, and the most stable run to run. The gap to the stock reference on safety (57.16 against 50.47) is nonetheless inside the stock arm’s own run-to-run scatter, so it is not separable at the benchmark level; the same multi-turn gap as elsewhere is present.

Limits

These are descriptive point estimates, not population guarantees. The evidence still needed above has not been produced, so this is not a deployment recommendation. Long-context evaluation remains incomplete across the family.

Where it sits in the family

See all four models side by side ›

Run it on hardware you own

Install OAICA, point it at a model you control, and run Claude Code and other agents on it. Local models are never gated and do not expire.

Install OAICA

What this page does and does not establish

  • Use measured values with their protocol, denominator, run and limitations.
  • Describe local deployment as a configuration to validate.
  • Describe moderation and drafting as supervised pilot candidates.

We do not claim safety superiority, regulatory approval, verified absence of contamination, a universal capability ceiling, or replacement of human reviewers. Full protocol notes, counts, run hashes and the score-recomputation code are published on the research site.

See also: the whole family · install and usage docs · pricing.