YFarmX logoYFarmX

Tools Model Security Capability muse-spark

Model record

Muse Spark: how it does on security work, and what it will answer

Lab
Meta
Access
api
Weights
closed
Safeguards
Meta's control is contractual: the acceptable use policy and deployer-side Llama Guard, Prompt Guard and Code Shield
Score rows
1
Confidence
CONFIRMED

What happened

Muse Spark scores 65.4% end-to-end on Cybench on all 40 tasks, the highest full-task-set card figure on the leaderboard read 11 September 2026.

Cybench, unguided, solved: 65.4% (vendor-card; vendor system card run; 40 of 40 tasks; safeguards: as run by the vendor; safeguard state per the card; run 2026-09-11; CONFIRMED). task count is what the vendor ran, so rows are comparable only with it beside them. Refusal gate (CONFIRMED, observed 2026-09-11): Acceptable use policy bars creating malicious code; no verification tier; deployer-side filters shipped as Purple Llama.

Sources

One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026