Tools Model Security Capability muse-spark
Model record
Muse Spark: how it does on security work, and what it will answer
- Lab
- Meta
- Access
- api
- Weights
- closed
- Safeguards
- Meta's control is contractual: the acceptable use policy and deployer-side Llama Guard, Prompt Guard and Code Shield
- Score rows
- 1
- Confidence
- CONFIRMED
What happened
Muse Spark scores 65.4% end-to-end on Cybench on all 40 tasks, the highest full-task-set card figure on the leaderboard read 11 September 2026.
Cybench, unguided, solved: 65.4% (vendor-card; vendor system card run; 40 of 40 tasks; safeguards: as run by the vendor; safeguard state per the card; run 2026-09-11; CONFIRMED). task count is what the vendor ran, so rows are comparable only with it beside them. Refusal gate (CONFIRMED, observed 2026-09-11): Acceptable use policy bars creating malicious code; no verification tier; deployer-side filters shipped as Purple Llama.
Sources
- Cybench leaderboard CSVraw.githubusercontent.com/cybench/cybench.github.io/main/dat…
- Refusal policy or licenceraw.githubusercontent.com/meta-llama/llama-models/main/model…
On YFarmX
- Reference pageyfarmx.com/ai/llms/muse-spark/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026