This Muse Spark 1.1 review examines Meta's new reasoning model as both a developer API and the engine behind a more action-oriented Meta AI. Released on July 9, 2026, the model adds a one-million-token context window, multimodal input, tool use, computer control, and multi-agent orchestration. On July 24, Meta also began rolling out consumer features that can work with email and calendar apps, prepare briefings, conduct research, and create slides.
Meta's own report shows real gains alongside limits on long-horizon autonomy. Independent testing finds an attractive price-to-performance profile but higher-than-average token use. The practical question is whether the model, product surface, permissions, and total task cost fit your work.
Muse Spark 1.1 Review: Quick Verdict
Muse Spark 1.1 is a strong value candidate for coding, research, multimodal analysis, and agent workflows that need a large context window. Independent results place its public-preview API among stronger reasoning models without making it the leader on every task.
Its biggest advantages are the one-million-token context window, fast output, tool and function calling, parallel subagent support, computer use, and integration into Meta AI. Its biggest cautions are uneven long-horizon performance, proprietary weights, a preview-stage developer platform, potential verbosity, and the risks created when an assistant can access personal apps or take actions.
Developers should pilot it on a measured workload. Consumers should start with low-risk tasks and review connected-app permissions before enabling recurring or consequential actions.
What Is Muse Spark 1.1?
Muse Spark 1.1 is a proprietary multimodal reasoning model from Meta Superintelligence Labs. It succeeds the original Muse Spark released in April and now powers Thinking mode in the Meta AI app and on meta.ai. Developers can access it through the new Meta Model API, which was in public preview at this review's July 27 update.
The model accepts text and images and produces text. The API exposes agent-oriented features such as developer prompts, tool and function calling, structured output, and parallel tool use.
This is a different strategy from Meta's earlier open-weight Llama releases. Muse Spark 1.1's model weights are not public, and Meta has not disclosed its parameter count. Teams should evaluate it as a hosted proprietary service rather than assume it can be downloaded or self-hosted.
Muse Spark 1.1 Features for AI Agents
Long Context and Context Management
Muse Spark 1.1 supports a one-million-token context window. Meta also says the model can retrieve information from earlier in a long task and compact its working context while preserving important steps. That is useful for large repositories, long research projects, and workflows that collect information across many tool calls.
A large context window is capacity, not perfect memory. Test retrieval after long tool traces and confirm that updated instructions replace stale notes.
Parallel Subagents and Tool Use
Meta describes Muse Spark 1.1 as able to plan work, delegate execution to parallel subagents, and bring results back to a main agent. It was also trained for common coding-agent patterns including planning mode, goal conditioning, delegation, and context compaction.
Parallel work can shorten tasks that split cleanly, but it can also multiply cost and review complexity. Record which subagents ran, what tools they used, and how conflicts were resolved.
Computer Use and Multimodal Work
The model can combine interface control with scripts. Meta demonstrates it debugging a web app through screenshots and creating a Marketplace listing from smartphone video. These are vendor demonstrations, not guarantees.
Computer use deserves narrow permissions. Keep human approval around purchases, publishing, destructive changes, external messages, and sensitive data transfers. SD's AI agent security guide offers a broader control checklist.
Muse Spark 1.1 Benchmarks and Real Limits
Meta's evaluation report shows substantial improvement over Muse Spark 1.0, especially in coding, agent tasks, calibration, and prompt-injection resistance. It also supplies an unusually useful limitation: on Terminal-Bench 2.1 and SWE-Bench Pro, Muse Spark 1.1 trails Claude Opus 4.8 and/or GPT-5.5. On long-horizon tests such as DeepSWE and DeepSearchQA, it is behind or roughly level with the best competitors rather than clearly ahead.
The report says the model solved 24 of 42 SWE-Bench Verified Hard tasks at least once. That is not the same as dependable first-try production performance.
Independent Artificial Analysis testing currently gives the highest-reasoning version an Intelligence Index score of 51, ranking it 18th among 187 comparable models on the live page. It measured about 121 output tokens per second and described the model as faster than average but somewhat verbose. The model generated 94 million output tokens across that benchmark suite, versus a 63 million median for its comparison class.
These results make Muse Spark 1.1 a serious test candidate, not a universal winner. Compare models with the same prompts, tools, acceptance tests, reasoning settings, and retry policy. SD's Claude Code vs Codex comparison explains why the agent harness matters as much as the model.
Muse Spark 1.1 Pricing and API Access
Artificial Analysis lists first-party Meta API pricing at $1.25 per million input tokens, $0.15 per million cached input tokens, and $4.25 per million output tokens. Those rates are competitive among current proprietary reasoning models, but they should be verified in the Meta Model API console before production use because the service is still in public preview and pricing can change.
Low token rates do not guarantee a low-cost agent. Reasoning traces, parallel subagents, tool calls, and retries can dominate the bill. Track cost per accepted task.
Record input, cached input, reasoning and output tokens; elapsed time; tool calls; retries; human correction; and final acceptance. The Kimi K3 review uses the same task-level approach to separate a rate card from workflow economics.
What the New Meta AI Agent Can Do
Meta's July 24 rollout turns Muse Spark 1.1 into a consumer agent. In select markets, Meta AI can plan tasks, connect with email and calendar apps, produce briefings, conduct research, create slides, and run recurring tasks. Users can steer work in progress.
These features were beginning to roll out in the Meta AI app and on meta.ai, with more countries and surfaces, including WhatsApp, planned for later. Availability may therefore differ by account and location.
Start with reversible tasks: summarize a calendar, prepare a sourced research brief, draft a plan, or assemble a mood board. Avoid automatic purchases, important messages, or work that depends on subtle private context.
Safety, Privacy, and Agent Controls
Meta's evaluation report says Muse Spark 1.1 reached or could not be ruled out from reaching “high risk” capability thresholds in chemical and biological work and cybersecurity before mitigations. Meta assessed the deployed residual risk as moderate or lower after layered safeguards. This is Meta's framework and conclusion, not an independent certification.
The report finds gains against prompt injection but notes weaker results in some scenarios, including file injection. Meta recommends strict tool allowlists, application safeguards, and workspace isolation. Connected agents can turn a bad instruction into an external action.
Meta says Incognito Chat is processed in a secure environment it cannot access and disappears by default. The feature was still rolling out and should not be confused with ordinary connected-app tasks. Check the active mode, shared data, retention, and action permissions.
Who Should Use Muse Spark 1.1?
Muse Spark 1.1 is worth testing for developers who need a lower-cost reasoning model with long context, multimodal input, and agent-friendly tool use. It may also suit teams seeking another provider for coding, research, document analysis, or computer-use experiments.
Consumers may find the planning, briefing, research, and slide features useful. Start with bounded tasks and inspect results before expanding access.
Wait or test cautiously if you require open weights, stable general availability, proven long-horizon reliability, strict data residency, or mature enterprise administration. Also compare end-to-end results with Claude, GPT, Gemini, and other models instead of choosing on price or one benchmark.
Conclusion
This Muse Spark 1.1 review finds a credible, competitively priced agent model with useful strengths in long context, multimodal input, tool use, parallel orchestration, coding, and computer control. The new Meta AI actions make it more relevant to everyday users, while the Model API gives developers a practical way to test it.
The case for caution is equally clear. The API is in preview, the model is proprietary, long-horizon results remain uneven, and agent access increases the cost of permission mistakes. Test Muse Spark 1.1 on representative tasks, measure total workflow cost, keep consequential actions behind approval, and treat benchmarks as evidence for a pilot rather than a reason to switch blindly.
Written by
Lena Ortiz
AI Tools Analyst
Lena tests AI products through the lens of creators, operators, and teams that need software to stay useful after launch week.
AI model reviews
Follow the models and agents changing how teams work.
Read more Syntax Dispatch coverage of AI models, coding agents, practical evaluations, and safe deployment workflows.
Browse AI toolsFAQ
When was Muse Spark 1.1 released?
Meta released Muse Spark 1.1 and the public-preview Meta Model API on July 9, 2026. New action-oriented Meta AI features began rolling out on July 24.
How much does Muse Spark 1.1 cost?
At publication, independent tracking listed Meta API pricing at $1.25 per million input tokens, $0.15 per million cached input tokens, and $4.25 per million output tokens. Check the provider console for current terms.
Is Muse Spark 1.1 open source?
No. Muse Spark 1.1 is proprietary, its model weights are not public, and Meta has not disclosed its parameter count.
Is Muse Spark 1.1 better than Claude or GPT?
Not across every task. Meta reports competitive results and major gains over Muse Spark 1.0, but its evaluation also shows Muse Spark 1.1 trailing Claude Opus 4.8 and/or GPT-5.5 on some coding tests.




