openPR Logo
Press release

Muse Spark 1.2 Benchmarks: Five Scoreboards, Five Different Answers

08-13-2026 07:41 AM CET | Business, Economy, Finances, Banking & Insurance

Press release from: publiera

/ PR Agency: Shakeel Ahmad
Muse Spark 1.2 Benchmarks: Five Scoreboards, Five Different

The Muse Spark 1.2 benchmarks https://www.orcarouter.ai/models/meta/muse-spark-1.2 that get quoted in headlines almost always come from one source - Meta's own deck - and that is the least independent of the five scoreboards now measuring this model. The other four disagree with it and with each other, which is the most useful thing to understand before trusting any number about it. This piece lays all five side by side; if you want the release context first, we covered it in our benchmark breakdown https://www.orcarouter.ai/blog/meta-muse-spark-1-2-explained

To see how wide the disagreement is: ask five people how good this model is and you'll get five defensible answers. It is second-best in the world at terminal coding tasks. It is twelfth of 185 models overall. It is fifth of 45 on an independent multi-domain index. It is fourteenth of 50 at coding under a neutral harness. It is not on the official coding leaderboard at all. Every one of those statements is currently true.

Scoreboard 1 - Meta's own deck

Meta reports three figures for Muse Spark 1.2:

• Terminal-Bench 2.1 - 82.9%, against a claimed 76.2% for the previous version

• DeepSWE v1.1 - 59.3%, up from 53.0%

• Meta internal coding bench - 70.6%, up from 68.3%

Methodology matters here more than the numbers. These are pass@1 over five attempts, run in Meta's harness, with each competing model paired to its own agent product. Meta's methodology note concedes that its setup may not be tuned for third-party systems. Nothing has been independently reproduced.

To Meta's credit, its chart shows Claude Opus 5 ahead at 86.7% - vendors don't usually publish the graph where they lose. But a vendor-run comparison where every model uses a different scaffold is a product comparison, not a model comparison, and it should be read that way.

Scoreboard 2 - Artificial Analysis

AA runs a fixed composite of reasoning, knowledge, maths and coding tasks across every model it tracks, which makes it the closest thing to an apples-to-apples general index.

Its live model page currently scores Muse Spark 1.2 (xhigh) at 57, ranking it #12 of 185. Two footnotes are important:

AA has revised this figure upward. Its launch-day write-up put the model at 54, which is the number nearly every article published in the last week still quotes. If you see 54, you're reading a snapshot from August 5.

The cost data is as interesting as the score. Running the index cost $639.27 and consumed 95 million output tokens, against a tier median near 70 million - the model burns a lot of thinking. On a per-task basis that's $0.40, up from $0.29 for version 1.1 at identical list pricing.

AA's agentic sub-results at launch showed the clearest generational movement: GDPval-AA v2 rose to 1631 Elo (#5) from 1371, Terminal-Bench v2.1 went from 78% to 80%, and τ3-Banking from 25% to 27%. There's also a revealing trade-off in its hallucination data - the model abstains more often now, cutting its hallucination rate from 38% to 28% but lowering raw accuracy from 41% to 38%. It became more cautious, not just smarter.

Scoreboard 3 - Vals AI, the common-harness test

Vals AI's distinguishing feature is that every model runs through the same harness. That removes exactly the variable that makes Meta's deck hard to interpret.

Under those conditions Muse Spark 1.2 lands at 71.88% ± 1.12 on the Vals Index - 5th of 45 models - at $0.69 per test, the lowest per-test cost of anything in the top five. That is a genuinely strong result, and it is the single most credible "this model is good" data point available.

Scoreboard 4 - Vals AI again, this time by domain

Here's where it gets strange. Vals publishes per-benchmark ranks, and Meta's *coding* model looks like this:

• Finance Agent (v2) - #1 of 44
• TaxEval v2 - #1 of 136
• Harvey's Legal Agent Benchmark - #1 of 31
• MedScribe - #2 of 80
• CorpFin v2 - #5 of 131
• Vals Multimodal Index - #6 of 32
• SWE-bench - #9 of 79
• Terminal-Bench 2.1 - #14 of 50
• MMLU Pro - #15 of 129

Read that list twice. On the one independent multi-domain harness available, a model Meta marketed almost exclusively on coding ranks first in finance, tax and legal agent work - and fourteenth on the benchmark its launch announcement led with.

That doesn't mean it's bad at coding; ninth of 79 on SWE-bench is a solid result. It means the positioning and the independent evidence point in different directions, and that the model's real strength may be long-horizon, tool-using, document-heavy professional work rather than terminal coding specifically. If you were only ever going to check one thing before adopting it, check it against your own domain rather than against the category on the box.

Scoreboard 5 - the official Terminal-Bench leaderboard

The board Meta's headline number references does not list Muse Spark 1.2. It does not list Muse Code either.

What it does list, at rank 8, is mini-SWE-agent + Muse Spark 1.1 - 76.2% ± 1.2%, xhigh, submitted by Princeton on July 9, 2026, at a run cost of $198.05. Above it: Claude Code + Claude Fable 5 at 83.8% ($552.67), Codex + GPT-5.5 at 83.1% ($2,059.19), Terminus 2 + Claude Fable 5 at 80.4%.

Now do the arithmetic on Meta's claim. 82.9 minus the claimed 6.7-point generational gain is 76.2 - precisely the board's figure for version 1.1 under Princeton's deliberately minimal scaffold. Meta never published the harness it used for its baseline, so this is suggestive rather than proven. But if the baseline really is that row, then the "+6.7 generational improvement" is measuring a bare third-party scaffold against Meta's own co-trained agent: a model upgrade and a harness upgrade fused into a single number.

There's a broader lesson here. Three different scales exist under one benchmark name: the official board tops out near 83.8%, Meta's deck shows 86.7%, and Artificial Analysis' own Terminal-Bench v2.1 implementation runs near 89.5%. A Terminal-Bench result is a *harness + model + effort* score. It is never a model score, and comparing across sources is meaningless.

What to actually conclude

Strip out the disagreements and a consistent picture survives:

• It is a strong upper-mid-tier model, not a frontier leader. Every independent source agrees on that band.

• It is unusually cost-effective per completed task - Vals' cheapest top-five model, and reasonable on AA's cost curve despite heavy token use.

• It made real gains on agentic, long-horizon work, which is where three separate sources moved in the same direction.

• Its coding leadership is vendor-asserted and independently unconfirmed. Fourteenth of 50 under a neutral harness is the number to plan around.

• It thinks a lot, which costs tokens and time.

The practical move is to stop treating public benchmarks as a purchasing decision and start treating them as a shortlist filter. Once a model is on your shortlist, the only benchmark that settles anything is your own task set. Running that comparison is easier when the candidates sit behind one API - OrcaRouter carries Muse Spark 1.2 among 200-plus models at 0% markup, so an A/B against your current model is a config change rather than a procurement cycle.

The takeaway

Muse Spark 1.2 is good, cheaper per task than most things near it, and better at long-horizon professional work than its coding-first marketing implies. But the five available scoreboards don't agree, and the one thing you should take from that is scepticism about any article - including a vendor's - that quotes a single figure as the answer. Check which harness produced the number, whether anyone outside the lab reproduced it, and whether the benchmark resembles your job. Then run your own.

Sourcing note: the 82.9% / 59.3% / 70.6% figures are Meta's own vendor-run results, unreproduced by any third party. Intelligence Index, cost and token figures are from Artificial Analysis; index and per-domain ranks from Vals AI; the 76.2% row from the official Terminal-Bench 2.1 leaderboard. The 82.9 - 6.7 = 76.2 observation is our inference, not a Meta statement. All checked August 7, 2026.

This release was published on openPR.

Permanent link to this press release:

Copy
Please set a link in the press area of your homepage to this press release on openPR. openPR disclaims liability for any content contained in this release.

You can edit or delete your press release Muse Spark 1.2 Benchmarks: Five Scoreboards, Five Different Answers here

News-ID: 4602660 • Views:

More Releases from publiera

Francophone IPTV Market Is Bigger Than Anyone Realizes
Francophone IPTV Market Is Bigger Than Anyone Realizes
Cable TV Is Collapsing. IPTV Is The Default Now. But Growth Comes With Problems. PARIS, August 2026 France: EUR 2.1B annual IPTV revenue. Quebec: EUR 890M. Belgium: EUR 480M. Switzerland: EUR 320M. Add Francophone Africa and you're at EUR 4.2B today. Projected EUR 5.2B by 2028 if current trends hold. This is real money. This is a real market. Yet the companies owning this market-Orange, Free, Videotron-are barely talking about growth. They're managing decline. The Cable
IPTV on Android TV: Installation Guide and Best Apps Compared
IPTV on Android TV: Installation Guide and Best Apps Compared
Android TV has emerged as serious alternative to Amazon Fire TV for IPTV streaming. Native Android operating system, Google Play Store access, and hardware flexibility create ecosystem uniquely suited for streaming enthusiasts. Whether you're considering Android TV box or already own one, this guide covers installation, app selection, optimization, and troubleshooting. Android TV's advantage is architectural. Pure Android means maximum app compatibility, flexibility in playlist formats, and zero restrictions from proprietary
Adedotun Olaoluwa Invites the World to Abuja as IGA Capital Hunts for $100 Million-Plus Opportunities
Adedotun Olaoluwa Invites the World to Abuja as IGA Capital Hunts for $100 Milli …
Adedotun Olaoluwa is inviting the world to Abuja - and, in doing so, making a larger argument about what Africa needs from an investment gathering now. The founder of Dotmount Communications is positioning the Africa Business and Investment Expo not as another conference built around speeches and ceremony, but as a marketplace for access: a place where investors, ministers, governors, entrepreneurs and project sponsors can meet around the harder question of
Godex Introduces Operational Readiness Program To Support Future Service Development
Godex Introduces Operational Readiness Program To Support Future Service Develop …
Summary: Godex has introduced a new operational readiness program focused on strengthening internal processes, technical coordination, and service continuity as the platform prepares for its next stage of development. Victoria, Bahamas, August 12, 2026 - Godex today announced the launch of an operational readiness program designed to support the next stage of its platform development. The initiative shifts the company's current focus from completing individual infrastructure milestones toward improving how technical

All 5 Releases


More Releases for Meta

AI Glasses Market Set to Boom Rapidly, Witnessing Strong Growth Through 2033 | M …
Coherent Market Insights has released a comprehensive study titled "AI Glasses Market: Industry Trends, Share, Size, Growth, Opportunity, and Forecast 2026-2033." The report provides validated market size estimates, percentage share analysis, competitive landscape evaluation, and detailed regional insights in an increasingly competitive and rapidly evolving market environment. It further examines industry performance indicators, growth drivers, restraints, cost structures, and investment feasibility metrics including projected returns and margin outlook. The study delivers
AI Glasses Market Set for Rapid Expansion Through 2033 | Meta (RayBan Meta), Goo …
Coherent Insights Reports has released a detailed research analysis on the Global "AI Glasses Market" 2026, highlighting key trends, growth dynamics, and forecast insights through 2033. This comprehensive report presents an in-depth evaluation of the landscape, analyzing the factors that influence industry growth, including manufacturers, suppliers, participants, and end users. It offers valuable insights into the core drivers fueling expansion across various segments such as product type, application, end-user, and
AI Glasses Market to See Thriving Worldwide| Meta (Ray‐Ban Meta), Google (Proj …
The AI Glasses market is estimated to be valued at USD 857.4 Mn in 2025 and is expected to reach USD 2,308.6 Mn by 2032, growing at a compound annual growth rate CAGR of 15.2% from 2025 to 2032. ➤ Coherent Market Insights has released a new research report to its portfolio of Market research titled Global AI Glasses Market: By Size, Trends, Share, Growth, Segments, Industry Analysis, and Forecast, 2032.
VR Wave Launches Premium Meta Oakley Lens for Oakley Meta Smart Glasses Users
United States, 7th Oct 2025 - VR Wave, a leader in precision-crafted lens technology, has announced the launch of its premium Meta Oakley Lens, designed to enhance the performance, comfort, and style of the Oakley Meta HSTN smart glasses. Featuring advanced photochromic technology, these lenses automatically adapt to changing light conditions, providing optimal vision whether indoors or outdoors. With the rise of Meta Oakley smart glasses and Meta Oakley AI glasses
VR Meta Universe Market Climbs on Positive Outlook of Booming Sales| Microsoft, …
Advance Market Analytics published a new research publication on "VR Meta Universe Market Insights, to 2030" with 232 pages and enriched with self-explained Tables and charts in presentable format. In the Study you will find new evolving Trends, Drivers, Restraints, Opportunities generated by targeting market associated stakeholders. The growth of the VR Meta Universe market was mainly driven by the increasing R&D spending across the world. Get Free Exclusive PDF Sample
VR Meta Universe Market May See a Big Move | Microsoft, Meta
Advance Market Analytics published a new research publication on "VR Meta Universe Market Insights, to 2028" with 232 pages and enriched with self-explained Tables and charts in presentable format. In the Study you will find new evolving Trends, Drivers, Restraints, Opportunities generated by targeting market associated stakeholders. The growth of the VR Meta Universe market was mainly driven by the increasing R&D spending across the world. Get inside Scoop of the