Muse Spark 1.1 and the Genius Behind Meta
Deep Dive: Muse Spark 1.1, Meta's Neocloud Pursuit, Overdone Selloff, Compute Oversupply Fears
Muse Spark 1.1
Meta’s newest most powerful API model is Muse Spark 1.1. On their first day of public access, it outperformed some of the leading frontier models on agentic workflows, such as Anthropic’s Opus 4.8 on MAX.
Source: Meta.AI
From the analysis above, Muse Spark 1.1 is the current leader in agent benchmarks in 4/6 categories. Where it fails to beat Claude and OpenAI is in hardcore coding (SWE-Bench, DeepSWE) and Multimodal categories, which test software engineering and the ability to analyze, reason across, and generate different formats.
For Meta, this is a generational breakthrough. For years, analysts have been scorning Zuckerberg for his lacklustre AI spending, committing hundreds of billions of dollars to a buildout that simply didn’t have a clear link to revenues. Their underlying ads and social media business was more than able to fund the buildout; the question of ROI, however, was always overhanging. We believe that Muse Spark 1.1 is the first of many strides forward that will crown Meta as the leader of AI.
Meta now has a competitive frontier model and is monetizing it through token usage. This is the first time in Meta’s history that they’ve made users pay for their AI models, and it’s clear what their objective is. They are trying to offer the most efficient agentic model that can operate at the lowest cost per token. Meta will pose a serious threat to both OpenAI and Anthropic simply because of their token costs and heightened efficiency.
Source: Meta.AI
The Tokenomics:
As we’ve come to understand, the newest, most powerful GPUs are reserved specifically for training large-parameter frontier models. As newer generations of chips get released, the older ones slowly lose their value. H100 and H200 equivalents used to be the most powerful on the market and have been pushed toward less intensive use cases. Simply put, these chips no longer compare to the current best on the market, but that doesn’t mean they’ve gone entirely obsolete. These older-generation chips are still perfectly viable for inference and agentic use cases.
Source: Silicon Data
This chart from Silicon Data models the H100 term rate curves, showing a substantial rise in prices through 2026 when compared to November, 2025. This chart makes economic sense as to why Meta is interested in entering the “neocloud” business, as the assets they’ve stockpiled are appreciating. What Meta is aiming to achieve does not signify a compute oversupply, but rather a generational shift in financing compute. The vast majority of chips that are being rented out from Meta will be their old stockpile of H100 and H200’s, while maintaining Blackwell and Rubin chips for their frontier model training. Those higher-end GPU’s will be supply-constrained through 2028, as every single company on the planet requires them to be competitive. What also becomes supply-constrained will be the older GPUs that are now obsolete for frontier training. As we know, leading-edge GPU’s like Rubin and Blackwell deliver roughly 3x tokens per megawatt versus H100, with Nvidia claiming up to 15x lower cost per token on optimized workloads. In a vacuum, the leading-edge GPUs are a no-brainer choice, but not when they are completely sold out. This AI super cycle is causing the strongest compute to be utilized specifically for training frontier models. Even Jensen Huang, while announcing that Nvidia had secured supply for ‘‘very, very robust growth,’’ admitted the company remains supply constrained.
Blackwell's headline advantages, such as FP4 compute, giant NVLink domains, and 8 TB/s bandwidth, matter most for serving massive frontier models at high interactivity. But the token volume concentrated at the frontier could be rather misleading: Beneath it sits a huge layer of everyday enterprise AI. Small fine-tuned models, embeddings, batch summarization, document-lookup pipelines, and ordinary chatbots that don’t need frontier models or frontier hardware at all. For a bandwidth-bound 70B decode workload, an H100 isn’t something you should frown at; it's actually correctly sized to operate under these parameters.
The trade-off is token efficiency; on SemiAnalysis's benchmarks, Blackwell serves tokens 4–9x cheaper than Hopper on optimized workloads. But per-token efficiency only matters if you can get the hardware, and for workloads that fit comfortably in an H100's 80GB, oftentimes the cheapest option becomes the most rational choice. Market data pulled from SemiAnalysis’s InferenceMAX platform shows that the cost per 1 million tokens for an H100 and a Blackwell chip per 1 million tokens is $0.14 and $0.03, respectively. The reason Hopper chips are even viable is the spread between token prices and token costs. API tokens for large open models retail at $0.20-$1.00 per million; SemiAnalysis's benchmarks put serving costs for these models at $0.03–0.55 per 1 million tokens, depending on the chip. At spreads like that, even hardware with several times worse token economics still runs at enormous gross margins. The efficiency gap between legacy and frontier chips decides who earns more, and not who earns.
The Hardware Bonanza:
Hardware was sold off hard over the past week due to the misconception around Meta’s announcement. Many individuals took the selling of “excess” compute as an indication of an AI oversupply. We are here to tell you that this is factually incorrect. To begin, if there were truly an oversupply of compute, then why is Meta renting 1.6 GW worth of capacity from Crusoe, committing $27 billion to Nebius, $35.2 billion to CoreWeave, and just recently having Google cap their Gemini usage? Companies that are drowning in “excess” capacity do not exhibit behaviours like this. So why are they? The answer is only found between the lines… Meta selling their older generation chips while simultaneously locking up neocloud capacity is not contradictory. They are selling yesterday’s viable compute, while locking up tomorrow’s scarce capacity. This is Meta’s way of monetizing every part of their AI business, while also funding the next round of Capex. One of Meta’s largest criticisms was their lacklustre spending with no real sign of ROI. Muse Spark 1.1 and their “neocloud” proposal refute all claims of “lack of ROI.” The FCF and narrative legitimacy they can acquire from this can help Meta raise Capex guidance without being punished for it; AKA fuelling the AI trade for even longer. We believe this will cause a snowball effect of Capex guidance being revised upward. We predict that Capex from hyperscalers could see $1.0-$1.3 trillion in 2027.
Source: Aria Research
Compute & Neoclouds:
The following section will talk thoroughly about the impacts on the compute and neocloud sectors.









