Marvin Labs
AI Capex vs Returns: Signs of the First Plateau
Live with Marvin

AI Capex vs Returns: Signs of the First Plateau

12 min readJames Yerkess, Senior Strategic Advisor

Demand for AI compute is not what changed. Nvidia reported $96.2B of revenue for the quarter ended 26 July 2026, with data center revenue of $89.0B, up 117% on the year, at a 75.0% gross margin. A supplier facing balanced supply and demand does not hold that margin through consecutive quarters of record volume.

What changed is who keeps the money. Roughly two thirds of data center capex goes to GPUs, and the company selling most of them keeps three quarters of every dollar as gross profit. The plateau worth analyzing sits at the return-on-investment level of an individual data center, and the constructive version of it involves those margins moving rather than the spending stopping.

This piece builds on a Marvin Labs LinkedIn Live discussion with Alex Hoffmann (Co-Founder & CEO, Marvin Labs), Max Stamakun, CFA (Co-founder & Portfolio Manager, Israilov Financial LLC), and moderator James Yerkess (Former Global Head of Transaction Banking & FX, HSBC Wealth Management), with the underlying numbers brought up to date.

Two and a half businesses earn the excess returns

Hoffmann's framing on the call was that there are "two and a half people making money" in AI infrastructure. The disclosed margins support it, and the ordering is the opposite of what most models assume.

Nvidia's 75.0% gross margin is the visible half. It is not the highest in the chain.

Micron and SK Hynix both earn more per dollar of revenue than Nvidia does. Samsung is the only one of the four below it, and its 69.6% is consolidated across the whole group, which understates the memory business: Samsung does not report cost of goods sold or gross profit separately for its Device Solutions division. SK Hynix reports under K-IFRS.

The pricing behind those margins shows up at the component level, and it is recent.

DDR5 sat between $2 and $3 per GB through the first three quarters of 2025, peaked at $13.28 in March 2026, and has ranged between $7 and $11 since. DDR3 is the useful control, because no AI build consumes it. It ended August 2026 at $2.83 against $2.42 a year earlier, a move of an entirely different order. These are cheapest listed retail prices, not contracted ones, so they overstate the volatility a hyperscaler actually pays. The timing and the direction are the signal.

The cause is capacity reallocation, not a surge in demand for conventional DRAM. Every wafer moved to HBM is a wafer not producing DDR5, and HBM keeps taking a larger share of what an accelerator costs to build.

On Epoch AI estimates compiled by Stanford DAM, HBM rose from 52% of accelerator component cost in 1Q24 to 63% in 4Q25, with absolute spend up from $1.7B to $11.0B per quarter. Every other component gave up ground to fund it: packaging fell from 18.9% to 14.8%, logic from 14.2% to 12.9%, and auxiliary components from 15.1% to 8.8%.

That moves the question of where competition would have to arrive. GPUs are roughly two thirds of data center spend, and memory is now roughly two thirds of the accelerator bill of materials, which puts memory somewhere near 40% of what a data center costs. Chaining two estimates that way is rough, but the conclusion does not depend on the precision. The largest single claim on the build is held by three suppliers, two of them reporting gross margins above 80% and the third not breaking out the division that would show it.

Google and OpenAI both run in-house chip programs, and Chinese labs are training on domestically developed silicon. None of it has dented Nvidia's margin yet, and even when it does, it addresses the shrinking part of the bill. The constructive version of the plateau now runs through memory pricing more than GPU pricing, which is the better of the two places for it to run, because memory has three suppliers competing rather than one.

AI capex keeps rising while the funding mix deteriorates

The 2026 capital plans are larger than last year's debate assumed. Alphabet alone guided to $180B to $190B of capital expenditure in 2026 and told investors 2027 will increase significantly again. Across the five largest US cloud and AI infrastructure providers, committed 2026 capex runs between $660B and $690B.

How it gets financed has changed more than the headline. Alphabet raised $84.75B of equity in June, its first raise of that kind in around two decades, including a $10B private placement from Berkshire Hathaway. Global AI-related debt issuance is on track for roughly $570B in 2026 on Morgan Stanley's estimate, more than double 2025.

On the call, the reading was that debt markets were supplying whatever the hyperscalers asked for, and that the absence of funding strain was itself a positive signal. The order books have since started to say something different. Coverage ratios on hyperscaler bond sales fell from nearly 5x in February 2026 to below 2x in July. Demand is still there, but the margin of comfort has more than halved inside six months.

For coverage models, that makes spreads the wrong place to look first. They have barely budged. Coverage ratios and new-issue concessions react earlier, and they already have.

Token pricing has split into two markets

Frontier model pricing has held. Everything below it is falling hard, and that split explains most of the commercial behavior worth tracking.

The first half of 2026 was defined by buyers discovering what AI actually cost them, including the case, raised on the call, of Uber consuming a full-year budget in the first quarter. That period is over. OpenAI discounted its non-frontier models by 80%, a level no vendor sustains at a reasonable margin. Anthropic has held headline pricing while folding Fable into its standard plans, which is a price cut expressed as a packaging change.

Buyers adapted in a way that compounds the pressure. Stamakun described the emerging pattern as reserving the strongest model for planning and delegating the rest to something cheaper, on the view that current models already handle the large majority of production tasks. Stripe's acquisition of OpenRouter is the clearest evidence of where this goes. Routing infrastructure exists so that token spend can move to the cheapest adequate model, which is an explicit bet against vendor lock-in by the people paying the bills.

The demand for the compute capacity is always going to be there. It is really a question, can you satisfy this demand in a financially sensible way?

Alex Hoffmann

Cheaper chips would put the neoclouds in a bind

A fall in GPU prices would not be uniformly good news, and the neoclouds are where that shows up first.

CoreWeave's total debt now exceeds $21B, against under $8B in 2024, raised through delayed draw term loans secured on GPUs and on customer contracts. That structure creates a specific problem. Cheaper chips improve the economics of operating the business while reducing the value of the collateral securing its debt.

The ratings show where the risk actually sits. The $8.5B facility closed in March 2026 was the first HPC infrastructure loan of its kind to reach investment grade, but what earns that rating is the creditworthiness of the customer on the other side of the contract, not the operator's own credit, which remains speculative grade. A $3.1B facility closed in May, backed by two non-investment-grade customer contracts, did not get an investment grade rating. There is also a duration mismatch worth modeling: the $2.6B facility closed in August carries a five-year maturity against customer contracts averaging around three years.

Visible stress is more likely to arrive here than at one of the large, heavily indebted buyers, and specifically not at Oracle, where a CDS around 120 is a reasonable level for a BBB- credit and gets misread whenever it appears on a comparison chart. The sequence to watch is a smaller operator failing when collateral values fall at the same time as a revenue commitment lapses.

Depreciation is where the earnings risk sits

Valuation scrutiny has concentrated on multiples. The more useful question is whether the denominator is right, and depreciation is where the answer sits. It is one more place where the scale of AI capex has changed what research has to measure.

Most hyperscalers depreciate AI server equipment over five to six years while Nvidia's product cadence implies a shorter economic life. One estimate puts the resulting understatement at around $176B of depreciation across the industry between 2026 and 2028, with Oracle's 2028 earnings potentially overstated by 26.9% and Meta's by 20.8%.

The disclosure signal is that companies have stopped moving in the same direction. Between 2020 and 2024 the large technology companies steadily extended useful lives. In 2025 that diverged: Amazon shortened the estimate for a subset of servers while Meta extended its own further. When two companies buying similar hardware for similar workloads move their assumptions in opposite directions, at least one estimate is describing something other than equipment.

Two adjustments belong in any model of this sector. Strip out vendor financing, where the supplier is funding its own demand. Strip out gains recognized on equity stakes in other AI companies, which reach GAAP earnings without reflecting operating performance. What remains is the number the multiple should be applied to.

The IPO question moved from access to pricing

Both candidates have already filed. Anthropic filed confidentially on 1 June 2026 and OpenAI followed on 8 June. The open question is timing and price.

The two took opposite routes there. Anthropic chose a narrow position early, optimized for software engineering, and reached enterprise customers across all three major clouds. OpenAI carried a broader customer base and an early Microsoft lock-in. The difference shows up in the timetable. Anthropic's bankers have been lining up investor meetings since July, while OpenAI's CFO told employees in August that the company will be public in 2027 or sooner, which is not the language of a company racing a closing window.

Anthropic raised capital at more than $965B and investors have discussed a float near $2T. Reporting on its filing suggests the risk factors will name AI backlash explicitly. That section deserves close reading when the prospectus becomes public, because risk factors are the one place a company is obliged to argue against itself.

The more hype there is around an IPO, the more likely that the pricing is overpriced.

Max Stamakun

The structural tension is the same in both cases. These businesses are valued like high-margin software companies while carrying the capital requirements of infrastructure companies. Anthropic holding headline pricing while competitors discount is consistent with defending a number into a listing, and it is difficult to sustain when comparable models sell at a fraction of the price.

What to track over the next twelve months

  • Frontier token pricing. The segment that has moved least, and the one still earning a premium. A cut here says the price war has reached it.
  • Memory prices against the 2028 expectation. All three suppliers guide that current pricing is temporary. Memory should normalize before GPU pricing does, so an early move is the first real sign of the supply chain loosening, and the single most useful number on this list.
  • Nvidia and memory gross margins. Erosion here transfers economics to the operators. Nothing in the reported numbers shows it yet.
  • Bond coverage ratios. The fall from 5x to below 2x happened while spreads stayed tight.
  • Useful-life assumptions. A discretionary lever. Track the changes and the direction peers move in.
  • Language in guidance. The shift from capacity constraints to utilization efficiency tends to precede the numbers, as Yerkess noted on the call.
  • The first withdrawal. No hyperscaler has yet stopped spending free cash flow on data centers. Whoever goes first will say something the disclosures have not.

Most of these start as guidance and only become evidence when checked against what the company later delivered. Capex efficiency, revenue run-rates, and useful-life estimates are all asserted by management before they are testable, and the gap between the assertion and the outcome is where the analytical work sits. Guidance Tracking makes that comparison systematic across a coverage list.

For the full conversation, including an audience question on whether AI vendors will price low to build adoption and raise prices once switching gets expensive, watch the video above.

James Yerkess

by James Yerkess

James is a Senior Strategic Advisor to Marvin Labs. He spent 10 years at HSBC, most recently as Global Head of Transaction Banking & FX. He served as an executive member responsible for the launch of two UK neo banks.

Start your free evaluation

Analyze 15 leading companies immediately. No registration required, and a demo is optional, not a gate.