GPU COMPUTE INTELLIGENCE

Where AI infrastructure
meets commodity markets.

ComputeAi is a London-based research and proprietary trading operation at the intersection of AI compute infrastructure and energy commodity markets.


WHO WE SERVE

Intelligence for participants
at the compute frontier.

The GPU compute futures market is new. The participants entering it come from different worlds — finance, energy, infrastructure, technology. We bridge them.

DATA CENTRE LENDERS
Credit exposure to GPU revenue
Banks and credit funds financing data centre builds need a GPU forward curve to underwrite their loans — the same way oil lenders use WTI futures to price E&P credit risk. We provide the basis intelligence that makes those underwriting assumptions defensible.
ENERGY TRADERS
GPU demand as a power market signal
Data centre load is now a measurable and growing component of ERCOT and PJM demand. The GPU compute forward curve is the demand signal that power developers and energy traders have been missing. We connect the compute market to the energy market it depends on.
INSTITUTIONAL PARTICIPANTS
Systematic intelligence for a new futures market
Quant funds and proprietary trading operations entering GPU1/GPU2 on CME need physical market context — supply, basis, index divergence, quality-adjusted pricing — that exchange data alone does not provide. We track the physical market so you understand what the futures are pricing.
NEO-CLOUD OPERATORS
Revenue certainty in a volatile market
GPU rental prices have moved from $10/hr at peak scarcity to $2.50 and back. Operators who understand the forward curve and the basis between their actual pricing and the settlement index can hedge revenue exposure with precision. We provide the market intelligence that makes that possible.
AI INFRASTRUCTURE TEAMS
Compute procurement intelligence
Finance and infrastructure teams at AI companies face GPU rental costs as their largest variable expense. Understanding where physical prices are relative to index benchmarks — and what the forward curve implies about future costs — is essential for budgeting and hedging.
SEMICONDUCTOR SUPPLY CHAIN
GPU rental as a leading demand indicator
GPU rental prices lead hardware order flow by 6-12 months. A sustained softening in on-demand GPU prices is an early signal of data centre capex deceleration. We provide the daily physical market readings that make that signal legible before it appears in quarterly earnings.

MARKET KNOWLEDGE

The language
of compute markets.

GPU compute is moving from a technology procurement decision to a commodity market. We speak both languages.

PHYSICAL MARKET

GPU-Hour
The base unit of compute pricing. One GPU made available for one hour, regardless of utilisation. The underlying commodity in all compute futures contracts.
Physical Basis
The difference between the on-demand spot price for a GPU and its settlement index price. Positive basis means physical is above the index. The basis is what hedgers cannot eliminate.
Quality-Adjusted Basis
The basis restated after controlling for GPU variant, interconnect type, and pod configuration. Compares like-for-like hardware rather than blended averages. The correct basis for trading analysis.
Grey Route
GPU rental listings from resellers and secondary market providers rather than hardware owners. Grey route supply creates price divergence within the spot market and distorts blended index calculations.
Generational Spread
The price ratio between current-generation GPU hardware and next-generation hardware. When the physical spread diverges from the index spread, a relative-value trade opportunity emerges.
GPU Breakeven
The minimum rental price at which a provider covers hardware amortisation, power, colocation, and operations. The floor below which providers exit the market. Varies significantly by geography and power contract.
Reserved Capacity
GPU capacity contracted for a fixed term at a discount to on-demand pricing. Provides revenue certainty for providers and cost predictability for buyers. Settlement indices typically blend reserved and on-demand price signals.
Spot Capacity
GPU capacity priced at real-time supply and demand, typically at a significant discount to on-demand but subject to interruption. Spot pricing creates intraday volatility that blended indices smooth over.
Utilisation Rate
The proportion of available GPU-hours actually rented and generating revenue. The most important operational metric for GPU rental economics — low utilisation compresses margins; high utilisation precedes price increases.
Burst Capacity
Temporary GPU capacity available above a reserved baseline, priced at spot or on-demand rates. When burst capacity is unavailable, spot prices spike — the mechanism that drives index volatility.
Neo-cloud
Independent GPU compute providers renting hardware on-demand to AI companies, researchers, and developers. The neo-cloud market is the physical market that GPU futures settlement indices primarily track.
Hyperscaler
A cloud provider operating at a scale that enables near-linear cost reduction with capacity. Hyperscalers own their own data centres and GPUs. Their internal transfer prices rarely appear in public spot markets.
8-GPU Pod
The standard unit of GPU compute supply — eight GPUs connected within a single server node. The pod is the natural unit for LLM training and the reference configuration for most settlement index calculations.
Colocation
A facility model where GPU hardware is housed in a third-party data centre, paying for space, power, and cooling. Colocation separates infrastructure ownership from GPU ownership — relevant for understanding provider cost structures.
Data Centre Credit
Debt financing secured against data centre assets and their GPU rental revenue streams. The GPU futures forward curve makes revenue assumptions in credit models testable for the first time.
Revenue Covenant
A lender protection in a data centre loan agreement requiring the borrower to maintain GPU rental revenue above a defined threshold. The compute futures forward curve provides the observable reference price for such covenants.
Inference Economics
The unit economics of running AI model inference at scale: revenue per million tokens versus cost per GPU-hour. As open-weight models compress token prices, the GPU rental cost structure inference revenue must support becomes the binding constraint.

HARDWARE

SXM
Server Extension Module — the high-performance GPU form factor for data centre use. SXM GPUs connect via NVLink and draw more power than PCIe variants. The standard form factor for AI training clusters — commands a significant price premium.
PCIe
Peripheral Component Interconnect Express — the standard interface for connecting GPUs to server motherboards. PCIe GPUs are cheaper and more widely available than SXM but offer lower inter-GPU bandwidth. Common in inference deployments.
HBM
High Bandwidth Memory — stacked DRAM dies mounted directly on the GPU package via CoWoS. Delivers over 8 TB/s memory bandwidth. Memory bandwidth, not raw compute, is the binding constraint in large-model inference.
CoWoS
Chip-on-Wafer-on-Substrate — the advanced semiconductor packaging technology required to integrate HBM memory with GPU dies. The binding production constraint on H100 and B200 supply. Cannot be expanded quickly.
NVLink Fabric
A high-speed proprietary interconnect linking multiple GPUs into a unified compute fabric. NVLink configurations command a significant premium over InfiniBand or Ethernet-connected pods — a core driver of quality basis.
NVSwitch
A dedicated switching chip enabling full all-to-all NVLink connectivity across all GPUs in a server. Required for large NVLink fabric configurations. Each NVSwitch adds cost and power overhead but enables collective communication at near-memory bandwidth speeds.
InfiniBand
A high-speed networking fabric used to interconnect GPU nodes within and between racks. Delivers lower latency and higher bandwidth than Ethernet for distributed AI workloads. The dominant interconnect in large AI training clusters.
RoCE
RDMA over Converged Ethernet — delivers remote direct memory access over standard Ethernet infrastructure. Lower cost than InfiniBand but higher latency. Increasingly used in GPU clusters where cost matters more than peak performance.
RDMA
Remote Direct Memory Access — enables one GPU to read or write the memory of another without involving the host CPU. Essential for high-performance collective communication in distributed training. Reduces latency by orders of magnitude versus standard networking.
Fat-Tree Topology
A network architecture where bandwidth increases toward the core. The standard topology for large GPU clusters — provides non-blocking all-to-all connectivity at the cost of significant switching infrastructure.
Thermal Design Power
The maximum sustained power a GPU draws under load — the input to every GPU power cost calculation. H100: 700W. B200: 1,000W. Rising TDP per generation is the primary driver of data centre power demand growth.
Rack Density
Power consumption per rack in kilowatts. Standard enterprise IT: 6–10 kW. GPU training clusters: 60–130+ kW. Rising rack density is the binding constraint on legacy facility upgrades and the primary driver of liquid cooling adoption.
Liquid Cooling
Direct liquid cooling of GPU chips and server components, required above approximately 40 kW per rack. Reduces PUE compared to air cooling. Next-generation GPU racks at 130+ kW require liquid cooling as standard.
Power Use Effectiveness
Total data centre energy consumption divided by IT equipment energy consumption. A PUE of 1.35 means 35% overhead for cooling and power delivery beyond the GPU's own draw. The multiplier that converts GPU power draw into total electricity cost.

ENERGY

Power Purchase Agreement
A long-term contract between a generator and a buyer fixing the price per megawatt-hour — typically 10–25 years. The primary instrument data centre operators use to manage power cost exposure independently of spot market volatility.
Nuclear PPA
A Power Purchase Agreement with a nuclear generator providing fixed-price baseload electricity. Increasingly signed by data centre operators to lock in predictable power costs independent of gas price volatility.
Spot Power
Electricity purchased at real-time market prices. Spot power prices are highly volatile — they spike during peak demand and can go negative during surplus generation. Data centres exposed to spot power face significant GPU breakeven risk.
Peaker Plant
A power generation facility that operates only during peak electricity demand. Typically gas-fired. Peaker plants set the marginal cost of electricity during high-demand periods — the price that data centre spot power buyers face when AI workloads spike demand.
Capacity Market
A mechanism in which generators are paid to ensure future capacity availability. Data centres are beginning to participate as flexible demand-side assets — receiving capacity payments in exchange for agreed load curtailment.
Demand Response
A programme in which consumers agree to reduce load during grid stress events in exchange for payments or lower tariffs. Large GPU clusters are emerging as significant demand response assets — their interruptible load is a new grid balancing tool.
Grid Interconnect
The physical connection between a data centre and the electricity transmission grid, measured in megawatts. Queue times of 3–7 years in constrained regions make grid interconnect a binding constraint on data centre expansion.
Load Factor
The ratio of actual energy consumed to the maximum possible energy consumption over a period. A data centre running GPU clusters at full utilisation approaches a load factor of one — the ideal configuration for power contract economics.
Uranium Spot
The price of uranium oxide for immediate delivery, quoted in USD per pound. A leading indicator of nuclear power economics. Rising uranium spot prices increase the long-run cost of nuclear PPAs — the baseload power source increasingly favoured by data centre operators.

AI WORKLOADS

GPU Compute
Rental access to graphics processing units — the specialised chips that power AI training and inference. Measured in GPU-hours, priced by supply and demand, now tradeable as a financial commodity.
Training
The process of adjusting a neural network's parameters by exposing it to large datasets. GPU-intensive, long-duration, and highly parallel. Training runs are the largest single consumer of GPU compute and the primary driver of frontier model development costs.
Inference
Running a trained model to generate outputs. Lower per-GPU intensity than training but higher aggregate demand due to volume. Inference now represents the majority of total GPU-hours consumed globally.
Fine-Tuning
Adapting a pre-trained foundation model to a specific task by training on a smaller, curated dataset. Less GPU-intensive than training from scratch. The dominant workload for enterprise AI teams.
FLOP
Floating Point Operation — the basic unit of neural network computation. Model training cost is typically expressed in FLOPs. A single frontier model training run may consume 10²⁴ to 10²⁵ FLOPs, requiring thousands of GPUs running for months.
Token
The atomic unit of language model input and output — roughly 0.75 words in English. Token volume is the denominator in inference cost calculations and the numerator in inference revenue calculations. The commodity unit of the AI output market.
Throughput
The number of tokens generated per second by an inference system. Higher throughput reduces cost per token but requires more GPU-hours. The trade-off between throughput and latency determines the appropriate GPU configuration for a given workload.
Context Window
The maximum number of tokens a language model can process in a single inference call. Longer context windows require more GPU memory. Models with 128k+ token contexts are memory-bound rather than compute-bound — favouring high-HBM GPU configurations.
Batch Size
The number of inference requests processed simultaneously by a GPU. Larger batch sizes improve GPU utilisation and reduce cost per token but increase latency. Batch size optimisation is the primary lever for inference cost management.

FUTURES & DERIVATIVES

Settlement Index
The daily benchmark price against which a GPU compute futures contract cash-settles. Different index methodologies — offer-based, transaction-based, provider-median — produce materially different settlement prices for the same hardware.
Forward Curve
The term structure of GPU compute prices across future delivery dates — 1 month, 6 months, 1 year, and beyond. The shape of the curve signals whether the market expects current conditions to persist or normalise.
Backwardation
A forward curve where near-term prices are higher than longer-dated prices. In GPU compute, backwardation signals that current scarcity is real and the market expects supply to normalise over time.
Contango
A forward curve where longer-dated prices are higher than near-term prices. Signals the market expects GPU rental costs to rise as demand grows faster than new supply comes online.
Calendar Spread
A position combining a long front-month futures contract and a short back-month contract on the same underlying. Profits when the shape of the forward curve changes — specifically when backwardation deepens.
Compute Swap
A bilateral OTC derivative in which one counterparty pays a fixed GPU rental rate and the other pays a floating rate tied to a published index. The instrument that preceded listed GPU futures and established the market.
Cash Settlement
Futures contract settlement by cash payment of the difference between the contract price and the settlement index price. No physical GPU hours are delivered — the contract is purely financial.
Open Interest
The total number of outstanding futures contracts that have not been settled or closed. Rising open interest signals new money entering the market. A key measure of market depth and participant commitment.
Roll Yield
The profit or loss generated by rolling a futures position from an expiring contract to the next delivery month. In backwardation, rolling generates positive yield. In contango, rolling generates negative yield.
Mark-to-Market
The daily settlement of futures positions at the closing price, with gains and losses transferred between counterparties via margin accounts. Eliminates counterparty credit risk by ensuring all positions are current at end of each trading day.
Initial Margin
The deposit required to open a futures position — typically 5–15% of notional value, set by the exchange clearinghouse. Reflects the maximum expected one-day price move.
Variation Margin
Daily cash flows between a futures position holder and the clearinghouse reflecting mark-to-market gains and losses. Unlike initial margin, variation margin is not returned at position close — it represents realised daily P&L.
Notional Value
The total value of a futures contract — contract size multiplied by settlement price. For GPU futures, notional equals GPU-hours per contract multiplied by the index price.
Futures Basis
The difference between the futures price and the spot index price. In a market with no storage — as GPU compute cannot be stored — the theoretical futures basis reflects only the term premium for price uncertainty over the contract period.
Hedge Ratio
The number of futures contracts required to offset a given physical exposure. Must account for basis risk — the difference between the physical price a hedger actually faces and the settlement index the futures contract tracks.
EFP Mechanism
Exchange for Physical — a mechanism allowing a futures position to be exchanged for an actual physical GPU compute delivery. Enables data centre operators to use the futures market and receive actual compute capacity.
Perpetual Futures
A derivative instrument with no expiry date. Price tracks spot via a periodic funding rate. The instrument used in crypto markets, now being applied to GPU compute — subject to ongoing regulatory review.
Funding Rate
In perpetual futures, a periodic payment between long and short holders that anchors the futures price to the spot index. When futures trade above spot, longs pay shorts. When futures trade below spot, shorts pay longs.

All definitions are provided for educational purposes only and do not constitute financial, investment, or legal advice.


ABOUT

ComputeAi Ltd

London-based. Research and proprietary trading. At the intersection of AI compute infrastructure and energy commodity markets.

Not investment advice. Not FCA regulated. Research is published for informational purposes only.


Get in touch

Data centre lenders, energy traders, neo-cloud operators, and institutional participants tracking the compute futures market. Research enquiries and partnership discussions welcome.

research@computeai.ai