Research dossier · September 2026AI infrastructure economics · metering and allocation

The operator already pays for it. The blind spot is attribution.

A neutral layer that joins compute telemetry to power, cooling and facility infrastructure, so an operator can meter, attribute, price, and plan the physical economics of AI compute per tenant.

Product
GridMind
Lab role
First paid design-partner pilot
Product status
Ready for pilots · read-only · one zone
Research status
Assessed internally · no external validation · data access is the gating risk

Customer: GPU clouds, neoclouds and colocation operators running multiple tenants. Everything below is a working argument for what to build next — not a claim about what already exists.

The lab decision

What we are actually deciding

A thesis is only useful if it changes what gets built. This is the call, the reasoning, and the condition that would reverse it.

Decision

Prioritize GridMind for the first paid pilot, conditional on securing one operator who grants read-only telemetry access and names a commercial decision that depends on a better number.

Reversal condition

If no operator grants read-only compute and power telemetry within roughly four weeks, the Agent Environment becomes the temporary primary validation track, because it can be tested faster with less infrastructure access.

Why
  • The buyer already carries the cost line; there is no need to create a new budget or new demand.
  • Value is provable before billing changes: a reconciled cost model is useful as evidence on its own.
  • The work integrates with systems operators already run, rather than asking them to adopt a new stack.
  • The moat — a normalized economic model and history — compounds with every zone measured.
Next gate
One cluster, one reconciliation report, one decision model, one paid conversion.

If the gate is not met, the honest move is to stop or narrow — not to build a broader product to justify the work already spent.

Head to head

Two theses, one sequence

Both products are incubated by the same lab. They share a way of working, not a codebase, a data plane, or a release cadence.

Dimension
Agent Operating Environment
GridMind
Problem clarity
Agent Operating EnvironmentClear problem, broad and still-forming category
GridMindClear infrastructure-economics problem with an existing budget line
Economic buyer
Agent Operating EnvironmentProduct, engineering or operations leader inside the business
GridMindCFO, COO, infrastructure or finance inside a GPU operator
Time to first proof
Agent Operating EnvironmentFaster: one API and one governed action is enough to show the idea
GridMindSlower: telemetry access, attribution and reconciliation come first
Revenue clarity
Agent Operating EnvironmentUncertain until real agent demand reaches the business
GridMindStronger: billing, margin, power cost and capacity decisions already exist
Main technical risk
Agent Operating EnvironmentDelegated identity, authorization, side effects and liability
GridMindHeterogeneous telemetry, attribution accuracy, facility integration
Platform risk
Agent Operating EnvironmentHigh: protocols and platforms can absorb generic connectivity
GridMindMedium: adjacent vendors can move, but facility data is sticky
Defensibility
Agent Operating EnvironmentPolicy, identity, integrations and activity history
GridMindNormalized economic model, history, facility-specific allocation
Lab role now
Agent Operating EnvironmentNarrow discovery and prototype track
GridMindFirst paid design-partner pilot, conditional on data access

Read the table as sequencing, not as a verdict on merit. GridMind has the stronger economic case; the Agent Environment has the faster test.

The problem

Tenants buy GPU hours. Nobody can price the rest.

The invoice shows compute. The bill shows electricity. Between them sit cooling, rack and facility capacity, network, storage and shared overhead — and in a multi-tenant GPU environment, none of it is cleanly attributable to the tenant that caused it. Operators answer commercial questions with a spreadsheet and a rule of thumb.

01Attribution

Knowing a tenant has 40 GPUs does not tell you what those GPUs cost in power, cooling and shared capacity.

02Reconciliation

Compute telemetry and facility meters rarely share a clock, a tenant identity, or a unit of account.

03Pricing

Flat GPU pricing hides the cost differences between power zones, cooling headroom and utilization.

04Capacity

Decisions about the next few megawatts are made without a defensible per-tenant cost baseline.

What the 2026 evidence says

Sources, and what they do not prove

Public evidence can validate a problem and still leave the business case open. Each source below is paired with the claim it cannot support.

Electricity 2026 — demand analysisInternational Energy Agency
Establishes

Data-center electricity demand is a material, quantified driver of grid demand, which makes power economics an operator problem, not a footnote.

Does not establish

It does not show that any operator will buy a metering layer for tenant-level allocation.

International Energy Agency source
20,000+ NVIDIA Blackwell Ultra GPUs at 60 MWYotta
Establishes

Commercial GPU capacity is being planned at a scale where power and cooling are first-order commercial constraints.

Does not establish

It is a company announcement; it does not describe an unmet tooling requirement.

Yotta source
OpenCost specificationOpenCost
Establishes

Cost allocation is a solved, standardized discipline in the Kubernetes layer — and it stops at compute.

Does not establish

It does not cover facility power, cooling or rack-level physical economics.

OpenCost source
Run:ai GPU orchestrationNVIDIA
Establishes

Scheduling and cluster utilization are already a mature, well-served product category.

Does not establish

It does not answer what a scheduled workload costs physically, or how to bill for it.

NVIDIA source
FinOps allocation capabilityFinOps Foundation
Establishes

Allocation, showback and chargeback are established disciplines with known methods and failure modes.

Does not establish

It does not validate allocation of physical infrastructure, which has different constraints.

FinOps Foundation source
Competitive read

Who owns which layer

Positioning is a map of who already owns what. A gap only matters if the customer feels the pain on the other side of it.

LayerWho is thereWhat they ownOur opening
DCIM
Who is thereSchneider EcoStruxure, Nlyte, Sunbird, Hyperview
What they ownFacility assets, capacity, power and environmental management.
Our openingHigh-value gap: they know the facility, not the workload or the tenant contract.
Cloud and Kubernetes cost
Who is thereOpenCost, Kubecost, CAST AI
What they ownCompute cost attribution for containers, nodes and GPU workloads.
Our openingThey stop at the edge of the rack; power, cooling and facility overhead are outside the model.
GPU orchestration
Who is thereNVIDIA Run:ai, Kubernetes GPU schedulers
What they ownWhere and when workloads run.
Our openingScheduling decisions are made without the physical cost of the decision.
Energy and facility monitoring
Who is thereBMS, PDU, meter and energy management platforms
What they ownWhat the facility physically consumed.
Our openingFacility consumption is not joined back to tenant identity and workload economics.
Neoclouds and AI clouds
Who is thereYotta, E2E Networks, Neysa, CoreWeave, Nscale and similar
What they ownSelling GPU capacity to customers.
Our openingThey are customers for this layer, not competitors — their margin depends on getting allocation right.
Developer view

What gets built, and what does not

Request path
Signed edge collector→Raw telemetry with lineage→Normalized tenant, workload and resource model→Attribution and allocation engine→Cost and reconciliation report→Operator dashboard and API
Non-negotiable principles
  • Read-only first: observe and reconcile before anything is recommended, and long before anything is controlled.
  • Every cost line is labelled measured, allocated, or not claimed — with its source, timestamp and confidence.
  • Cooling and overhead are apportioned under a rule the operator approves, shown under alternatives.
  • Facility-specific allocation is configuration, not a rewrite — the engine is shared, the model is per site.
  • Signed agents, encrypted telemetry, tenant isolation, RBAC, audit, and edge operation where the network requires it.
MVP shape

One cluster or zone, one or two tenants, 8–32 GPUs, at least fourteen days of representative load, reconciled against existing meters and current allocation rules — three to four weeks.

In scope
  • Collectors for existing GPU, workload and tenant telemetry
  • At least one layer of power measurement: PDU, rack, UPS or facility meter
  • Normalization into tenant, workload, resource, power, cooling and facility lines
  • An allocation methodology agreed in writing before any number is produced
  • Reconciliation against invoices and meters, with every discrepancy listed
  • A cost model for one named commercial decision, with confidence ranges
Explicit non-goals
Any control action, workload migration, or change to power distributionGenerating customer invoices or changing contractual termsReplacing DCIM, scheduler or billing platformAccredited financial measurement claimsWhole-facility modelling before one zone is proven
Business view

Who pays, for what

Buyer

A GPU cloud, neocloud, managed GPU provider or high-density colocation operator running multiple tenants, with a pending decision on pricing, power pass-through or expansion.

Value
  • A tenant-level cost view that separates compute from power, cooling and overhead.
  • Defensible allocation rules that survive a customer negotiation.
  • Evidence for power pass-through, energy-aware pricing or premium tiers.
  • Visibility into underutilized capacity and where the next constraint will appear.
  • A shared economic model for infrastructure, finance and commercial teams.
Pricing
  • Fixed-fee pilot scoped to one zone, with a written reconciliation report
  • Monthly platform fee per monitored cluster or infrastructure zone
  • Optional modules: energy-aware billing, power pass-through, capacity planning, optimization
  • No transaction take rate; customer contracts and tariffs stay with the operator
Moat
  • The normalized tenant to workload to compute to power to cooling to cost model
  • Facility-specific allocation models and the history behind them
  • Deep integrations with GPU, Kubernetes, DCIM, PDU, BMS and meter systems
  • Operator workflow lock-in once billing and planning depend on the numbers
Risks

What would make this wrong

Poor telemetry produces confident, wrong economics.
Response · Expose lineage, confidence and gaps; label every line measured, allocated or not claimed.
Every facility is different, so every pilot is custom.
Response · Adapters plus configuration; a narrow zone; read-only first; no bespoke engine forks.
DCIM, FinOps or orchestration vendors move into attribution.
Response · Own the workload-to-physical workflow and stay hardware-neutral and integrator-friendly.
The pilot becomes a data consultancy.
Response · Fixed scope, paid pilot, standard collectors, reusable report, explicit production conversion path.
A control feature touches physical infrastructure.
Response · Control stays out of scope until authorization, rollback and safety review are complete.
Validation

The next four weeks

A thesis becomes real when a stranger pays for the outcome. These are the steps, the gates, and the conditions under which we stop.

Week 0
Access and method

Source inventory, meter boundaries, tenant labels and a written methodology agreed before any number is produced.

01
Weeks 1–2
Collect and normalize

Read-only collection across the zone. No writes, no control, no invoices.

02
Week 3
Reconcile

Compare the model against meters and invoices, and list everything that cannot be explained.

03
Week 4
Decide

Model the named commercial decision, state the conditions under which it holds, and get a paid-conversion decision.

04
Gates that must pass
  • At least fourteen days of representative load with trustworthy meters
  • A named operations or finance owner who signs off on the output
  • A meaningful share of consumption reconciled, with the remainder explained
  • One commercial decision with a quantified financial or capacity impact
  • Agreement to a paid production phase, not just an enthusiastic readout
Stop or narrow if
  • No read-only access to both compute telemetry and one layer of power measurement
  • No decision is waiting on a better number — the data would be interesting but useless
  • Unattributable usage dominates the zone and cannot be reduced
  • The operator insists on billing-grade accuracy the instrumentation cannot support

Evidence here is public energy reporting, vendor specifications and one operator scale announcement. It establishes that physical constraints are real and that adjacent categories stop short of tenant-level physical economics. It does not establish that operators will buy — that remains the open question, and the first paid pilot is the test.

GridMind

Read the argument, then test it with us

Every claim on this page is falsifiable. If you run the infrastructure, or the business, this is where we would rather be corrected than impressed.

RelayForge is a product lab. SentraZero, PowerIQ and the Outbound Engine are case studies, not products.