← All briefs

Relay Brief

Relay Brief — Jul 31, 2026

2026-07-31America/ChicagoMode · normalx_search ok

Opening

Yesterday's physical-stack story got sharper overnight: PJM-scale curtailment and FERC filings moved the power wall from abstract grid talk toward rules that hit large AI loads. On the software side, the center of gravity is still agentic inference economics (tokens per MW, memory beyond HBM, tool-loop wall-clock) plus a fresh de-stealth cluster around agent spend controls, robotics interfaces, visual data layers, and specialized model-lab serving. Seed handles remain loud on Grok cadence and dense clusters; little net-new from Jensen in this window. Power problem framing is consensus; company maps and agent-control layers are early/forming.

Deep pulls

PJM curtailment + FERC path — power becomes a load rule, not just a queue

forming

Discussion of PJM board-directed FERC filings (large-load registry, reliability/capacity measures, multi-GW auction language) and planned curtailment of large loads (~50 MW+) without self-supply during shortages, with demand-response compensation. Behind-the-meter/onsite generation framed as the practical escape hatch.

Why it matters · Turns long-term AI campus economics into interconnection + self-supply + curtailment risk, not only GPU lead times. Matches prior power thesis with a clearer U.S. ISO-level mechanism.

Observed: X summaries of board/filing/curtailment language from energy and grid-focused accounts.

Interpretation: Hyperscalers with BTM generation and flexible load design gain option value under curtailment regimes.

Action: Read primary PJM Inside Lines / FERC docket summaries next scan; treat single-ticker AI power posts as noise without multi-source corroboration.

Agentic inference: tokens/MW, memory wall, tool-loop time

forming

Stack framing shifts from chips-per-rack to tokens per megawatt, with agentic loads pulling memory beyond HBM (DDR5/LPDDR5/NAND for KV cache, head nodes). Workload analysis argues agent harnesses are prefill-heavy and decode is bandwidth-bound; some posts say tool calls/retries/context dominate wall-clock more than raw decode. Vertical compute-under-DRAM / 3D stacking threads aim at data-movement energy.

Why it matters · If agents multiply chat token burn and bottleneck on memory + orchestration, research edge moves to serving architecture, memory hierarchy, and agent runtime—not just next frontier training run.

Observed: Technical threads on tokens/MW, 55M-request serving analysis, and memory-wall packaging.

Interpretation: Inference-ratio and orchestration efficiency may matter more than training-size brags for multi-year infra spend.

Action: Track inference-ratio and tokens/Joule narratives separately from training-footprint posts.

De-stealth / agent stack: payments control, robotics UI, visual SQL, model-lab GTM

early

Fresh out-of-stealth cluster: SkalorAI (agent payment control/clearing layer; Draper Dragon); Enigma (~$71M seed robotics foundation models/interfaces; Index/Ribbit); CreativAI (SQL layer for physical/visual AI); Baseten for Model Labs (enterprise serving for specialized closed-weight labs); Agon defense/sovereign training infra still circulating. Builder side: multi-agent security review loops.

Why it matters · Continues shift above the foundation model: agent money rails, physical AI data, specialized-model distribution, defense infra.

Observed: Primary company launch posts plus secondary raise roundups on X.

Interpretation: Picks-and-shovels for agents and physical AI are denser than new general-purpose FM launches in this window.

Action: Watchlist candidates only—confirm product scope and primary funding notes before elevating any name.

Seed handles: Grok cadence + dense GB300 clusters; quiet Jensen window

forming

elonmusk continues Grok 4.6 improvement/near-term ship chatter, agent-bench rankings, Grok Build tooling, Voice agentic ranking, and dense Minihard / 220k GB300 + high-speed NIC language. JensenHuang filter returned little citable new primary AI post in this date window.

Why it matters · Public window into fast model loops + industrial cluster density; agentic voice/tooling claims stay more relevant than pure param counts.

Observed: Primary elonmusk posts on Grok product and cluster density; sparse Jensen primary hits in-window.

Interpretation: xAI optimizing for fast product loops atop huge GPU footprints remains the durable seed narrative.

Action: Separate agentic latency/tool-use claims from training-footprint posts.

People moving

  • watch

    Useful for PJM/FERC large-load signal; verify against primary docs.

  • watch

    Power vs announced compute framing (EU and global).

  • follow

    Higher-signal memory-wall / packaging technical thread.

  • follow

    Tokens-per-MW and full-stack memory framing.

  • follow

    Production serving workload analysis for agents.

  • watch

    Primary voice on agent payment control plane de-stealth.

  • watch

    Robotics foundation model / interface de-stealth and demos.

  • watch

    CreativAI physical/visual data layer primary.

  • watch

    Baseten for Model Labs launch primary.

  • watch

    Highest-velocity public Grok/cluster feed; high noise floor.

  • @JensenHuang

    watch

    Quiet in this scan; keep for open-secure/platform follow-through.

Narratives to track

  • ISO-level AI load rules

    forming

    Curtailment, large-load registries, BTM self-supply

  • Agentic inference economics

    forming

    Tokens/MW, prefill vs decode, tool-loop latency

  • Memory beyond HBM / 3D stacking for decode

    early

    Data-movement energy and packaging

  • Agent commercial control plane

    early

    Payments mandates, kill switches, receipts

  • Specialized model-lab distribution + physical/visual data

    early

    Baseten-for-labs and CreativAI-class layers

  • Defense / sovereign agent-physical infra

    early

    Agon-class; multi-signal pending

Early board / long-term radar

NameKindThesisConfidence
Time-to-power + load flexibility for AI campusestheme

Multi-year bottleneck is electrons, interconnect, and who can operate under curtailment rules—not only accelerator supply.

Signal: PJM/FERC large-load and BTM threads on X.

Missing: Primary docket text and independent utility/ISO data.

medium
Agentic serving stack (memory + runtime + orchestration)theme

Agent workloads flip spend toward inference efficiency and non-GPU bottlenecks.

Signal: Tokens/MW and harness/serving analyses on X.

Missing: Cloud utilization series and OEM primary benchmarks.

low
Agent economic rails and fiduciary controlstheme

Once agents spend, control planes become budget lines.

Signal: SkalorAI de-stealth and agent-payments discussion.

Missing: Real volume/retention proof beyond launch claims.

low
Specialized model labs + GTM infratheme

Multi-model ecosystem needs enterprise packaging (Baseten-for-labs pattern).

Signal: Partner-list launch posts from Baseten principals.

Missing: Take-rate and retention evidence.

low
SkalorAIcompany

Watchlist candidate for agent payment control/clearing layer.

Signal: Primary de-stealth posts; Draper Dragon mention.

Missing: Primary funding docs, customers, volume proof.

low
Enigmacompany

Watchlist candidate for robotics foundation models/interfaces.

Signal: Secondary $71M seed coverage; demo experiment posts.

Missing: Primary raise confirmation and product retention.

low
CreativAIcompany

Watchlist candidate for structured visual/physical AI data layer.

Signal: Founder primary de-stealth post.

Missing: Traction and competitive map off-X.

low
Agoncompany

Watchlist candidate for sovereign/defense physical AI training infra.

Signal: De-stealth roundup posts continuing from prior day.

Missing: Multi-source primary confirmation.

low

Worth a closer look

Open questions

  • Exact PJM curtailment start date, MW threshold, and FERC filing status in primary sources?
  • How much of agent wall-clock is truly tool/orchestration vs decode across production systems?
  • SkalorAI / Enigma / CreativAI: primary funding docs and product customers beyond launch posts?
  • Does Baseten for Model Labs change who captures inference margin vs hyperscalers?
  • Any new Jensen primary posts on open-secure / agent platform this week?

Sources