Relay Brief
Relay Brief — Jul 31, 2026
Opening
Yesterday's physical-stack story got sharper overnight: PJM-scale curtailment and FERC filings moved the power wall from abstract grid talk toward rules that hit large AI loads. On the software side, the center of gravity is still agentic inference economics (tokens per MW, memory beyond HBM, tool-loop wall-clock) plus a fresh de-stealth cluster around agent spend controls, robotics interfaces, visual data layers, and specialized model-lab serving. Seed handles remain loud on Grok cadence and dense clusters; little net-new from Jensen in this window. Power problem framing is consensus; company maps and agent-control layers are early/forming.
Deep pulls
PJM curtailment + FERC path — power becomes a load rule, not just a queue
formingDiscussion of PJM board-directed FERC filings (large-load registry, reliability/capacity measures, multi-GW auction language) and planned curtailment of large loads (~50 MW+) without self-supply during shortages, with demand-response compensation. Behind-the-meter/onsite generation framed as the practical escape hatch.
Why it matters · Turns long-term AI campus economics into interconnection + self-supply + curtailment risk, not only GPU lead times. Matches prior power thesis with a clearer U.S. ISO-level mechanism.
Observed: X summaries of board/filing/curtailment language from energy and grid-focused accounts.
Interpretation: Hyperscalers with BTM generation and flexible load design gain option value under curtailment regimes.
Action: Read primary PJM Inside Lines / FERC docket summaries next scan; treat single-ticker AI power posts as noise without multi-source corroboration.
Agentic inference: tokens/MW, memory wall, tool-loop time
formingStack framing shifts from chips-per-rack to tokens per megawatt, with agentic loads pulling memory beyond HBM (DDR5/LPDDR5/NAND for KV cache, head nodes). Workload analysis argues agent harnesses are prefill-heavy and decode is bandwidth-bound; some posts say tool calls/retries/context dominate wall-clock more than raw decode. Vertical compute-under-DRAM / 3D stacking threads aim at data-movement energy.
Why it matters · If agents multiply chat token burn and bottleneck on memory + orchestration, research edge moves to serving architecture, memory hierarchy, and agent runtime—not just next frontier training run.
Observed: Technical threads on tokens/MW, 55M-request serving analysis, and memory-wall packaging.
Interpretation: Inference-ratio and orchestration efficiency may matter more than training-size brags for multi-year infra spend.
Action: Track inference-ratio and tokens/Joule narratives separately from training-footprint posts.
De-stealth / agent stack: payments control, robotics UI, visual SQL, model-lab GTM
earlyFresh out-of-stealth cluster: SkalorAI (agent payment control/clearing layer; Draper Dragon); Enigma (~$71M seed robotics foundation models/interfaces; Index/Ribbit); CreativAI (SQL layer for physical/visual AI); Baseten for Model Labs (enterprise serving for specialized closed-weight labs); Agon defense/sovereign training infra still circulating. Builder side: multi-agent security review loops.
Why it matters · Continues shift above the foundation model: agent money rails, physical AI data, specialized-model distribution, defense infra.
Observed: Primary company launch posts plus secondary raise roundups on X.
Interpretation: Picks-and-shovels for agents and physical AI are denser than new general-purpose FM launches in this window.
Action: Watchlist candidates only—confirm product scope and primary funding notes before elevating any name.
Seed handles: Grok cadence + dense GB300 clusters; quiet Jensen window
formingelonmusk continues Grok 4.6 improvement/near-term ship chatter, agent-bench rankings, Grok Build tooling, Voice agentic ranking, and dense Minihard / 220k GB300 + high-speed NIC language. JensenHuang filter returned little citable new primary AI post in this date window.
Why it matters · Public window into fast model loops + industrial cluster density; agentic voice/tooling claims stay more relevant than pure param counts.
Observed: Primary elonmusk posts on Grok product and cluster density; sparse Jensen primary hits in-window.
Interpretation: xAI optimizing for fast product loops atop huge GPU footprints remains the durable seed narrative.
Action: Separate agentic latency/tool-use claims from training-footprint posts.
People moving
watch
Useful for PJM/FERC large-load signal; verify against primary docs.
watch
Power vs announced compute framing (EU and global).
follow
Higher-signal memory-wall / packaging technical thread.
follow
Tokens-per-MW and full-stack memory framing.
follow
Production serving workload analysis for agents.
watch
Primary voice on agent payment control plane de-stealth.
watch
Robotics foundation model / interface de-stealth and demos.
watch
CreativAI physical/visual data layer primary.
watch
Baseten for Model Labs launch primary.
watch
Highest-velocity public Grok/cluster feed; high noise floor.
- @JensenHuang
watch
Quiet in this scan; keep for open-secure/platform follow-through.
Narratives to track
ISO-level AI load rules
formingCurtailment, large-load registries, BTM self-supply
Agentic inference economics
formingTokens/MW, prefill vs decode, tool-loop latency
Memory beyond HBM / 3D stacking for decode
earlyData-movement energy and packaging
Agent commercial control plane
earlyPayments mandates, kill switches, receipts
Specialized model-lab distribution + physical/visual data
earlyBaseten-for-labs and CreativAI-class layers
Defense / sovereign agent-physical infra
earlyAgon-class; multi-signal pending
Early board / long-term radar
| Name | Kind | Thesis | Confidence |
|---|---|---|---|
| Time-to-power + load flexibility for AI campuses | theme | Multi-year bottleneck is electrons, interconnect, and who can operate under curtailment rules—not only accelerator supply. Signal: PJM/FERC large-load and BTM threads on X. Missing: Primary docket text and independent utility/ISO data. | medium |
| Agentic serving stack (memory + runtime + orchestration) | theme | Agent workloads flip spend toward inference efficiency and non-GPU bottlenecks. Signal: Tokens/MW and harness/serving analyses on X. Missing: Cloud utilization series and OEM primary benchmarks. | low |
| Agent economic rails and fiduciary controls | theme | Once agents spend, control planes become budget lines. Signal: SkalorAI de-stealth and agent-payments discussion. Missing: Real volume/retention proof beyond launch claims. | low |
| Specialized model labs + GTM infra | theme | Multi-model ecosystem needs enterprise packaging (Baseten-for-labs pattern). Signal: Partner-list launch posts from Baseten principals. Missing: Take-rate and retention evidence. | low |
| SkalorAI | company | Watchlist candidate for agent payment control/clearing layer. Signal: Primary de-stealth posts; Draper Dragon mention. Missing: Primary funding docs, customers, volume proof. | low |
| Enigma | company | Watchlist candidate for robotics foundation models/interfaces. Signal: Secondary $71M seed coverage; demo experiment posts. Missing: Primary raise confirmation and product retention. | low |
| CreativAI | company | Watchlist candidate for structured visual/physical AI data layer. Signal: Founder primary de-stealth post. Missing: Traction and competitive map off-X. | low |
| Agon | company | Watchlist candidate for sovereign/defense physical AI training infra. Signal: De-stealth roundup posts continuing from prior day. Missing: Multi-source primary confirmation. | low |
Worth a closer look
- PJM large-load / curtailment primary docs
Use as pointer to ISO/FERC sources
- wafer_ai memory-wall / 3D stacking thread
Technical packaging angle
- jacobeverly 55M-request serving analysis
Agent harness workload shape
- SkalorAI primary announcements
Agent spend control plane
- CreativAI de-stealth
Visual SQL layer
- Baseten for Model Labs
Specialized lab GTM/serving
- Enigma robotics experiment coverage
Seed + robots.online narrative
Open questions
- Exact PJM curtailment start date, MW threshold, and FERC filing status in primary sources?
- How much of agent wall-clock is truly tool/orchestration vs decode across production systems?
- SkalorAI / Enigma / CreativAI: primary funding docs and product customers beyond launch posts?
- Does Baseten for Model Labs change who captures inference margin vs hyperscalers?
- Any new Jensen primary posts on open-secure / agent platform this week?