Skip to the content.

Frontier infrastructure and workload expectations index

Latest slice: operator-level autoscaling envelope (#9387)

The existing OpScale row now separates four evidence modes from arXiv v1: an offline analytical opportunity model, production-trace replay on a research prototype, SLO-constrained static/provisioning sweeps, and profiling/model- accuracy microbenchmarks. The prototype uses a common nano-vLLM data plane for OpScale and the ported baselines, evaluates Qwen2-7B and Qwen2-57B-A14B, and reports 929K replayed requests plus 1.5B prompt tokens per model. Most prototype experiments use five Azure VMs with 40 A100-80GB GPUs total, NVLink within each eight-GPU VM, and InfiniBand across VMs; a separate sensitivity study uses 24 GB200 GPUs in one NVLink/NVL domain.

In the displayed one-hour dynamic windows, OpScale runs a one-second control interval while the model-level baselines scale every 20 seconds. It reports average allocations/SLO attainment of 7.1 GPUs / 98.4% for Qwen2-7B under a one-second P99 TTFT target, versus 11.2/14.3/13.0 GPUs for DynamoLLM/AIBrix/Production Stack, and 10.8 GPUs / 98.1% for Qwen2-57B-A14B under a two-second P99 TTFT target, versus 17.3/23.1/18.0 GPUs. On-demand Qwen2-7B scale-up is 10.68 seconds average at model level, versus 0.03/0.18/0.33 seconds for one operator, 50% of operators, or all operators; the corresponding P99s are 11.55 seconds versus 0.10/0.42/0.45 seconds.

The larger maxima remain experiment-bound and use the model-level configuration as their comparison denominator: 20.1% and 35.7% average GPU reductions at 1K sequence length for the dense and MoE models, 36.3% at 4K and 22.1% at 8K in the disclosed SLO-constrained sweeps, 14-28% cluster-power reduction at high load, 3-38% higher dense-model input TPS on the same GPU budget, and up to 44% higher MoE input TPS at 40 GPUs. No achieved active-batch value, numeric queue depth or wait, exact cluster GPU-utilization value, realized operator-replica count/placement series, configured per-model tensor/pipeline parallelism, numeric TBT target, failures/retries, explicit goodput metric, or currency/GPU-hour cost is disclosed. The separate granularity/topology figure reports request throughput in RPS, but does not state its model, sequence length, arrival/trace input, or SLO. arXiv v1 and its TeX source link no OpScale code, trace, configuration, or result-data repository; linked nano-vLLM and other repositories are dependencies or baselines, not an OpScale artifact. Counts remain 272 entries, 265 unique source URLs, and 224 entity labels.

Latest slice: Google AI Hypercomputer control plane and TPU topology (#9384)

Five official Google Cloud records now separate supported topology, configured maxima, admission semantics, and achieved active work. Cluster Director documents a host→single-rack sub-block→block→cluster GPU network hierarchy, with one hop inside a sub-block, at most two hops inside a block, and support for clusters at the “thousands of GPUs” scale. Current GKE Multislice supports homogeneous multi-host TPU slices and atomic slice-node-pool scaling, while the maintenance-only Cloud TPU API separately documents a 256-slice queued-resource ceiling; that legacy ceiling is not silently transferred to GKE. Dynamic Workload Scheduler distinguishes best-effort Flex-start admission from approved reservation-bound capacity. Finally, Google’s November 2023 XPK/GKE report is retained as an achieved 50,944-chip TPU v5e training workload across 199 pods, not as a schedulable fleet or supported maximum.

The current corpus contains 272 entries, 265 unique source URLs, 224 entity labels, and 3 explicit rumors. None of the five sources supplies a production queue-wait distribution, utilization, provider-wide active scale, failure/retry distribution, power, or total workload cost.

Latest slice: NVIDIA / Hugging Face acquisition-report lifecycle (#9383)

One open rumor record now preserves the August 2026 report lifecycle without promoting it to transaction fact. Business Insider reported NVIDIA–Hugging Face talks above $13B and said no deal had been reached; The Information later reported a $12.9B agreement. Neither company had announced or confirmed signing as of 2026-08-27, and a reported agreement is not a completed acquisition. Signing or party announcement, transaction structure and terms, regulatory conditions, and close remain explicitly unresolved through the 2026-11-30 review date. The dated source and independence audit is in source-register.md.

Latest slice: production distribution evidence (#9381)

Four existing records now carry variable- and denominator-specific audits for BurstGPT/Azure OpenAI, Chutes Year-in-Serving, ServeGen, and Alibaba Cloud Model Studio KV-cache. The audit distinguishes empirical distributions, explicit fitted models, statistical fit/tests, and synthetic generator assumptions. No entries or entities were added. ServeGen v3 supplies bounded, variable-specific fitted families, but public evidence still does not support one universal Zipf, lognormal, Pareto, Poisson, Hawkes, or MMPP workload law.

Status: initial end-to-end spine, incomplete by design. As of: 2026-08-27. Tracker: #9269.

This index records what frontier labs, hyperscalers, AI clouds, datacenter operators, accelerator vendors, serving-system builders, researchers, reporters, and market actors say or measure about the infrastructure and workloads behind frontier AI. It exists to stop fak from optimizing against an invented cluster, a stationary synthetic workload, or an unlabeled market rumor.

The machine-readable source of truth is index.json. The derived views are:

Empty or missing slices are coverage debt, not evidence that a category has no activity.

Latest slice: provider-scale demand denominators (#9379)

This slice adds official first-party scale disclosures for Google Gemini, Meta AI, and Microsoft 365 Copilot, and refreshes the existing OpenAI and Anthropic user-scale records. The five records deliberately preserve unlike denominators: Gemini app monthly active users, Antigravity weekly active users, monthly model developers, model-API token throughput, Google Cloud token-threshold customer cohorts, Meta AI monthly actives, weekly ChatGPT people, paid Microsoft 365 Copilot seats, GitHub Copilot users, Microsoft Foundry customers, Agent 365 registered agents, Purview-audited Copilot interactions, and Claude-on-Amazon-Bedrock customers. They are not interchangeable and do not establish requests, queries, sessions, messages, tokens, concurrency, interarrival laws, geography, market share, or Zipf behavior.

Latest slice: New York and ERCOT large-load controls (#9373)

Four bounded official-source records now cover the PUCT-granted ERCOT Batch Zero deadline extension plus three directly inspected sources: New York Executive Order 62, ERCOT’s June Batch Zero approval release, and ERCOT’s August transmission-planning response. The slice separates New York’s qualified state-application abeyance from local permissions and a blanket construction ban; tracked ERCOT request nameplate from verified or livable demand; and plan assumptions, sensitivities, and forecasts from actual load. It also keeps planned or recommended 765-kV facilities distinct from constructed and energized assets. The New York GEIS/report and 60/90-day and twelve-month processes, ERCOT’s final Batch Zero classification and Fall 2027 statewide plan, and the final end-2026 forecast including Batch loads remain future work. The Texas governor directive/~474 GW claim remains coverage debt because no direct governor source was supplied. The corpus now contains 266 entries, 259 unique URLs, and 220 entity labels.

Latest slice: Chinese platform envelopes (#9362)

Six bounded records add Baichuan 2 training, iFLYTEK ecosystem denominators, Meituan LongCat training/inference, production-scale asynchronous RL, stateful generative- recommendation caching, and one-week recommendation-training sequences. The corpus now contains 213 entries, 208 unique source URLs, and 176 entity labels. Developer teams, applications, developers, API growth, agents, sequences, users, requests, and tokens remain separate denominators; internal/vendor maxima are not universal.

Latest slice: realized failures and cancellations (#9363)

This slice reconciles the existing Untether AI bankruptcy record and adds Builder.ai insolvency, a bare-metal cloud wind-down, semiconductor-project cancellation and delay, one named data-center withdrawal, and a 25-project lower-bound U.S. cancellation cohort. The corpus now contains 218 entries, 213 unique source URLs, and 181 entity labels. Bankruptcy, insolvency, wind-down, impairment, withdrawal, cancellation, delay, operating shutdown, and lost live MW remain distinct states.

Latest slice: named grid and utility outcomes (#9364)

Four records connect large-load policy to service reality: FERC’s rejection of the Susquehanna 300-to-480 MW co-location amendment, AEP Ohio’s collateralized data-center tariff, Dominion’s GS-5 large-load class, and LPSC approval of 2,262 MW plus 500-kV transmission for Meta’s named Louisiana project. The corpus now contains 222 entries, 217 unique source URLs, and 185 entity labels. Tariff, contract, collateral, queue, regulatory approval, construction, energization, actual load, and live IT MW remain distinct.

Latest slice: component shipment receipts (#9365)

Five primary-source records add Micron HBM3E volume production, Micron HBM4 high-volume shipments, Broadcom Tomahawk 6 production-volume shipments, TSMC CoWoS-L production and customer qualification, and TSMC Arizona N4 high-volume wafer production. The corpus now contains 227 entries, 222 unique source URLs, and 190 entity labels. Samples, qualification, volume production, shipped components, assembled systems, deployed clusters, and useful goodput remain separate lifecycle states.

Latest slice: speculative acceptance envelopes (#9366)

Five dedicated records add foundational draft/verify sampling, SpecInfer token trees, Medusa decoding heads, EAGLE feature speculation, and MagicDec long-context batching. The corpus now contains 232 entries, 227 unique source URLs, and 195 entity labels. Drafted, verified, accepted, rejected, fallback, and emitted tokens remain separate work; benchmark speedups are task/model/hardware/load envelopes, not stackable production multipliers.

Latest slice: non-coding agent workloads (#9367)

Five benchmark families add browser, enterprise, customer-service API, general-assistant, and desktop-computer workloads. The corpus now contains 237 entries, 232 unique source URLs, and 200 entity labels. Task counts, generated instances, websites, tools, capability tags, observations, actions, turns, users, sessions, and production arrivals remain separate denominators; benchmark success is not a production workload distribution.

Latest slice: speculative-decoding production observability (#9371)

Three pinned official-repository records add bounded vLLM, SGLang, and NVIDIA TensorRT-LLM speculative-decoding telemetry/control surfaces. The corpus now contains 244 entries, 239 unique source URLs, and 207 entity labels. Available instrumentation is not evidence that it is enabled, collected, retained, or representative: request accepted-draft-length histograms remain distinct from fleet aggregate counters; accepted, drafted, emitted/output tokens, and iterations remain distinct denominators; adaptive-controller inputs remain distinct from published outcomes; and benchmark acceptance or speedup remains distinct from production prevalence or goodput. No verified retry/fallback/client-side production distribution was found in this bounded primary-source slice.

Latest slice: public conversation populations (#9370)

Four direct population records add LMSYS-Chat-1M, WildChat, Chatbot Arena preferences, and OpenAssistant conversation trees. The corpus now contains 241 entries, 236 unique source URLs, and 204 entity labels. Conversations, messages/turns, trees, paths, votes, users or anonymized IDs, languages, countries, timestamps, model appearances, and production requests remain separate denominators; public and crowdsourced datasets are not provider-wide traffic.

Result first

The initial evidence already rejects several convenient defaults:

  1. The datacenter is becoming the system boundary. Public plans and product designs are expressed in racks, pods, sites, multi-site fleets, hundreds of megawatts, and gigawatts—not one eight-GPU server.
  2. Power and delivery time are first-class constraints. Chip supply is not enough; grid connection, firm generation, cooling, permitting, water, construction, and local consent can determine when nominal capacity becomes usable.
  3. Serving fleets are heterogeneous. Frontier providers publicly describe mixes of NVIDIA GPUs, TPUs, Trainium, custom silicon, multiple clouds, regions, and product channels. A receipt that names only the model is operationally incomplete.
  4. Inference is splitting into phases and state tiers. Prefill/decode disaggregation, KV-aware routing, cache offload, multi-tier memory, long-context specialization, and traffic-aware scaling are moving from papers into vendor and open-source platforms.
  5. One stationary request distribution is not credible. Production-trace work finds client-specific, modality-specific, reasoning-specific, bursty, nonstationary arrival and token behavior. “Poisson + fixed input/output length” is a test fixture, not a production claim.
  6. Nameplate capacity is not useful capacity. Plans, contracted gigawatts, installed chips, healthy schedulable accelerators, utilization, cluster goodput, SLO attainment, and accepted-token output are different quantities.
  7. The public record is rich on supply and poor on demand shape. Providers disclose capital, chip counts, regions, and power more often than tenant skew, prefix popularity, token distributions, cache-hit opportunity, concurrency, geography, or retries. Those missing distributions are high-priority unknowns for fak.

Evidence contract

Every entry must include:

Evidence classes are intentionally not interchangeable:

Class What it can prove
production_measurement Behavior measured on a production population, within its disclosed sample limits.
production_observation Operational behavior disclosed without a complete measurement dataset.
benchmark_measurement Performance inside a stated experimental envelope.
synthetic_experiment Behavior under a generated model; never proof of production prevalence.
official_statement What an organization says it did, plans, or contracted.
vendor_claim A product or performance assertion needing matched independent validation.
analyst_estimate / reported_estimate A third party’s bounded estimate, with methodology risk.
reported_observation Credible reporting of events or constraints, not a controlled measurement.
inference A conclusion drawn from cited evidence and labeled as such.
rumor Unconfirmed information with origin, corroboration, and confidence recorded.

A plan is not delivered capacity. Peak FLOPS are not goodput. Registered developers are not active users. A benchmark histogram is not a production distribution. A repeated rumor is not independently corroborated merely because many sites copied one origin.

Taxonomy

Actors

Expectations and assumptions

What fak should assume now

Until stronger evidence exists, use these as conservative design defaults—not universal facts:

Coverage ledger

The current spine contains 272 dated entries, 265 unique source URLs, and 224 distinct entity labels across 13 categories spanning frontier labs, hyperscalers, AI clouds, datacenter supply, accelerators, serving systems, workload traces, market signals, and 3 explicit rumors. FineServe, ServeGen, the one-year Chutes trace, OpenRouter geography, and SkyLB/SkyWalker now provide source-bounded production parameters. Azure OpenAI / BurstGPT adds 10.31M requests over 213 days with daily/weekly periodicity, token tails, separated failures, and multi-duration burst examples; Splitwise adds one-day Conversation and Coding traces; Huawei Atlas 900 A3 adds a vendor reference topology rather than a deployment receipt. Confidence intervals, geography, retries, speculative acceptance, session/tool-call distributions, and comparable denominators remain explicit gaps. It is broad, but it is not entity-complete. The requirement-level verdict is in coverage-audit.md; the detailed missing slices remain machine-readable under coverage.explicit_gaps in index.json.

Immediate next slices:

  1. finish the checked/unchecked slices/frontier-lab-census.md, including Chinese and regional labs;
  2. extend workload-parameters.md and geography-session-locality.md with direct provider tenant/timezone/session distributions;
  3. extend the filings-ledger.md beyond the four largest U.S. hyperscalers;
  4. extend the supply-chain-ledger.md into a named site/vendor delivery census;
  5. expand the market-chronology.md with startup failures, cancellations, and resolved rumors;
  6. add committed schema/link validation and scheduled refresh after the ledger schema stabilizes.

The clearest workload-distribution evidence so far is not a universal law. It is a set of category-dependent, heavy-tailed, bursty, multimodal, nonstationary behaviors: model popularity and user-model affinity evolve over months; KV-prefix reuse is skewed but differs between consumer and API traffic; coding-agent tool calls are heavily tailed; output lengths are strongly right-skewed even for identical prompts; request-pod cache affinity can be roughly bimodal; daily and weekly periodicity coexists with short and sustained bursts; and traffic changes abruptly after model releases. When a benchmark uses a Zipf, Poisson, lognormal, Pareto, Hawkes, or MMPP assumption, it must name the fitted population or label the distribution as synthetic.

Bounded remote-browser execution-envelope slice (#9375)

The pinned official source is Steel BrowserBench commit 847e4ed604764ce8b887265709ddb2d1c3c5f020, authored 2026-01-15. Its README, runner, package lock, and committed result files describe sequential lifecycle samples from an AWS EC2 client in us-east-1, navigating first to https://google.com/, with 10 warm-ups per provider excluded. The five indexed records retain their distinct provider SDK and test windows: Steel (steel-sdk 0.14.0, 2025-11-06), Kernel (@onkernel/sdk 0.18.0, 2026-01-10 to 2026-01-11), Browserbase (@browserbasehq/sdk 2.6.0, 2026-01-12 to 2026-01-13), Hyperbrowser (@hyperbrowser/sdk 0.71.0, 2025-11-05 to 2025-11-06), and AnchorBrowser (anchorbrowser 0.8.3, 2025-11-05 to 2025-11-06). Each committed file has 5,000 attempts. Steel, Kernel, Browserbase, and Hyperbrowser report 5,000 successful samples; AnchorBrowser reports 4,867 successful samples and 133 session_create failures (97.34%).

The record boundary is deliberately narrow:

The browser/desktop production gap remains explicit: no reviewed primary source supplies action/tool-call distributions, observation bytes, reset distributions, full session duration, escalations, user populations, or production retry/failure/tail distributions.

Validation

python3 -m json.tool docs/research/frontier-infrastructure/index.json >/dev/null
python3 - <<'PY'
import json
from pathlib import Path
p = Path('docs/research/frontier-infrastructure/index.json')
d = json.loads(p.read_text())
required = set(d['required_entry_fields'])
ids = set()
for i, row in enumerate(d['entries']):
    missing = required - row.keys()
    assert not missing, (i, sorted(missing))
    assert row['id'] not in ids, row['id']
    ids.add(row['id'])
    assert row['source_url'].startswith('https://')
    assert isinstance(row['evidence_class'], str) and row['evidence_class'].strip()
    assert isinstance(row['confidence'], str) and row['confidence'].strip()
assert d['coverage']['entry_count'] == len(d['entries'])
assert d['coverage']['entities'] == sorted({row['entity'] for row in d['entries']})
print(len(d['entries']), 'valid entries')
PY

Bounded remote-browser operational-envelope slice (#9376)

Issue #9376 extends the prior lifecycle slice with official operational boundaries for Steel, Browserbase, Hyperbrowser, Kernel, and Anchor Browser. The five added browserops-* records separate:

Direct quantities now include Steel’s plan table; Browserbase’s separate plan limits for active concurrency (Free 3, Developer 25, Startup 100, Scale 250+) and session creations per minute (Free 5, Developer 25, Startup 50, Scale 150+), with HTTP 429 for either excess and over-limit creation described as dropped; Browserbase’s configurable project-default session duration, six-hour maximum duration, and separate 10-minute CDP connection-inactivity timeout; Kernel’s standby timeout and pool long-poll result; and Anchor’s independent idle/hard timers. Hyperbrowser contributes only a configuration/lifecycle example: a per-session timeout override whose code sample sets 60 minutes. The reviewed Hyperbrowser guide provides no numeric concurrency, queue, rate, default-timeout, or maximum-timeout boundary. All sources were accessed 2026-08-27; Steel additionally identifies a 2026-06-30 last edit.

Use workload-assumptions.md for benchmark admission rules and coverage-audit.md for provider-by-boundary coverage and unresolved gaps.

Current serving-configuration slice

Issue #9382 adds five commit-pinned MLPerf configurations/results with explicit topology and batching omissions. The compact comparison is in workload-assumptions.md, and the usable parameter rules are in workload-parameters.md. These records are benchmark fixtures only: they do not describe production deployments, schedulable fleets, replica counts, achieved active batches, or prevalence.