NORTH-SOUTH WAN SIZING
AI Data Centre Bandwidth Calculator
North-south traffic at an AI site is a function of installed IT load, the size of the largest model, and how often checkpoints drain. Everything below is normalised per megawatt. Five independent flows are modelled; average and peak are dominated by different flows.
Indicative model · published assumptions, not a measurement or a capacity commitment. Route-specific design available on request.
Three things worth knowing before you slide
Megawatts, not GPUs.
Traffic doesn’t scale with GPU count — it scales with megawatts. Card generations turn over every eighteen months and their count per watt shifts each generation; traffic per watt barely moves. Sizing by GPU dates fast.
Peak and average are driven by different things.
Average is inference. Peak is almost entirely checkpoint drains. Mix them and you’re wrong in both directions at once.
The Φ coefficient.
Independent tenants don’t checkpoint on the same second, so their peaks don’t add up. Three tenants → Φ ≈ 0.70; eight tenants → Φ ≈ 0.40. The practical takeaway: contractually staggering checkpoint windows is the cheapest way to buy yourself capacity.
Total IT power at the site, not the utility feed.
Trillions of parameters. Drives checkpoint volume: bytes per parameter × retention factor.
Seconds to drain one checkpoint off-site. Shorter drains raise peak; longer drains average out.
Minutes between checkpoint events per burster. Sets the duty cycle for average bandwidth.
Concurrent training runs. Independent means their drain windows are uncorrelated — that is what lets Φ come down from 1.
Average
0.39 Tb/s
4 Gb/s per MW
Peak after phi
3.1 Tb/s
31 Gb/s per MW
Required link at 65% utilisation
4.8 Tb/s
Peak divided by target utilisation.
Lambdas across two diverse routes
26 ×400G
or 14 ×800G
Peak flow breakdown
Share of peak Tb/s per flow.
- Checkpoint drain: 2.0 T at peak, 63% of total.
- Dataset ingest: 360 G at peak, 11% of total.
- Inference I/O: 473 G at peak, 15% of total.
- Distill · copy-out: 300 G at peak, 10% of total.
- Ops · control plane: 40 G at peak, 1% of total.
Methodology notes
Why megawatts and not GPUs
GPU generations change count per watt roughly every 18 months. Hopper to Blackwell to whatever comes next moves the FLOPs and the memory per watt, but north-south traffic per watt barely moves — checkpoints, dataset ingest and inference all scale with power delivered to compute, not with the accelerator SKU. Sizing on GPU count ages badly the moment a new generation lands on the floor.
Why Φ exists
Independent tenants' checkpoint drain windows are uncorrelated, so their peaks do not sum — the site sees the convolution, not the addition. Φ is the effective correlation factor across N bursters, defaulted here to 1.00 at N=1, 0.75 at N=2, and dropping 0.05 per additional burster to a floor of 0.4. Contractually staggering drain windows in the tenant SLA is the cheapest capacity upgrade available: with three or more bursters it drives Φ toward 0.4, i.e. more than halves peak checkpoint bandwidth without touching physical infrastructure.
Model limits
The model assumes independent bursters. A shared cluster scheduler correlates them and Φ returns to 1. Cross-site distributed training is not included — it is a separate 10–100 Tb/s term for parallelism between data centres and belongs in a route-level design conversation, not a per-site sizing. Inference coefficients are calibrated for text plus light retrieval-augmented generation; media generation (image, and especially video) breaks the scale by an order of magnitude and should be sized directly against measured egress from the serving stack.
We sell the lambdas this calculator counts.
For a route-specific design — diverse paths, cross-connects, latency and delivery terms — open the configurator or talk to the network team. Both take the same brief.
The configurator is not pre-filled from this page — the model output is indicative and the configurator quotes real capacity. Bring the number over deliberately.