Signal & Supply
โ† Archive
September 7, 2026 PHOTONICS

An AI GPU Cluster Talks to Itself in Light. The Chip That Makes the Light Is Sold Out Through 2027.

Every GPU in a modern AI cluster is wired to its neighbors with pulses of laser light, not copper. The compound-semiconductor material that generates that light, indium phosphide, is now scarcer than the memory chips everyone panicked about a year ago.

Key takeaway Optical-transceiver demand is running roughly 30% ahead of what indium phosphide (InP) laser fabs can supply, and lead times on the fastest modules have stretched past 40 weeks. Lumentum's CEO said in July the shortage would be worse than the HBM memory crunch. Nvidia has already spent $4 billion locking up two of the world's few InP laser suppliers โ€” a bet that whoever secures capacity first sets the pace for everyone else's AI buildout through at least 2027.

The fabric multiplier

Drag the slider to scale an AI training cluster. Top row: the network layers a fabric needs at that scale โ€” each extra layer (Leaf โ†’ Spine โ†’ Core โ†’ Superspine) is another full stage of switch-to-switch optical links stacked on top of the GPU links. Bottom bars: GPU count, held as a fixed 1ร— reference, versus optical transceivers needed per GPU at that scale. Simplified, illustrative model: real topologies and multipliers vary by vendor, but the core relationship โ€” links growing faster than GPU count as clusters add layers โ€” is a real, widely reported property of AI network fabrics. "Laser chips needed" is treated as 1-per-transceiver here for simplicity; real high-speed modules often use several lasers per transceiver, so actual chip demand runs higher than shown.

NETWORK LAYERS AT THIS SCALE GPU LEAF SPINE CORE SUPER‑SPINE GPU COUNT (reference) 1× OPTICAL TRANSCEIVERS NEEDED, PER GPU 6× 0× 4× 8× optical transceivers needed per GPU, indexed to a 0–8× scale

Scale the cluster:

131,072 GPUs

A large training cluster of 131,072 GPUs needs roughly 786,432 optical transceivers โ€” about 1.25% of all 800G+ transceivers expected to ship worldwide in 2026.

GPUs in cluster
131,072
Network layers active
4
Transceivers needed
786,432
Share of 2026 global supply
1.25%

The plain version

Think of a giant AI data center as a stadium full of GPUs that all need to talk to each other constantly, trading enormous amounts of data as they train one model together. Ordinary copper wires can't carry a signal far or fast enough at that scale, so engineers switched to fiber optics โ€” the same technology that carries the internet across oceans. Every optical cable needs a small device called a transceiver, and every transceiver needs a laser to actually generate the light. Those lasers aren't made from ordinary silicon; they're grown from a rarer, trickier material called indium phosphide, on wafers a fraction of the size of a normal chip wafer, in a handful of highly specialized factories.

For years, indium phosphide lasers were a modest business, sold by the thousands to phone and internet companies. Then AI happened. Hyperscalers now need them by the millions, partly because every extra layer of switches added to scale a cluster up multiplies how many optical links โ€” and therefore how many lasers โ€” the whole system needs. Demand for these transceivers is running about 30% ahead of what suppliers can produce, and delivery times have stretched past 40 weeks for the fastest modules, four to five times longer than ordinary networking gear. Lumentum's CEO, whose company runs five of these specialty wafer fabs, said in July the shortage would end up worse than the memory-chip crunch that doubled RAM prices last year.

Nvidia isn't waiting around: in March it committed $4 billion to two laser suppliers, Lumentum and Coherent, buying up years of future capacity before its rivals could get in line. The long-term fix โ€” soldering the lasers directly onto the switch chip instead of onto separate cable ends, called co-packaged optics โ€” won't ship at real scale until 2027 at the earliest. Until then, the thing throttling how fast AI factories can grow isn't the GPU. It's the light that connects them.

The expert version

Modern AI training clusters connect thousands to hundreds of thousands of GPUs over non-blocking, multi-tier network fabrics โ€” commonly leaf-spine at modest scale, extending to leaf-spine-core-superspine at frontier scale. Because copper's reach is limited to roughly 2-3 meters at 800Gb/s-class signaling, nearly every inter-rack and scale-out link runs over pluggable optical transceivers (currently 800G, migrating to 1.6T) built around PAM4-modulated lasers. Each additional switching tier a cluster crosses adds a full stage of switch-to-switch optical links on top of GPU-to-leaf links, so transceiver โ€” and laser-chip โ€” demand scales faster than GPU count as clusters grow, not linearly with it.

The laser die itself is grown from indium phosphide (InP), a III-V compound semiconductor, on wafers typically 4 or 6 inches in diameter โ€” far smaller and lower-yielding than silicon's 12-inch standard, and concentrated in a handful of specialty fabs (Lumentum alone runs five). Global shipments of 800G+ transceivers are projected to jump roughly 2.6x year over year in 2026, toward the tens of millions of units, but InP laser output isn't keeping pace: industry estimates put transceiver demand running about 30% ahead of supply, with 800G/1.6T lead times exceeding 40 weeks versus 8-14 weeks for mature 100G parts. Lumentum CEO Michael Hurlston said in July, at the RAISE Summit in Paris, that the InP supply-demand gap has now overtaken DRAM's, calling it a bigger risk than the HBM memory shortage.

Nvidia responded by directly securing upstream capacity: in March it invested $2 billion each into Lumentum and Coherent โ€” a combined $4 billion โ€” paired with multi-year, multibillion-dollar purchase commitments and future capacity rights, effectively pre-buying laser output as both build new U.S. fabs. The structural fix is co-packaged optics (CPO): integrating silicon-photonic engines directly into the switch package (Nvidia's Quantum-X/Spectrum-X, Broadcom's 51.2Tbps CPO switch) to cut the number of discrete lasers and connectors per bit moved. Broadcom's CPO switch isn't shipping in volume until late 2027, and analysts expect the InP supply-demand imbalance to persist into 2027-2028.