Chipmaker Ships Model-Locked Silicon
A new inference wafer hard-wires one foundation model into transistors, and operators must file a civil-style petition before any other weights may occupy the same rack slot.
TORONTO — A chipmaker this week demonstrated inference silicon that permanently hard-wires a single foundation model’s weights into the die, eliminating memory fetches by treating the model as factory-installed furniture rather than software.
In a briefing for partners, engineers said the part cannot load alternate checkpoints without a physical board replacement. “Flexibility was the bottleneck,” a product manager said. “We removed it.”
Under accompanying guidance circulated to cloud buyers, swapping to a different model family requires a notarized Hardware Divorce filing with the vendor’s registry, naming the retiring weights, the successor model, and a cooling-off period during which the old die must remain in a sealed bin.
“You do not reflash a marriage,” the manager clarified, when asked about over-the-air updates. “You file, you wait, you rack a new spouse.”
Analysts said the approach targets memory-bound decoding costs that dominate large-model serving. Early adopters at two unnamed hyperscalers were said to be piloting paired racks: general GPUs for prompt prefill, locked dies for token generation.
A quiet comment window on the divorce form’s fee schedule runs through quarter-end. Until then, operators are advised to label cages with the model name stamped at the foundry, in case auditors ask which personality lives in the metal.