AI data center security has a timing problem: the racks are going up faster than anyone can audit what's inside them. That's the warning in a recent SecurityWeek piece, "AI Data Centers Are Being Built Faster Than They Can Be Secured," by Kevin Townsend, covering new research from into the top ten AI infrastructure security risks — a framework the researchers call "Forge," built around hardening the metal beneath the model.
Here's the part that should worry every CISO signing off on a new GPU cluster: the five most severe Forge risks — firmware and hardware integrity compromise, network and interconnect vulnerabilities, unsafe multi-tenant isolation, insecure out-of-band management planes, and AI supply chain compromise — all sit below the operating system. They're hard to detect with conventional tools, and when something goes wrong, the blast radius is cluster-wide, not host-wide.
Why does this keep happening? Because AI data centers aren't traditional data centers with more horsepower. Traditional facilities are data-processing warehouses for a known set of tenants. AI facilities are compute factories running dense GPU clusters as a single parallel engine, often serving unrelated, high-value multi-tenant workloads over high-performance fabrics — InfiniBand, RoCE, RDMA, NVLink — that are frequently unencrypted and under-monitored. Layer on heavy reliance on BMC automation, Redfish, and IPMI for remote hardware control, and you get a concentration of privilege sitting in the out-of-band management plane that most security programs were never built to watch.
In other words: firmware integrity, hardware health, and out-of-band management are becoming the biggest blind spot in AI infrastructure — and static, point-in-time audits can't keep pace with builds moving this fast. What's needed is continuous, predictive visibility into that layer, not an annual assessment that's stale before the ink dries.
Where the Forge framework flags "firmware and hardware integrity compromise" and "insecure out-of-band management plane" as the top-severity, below-the-OS risks, that's precisely the class of anomaly a 24x7 SOC is positioned to catch: unauthorized configuration changes, unexpected privilege escalation through management interfaces, and behavioral deviations that a quarterly scan would miss entirely. Reducing dwell time matters more, not less, in AI infrastructure — when the blast radius of a compromise is cluster-wide rather than host-wide, the difference between catching an anomaly in minutes versus discovering it at the next audit cycle is the difference between an incident and a headline.
If your AI data center build is scaling faster than your team's ability to monitor firmware integrity, hardware health, and out-of-band management continuously, that's the exposure gap worth closing first — with always-on detection, not an annual checklist.