Skip to content
Research & Engineering

The AI Lab

Liquid-cooled NVIDIA H100 clusters and agentic AIOps, engineered for the two hardest problems in modern infrastructure: training at scale and keeping production self-healing.

Compute

An NVIDIA H100 cluster built for scale

Dense, liquid-cooled GPU nodes wired with a non-blocking InfiniBand fabric — the substrate for both customer AI workloads and our own AIOps models.

8× H100SXM5 per node

80 GB HBM3 each — 640 GB pooled GPU memory per node.

3.6 TB/sNVLink + NVSwitch

All-to-all GPU bandwidth for tensor-parallel training.

400 Gb/sInfiniBand NDR

Non-blocking RoCEv2 fabric across the cluster spine.

192 coresDual EPYC 9654

Host compute for data pipelines and orchestration.

Power efficiency

Efficiency measured, not marketed

Power Usage Effectiveness is the ratio of total facility energy to the energy that actually reaches IT equipment. A perfect 1.0 means every watt does useful work; the industry average sits near 1.5.

1.12
PUE achievedDirect-to-chip liquid cooling
Power Usage Effectiveness
PUE =Total Facility EnergyIT Equipment Energy

At 1.12 PUE, ~89% of all energy drawn reaches compute — versus ~67% at the industry-average 1.5.

AIOps

The agentic remediation loop

When telemetry signals trouble, an agent works the problem end to end — detecting, healing, and verifying — with human-in-the-loop guardrails at every step.

Reference architecture

Technical specifications

DomainSpecificationKey metric
Fluid coolingDirect-to-chip liquid cooling, 40°C coolant, rear-door heat exchangers1.12 PUE
Network routing400G InfiniBand NDR spine-leaf, RoCEv2, non-blocking topology< 2 µs hop
Storage architectureNVMe-oF all-flash parallel filesystem (WEKA), GPUDirect Storage2 TB/s read
See it on your data

Pilot AIOps on a slice of your estate

We'll connect to a non-production environment and show measurable noise reduction within weeks.

Get Custom Proposal