Short answer
The FirstCoreAI AI Factory is hosted in FirstNet's data centre in South Africa. FirstCoreAI, FirstNet's AI business unit, owns and operates the GPU cluster, and inference on the AI Factory happens inside South Africa.
In detail
The AI Factory is physical compute, tuned for production throughput and concurrency rather than benchmark leaderboards, and it runs on networks FirstNet operates. Network tuning, rack scaling and cross-connects are handled by one team.
What is inside:
- Current-generation NVIDIA Blackwell-class GPUs with AMD EPYC processors and more than a terabyte of system RAM per node.
- A multi-terabyte NVMe model cache shared over NFS, a high-throughput in-band fabric between nodes and segregated out-of-band management.
- Kubernetes with Run.ai for GPU-aware scheduling, vLLM for inference, HAProxy and a Knative gateway at the edge, and Grafana and Prometheus for monitoring.
- Six live production model endpoints exposed through an OpenAI-compatible API.
Because processing stays local, workloads on the AI Factory involve no cross-border transfer of data. FirstCoreAI runs 60 to 90 minute tours covering the hardware, the model catalogue and a live workload, booked directly with FirstCoreAI.
Source: FirstNet GPU & Inference service page →
Didn’t answer your question?
