
Ask FirstNet about Sovereign AI
GPU & Inference as a Service
InfraAI delivers FirstCoreAI's AI Factory as a service: private inference endpoints, dedicated or shared GPU capacity and a curated set of open-weight models, all running in South Africa and reachable through familiar OpenAI-style APIs.
What's included
The InfraAI service
- Private inference endpoints hosted locally; nothing leaves the country
- Six live production model endpoints on the AI Factory
- Dedicated or shared GPU options sized to your workload
- Optional frontier routing, with notice when a request goes offshore
- Rand-denominated, forecastable commercials with no FX surprises
Quick answers
Where is the AI Factory hosted?
In FirstNet's data centre in South Africa. FirstCoreAI owns and operates the GPU cluster, and inference happens inside South Africa.
Do I need to rewrite my application to use InfraAI?
No. Every model is exposed as an OpenAI-compatible API, so any SDK that already works with OpenAI works with InfraAI. You only swap in the endpoint URLs and credentials issued at onboarding.
Which models are available?
Six production endpoints are live: an 80B mixture-of-experts coding model, a 31B general chat model, vision-language, speech-to-text, multilingual embeddings and a retrieval reranker. Frontier APIs can be routed to where a workload needs one.
How is access secured?
Clients authenticate with OAuth2 client credentials and receive short-lived bearer tokens. TLS is terminated at the cluster edge and OAuth2 proxies validate each request.
Inside the AI Factory
Hardware and models
The AI Factory is physical compute that FirstCoreAI operates in FirstNet's South African data centre, tuned for production throughput and concurrency rather than benchmark leaderboards.
Compute
Current-generation NVIDIA Blackwell-class GPUs with AMD EPYC processors and more than a terabyte of system RAM per node.
Models
An 80B mixture-of-experts coding model (131K context), a 31B chat model (128K context), vision-language, speech-to-text, multilingual embeddings and a retrieval reranker.
Storage and fabric
A multi-terabyte NVMe model cache shared over NFS, a high-throughput in-band fabric between nodes and segregated out-of-band management.
Software stack
Kubernetes, Run.ai for GPU-aware scheduling, vLLM for inference, HAProxy and a Knative gateway (the open-source Kubernetes serving layer) at the edge, plus Grafana and Prometheus.
How you connect
Drop-in OpenAI compatibility
Your application authenticates with OAuth2 client credentials, receives a short-lived bearer token and calls standard endpoints. TLS terminates at the cluster edge and each request is validated. Existing OpenAI SDKs work without code changes; your token endpoint and URLs are issued at onboarding.
- /v1/chat/completions for chat and coding
- /v1/embeddings for search and retrieval
- /v1/audio/transcriptions for speech-to-text
- /v1/rerank for retrieval reranking
Explore the details
Performance on the production clusterFirstCoreAI stress-tested the live cluster under real serving conditions. Every model is also benchmarked…
Measured, not promised
Performance on the production cluster
FirstCoreAI stress-tested the live cluster under real serving conditions. Every model is also benchmarked on the same hardware against frontier APIs for accuracy, latency, concurrency and cost per unit of work, so you see the evidence before committing.
~0.6s
Median time to first token on chat and coding models.
271
Concurrent requests sustained at peak during the full stress run.
0
Errors across all stress tests, at every concurrency level.
Is InfraAI the right fit?InfraAI suits high-volume document and call processing, RAG and retrieval back-ends, internal assistants…
Best for
Is InfraAI the right fit?
InfraAI suits high-volume document and call processing, RAG and retrieval back-ends, internal assistants, classification and extraction, and any workload where data sensitivity or cost rules out a foreign API. To see your workload on the Factory, book a demo or tour with FirstCoreAI.
- Dedicated capacity instead of a shared public queue
- Every inference stays inside the South African perimeter
- Support from the engineers who run the Factory
Can I get dedicated GPU capacity?
Yes. InfraAI offers both dedicated and shared GPU options, sized to the workload, with rand-denominated billing.
Can I tour the AI Factory?
Yes. FirstCoreAI runs 60 to 90 minute tours covering the hardware, the model catalogue and a live workload. Book through FirstCoreAI.
From the Knowledge Hub
In-depth answers about GPU & Inference
- Can I tour FirstNet's FirstCoreAI AI Factory GPU cluster?
- Can I get dedicated GPU capacity in South Africa from FirstNet's InfraAI service?
- How is access to FirstNet's InfraAI inference endpoints on the FirstCoreAI AI Factory secured?
- Which AI models are available on FirstNet's InfraAI service and FirstCoreAI AI Factory?
- Do I need to rewrite my application to use FirstNet's InfraAI GPU and inference service?
- Where is FirstNet's FirstCoreAI AI Factory hosted, and does inference stay in South Africa?
See it running on home soil
Your place in the stack
Sovereign AI is layer 5 of 5.
The intelligence layer sits on top of everything FirstNet runs, and can work with data from every layer below it. See the full stack →
Bring us a problem.
We’ll show you the stack that solves it.
Ask FirstNet for an instant answer, or leave your details in the chat and the right specialist will contact you.
