change to 16 9 aspect ratio 202605270818 scaled

Ask FirstNet about Sovereign AI

GPU & Inference as a Service

InfraAI delivers FirstCoreAI's AI Factory as a service: private inference endpoints, dedicated or shared GPU capacity and a curated set of open-weight models, all running in South Africa and reachable through familiar OpenAI-style APIs.

What's included

The InfraAI service

  • Private inference endpoints hosted locally; nothing leaves the country
  • Six live production model endpoints on the AI Factory
  • Dedicated or shared GPU options sized to your workload
  • Optional frontier routing, with notice when a request goes offshore
  • Rand-denominated, forecastable commercials with no FX surprises

Quick answers

Where is the AI Factory hosted?

In FirstNet's data centre in South Africa. FirstCoreAI owns and operates the GPU cluster, and inference happens inside South Africa.

Do I need to rewrite my application to use InfraAI?

No. Every model is exposed as an OpenAI-compatible API, so any SDK that already works with OpenAI works with InfraAI. You only swap in the endpoint URLs and credentials issued at onboarding.

Which models are available?

Six production endpoints are live: an 80B mixture-of-experts coding model, a 31B general chat model, vision-language, speech-to-text, multilingual embeddings and a retrieval reranker. Frontier APIs can be routed to where a workload needs one.

How is access secured?

Clients authenticate with OAuth2 client credentials and receive short-lived bearer tokens. TLS is terminated at the cluster edge and OAuth2 proxies validate each request.

Inside the AI Factory

Hardware and models

The AI Factory is physical compute that FirstCoreAI operates in FirstNet's South African data centre, tuned for production throughput and concurrency rather than benchmark leaderboards.

Compute

Current-generation NVIDIA Blackwell-class GPUs with AMD EPYC processors and more than a terabyte of system RAM per node.

Models

An 80B mixture-of-experts coding model (131K context), a 31B chat model (128K context), vision-language, speech-to-text, multilingual embeddings and a retrieval reranker.

Storage and fabric

A multi-terabyte NVMe model cache shared over NFS, a high-throughput in-band fabric between nodes and segregated out-of-band management.

Software stack

Kubernetes, Run.ai for GPU-aware scheduling, vLLM for inference, HAProxy and a Knative gateway (the open-source Kubernetes serving layer) at the edge, plus Grafana and Prometheus.

How you connect

Drop-in OpenAI compatibility

Your application authenticates with OAuth2 client credentials, receives a short-lived bearer token and calls standard endpoints. TLS terminates at the cluster edge and each request is validated. Existing OpenAI SDKs work without code changes; your token endpoint and URLs are issued at onboarding.

  • /v1/chat/completions for chat and coding
  • /v1/embeddings for search and retrieval
  • /v1/audio/transcriptions for speech-to-text
  • /v1/rerank for retrieval reranking

Explore the details

Performance on the production clusterFirstCoreAI stress-tested the live cluster under real serving conditions. Every model is also benchmarked…

Measured, not promised

Performance on the production cluster

FirstCoreAI stress-tested the live cluster under real serving conditions. Every model is also benchmarked on the same hardware against frontier APIs for accuracy, latency, concurrency and cost per unit of work, so you see the evidence before committing.

~0.6s

Median time to first token on chat and coding models.

271

Concurrent requests sustained at peak during the full stress run.

0

Errors across all stress tests, at every concurrency level.

Is InfraAI the right fit?InfraAI suits high-volume document and call processing, RAG and retrieval back-ends, internal assistants…

Best for

Is InfraAI the right fit?

InfraAI suits high-volume document and call processing, RAG and retrieval back-ends, internal assistants, classification and extraction, and any workload where data sensitivity or cost rules out a foreign API. To see your workload on the Factory, book a demo or tour with FirstCoreAI.

  • Dedicated capacity instead of a shared public queue
  • Every inference stays inside the South African perimeter
  • Support from the engineers who run the Factory

FAQs

Questions,
answered.

Straight answers from the FirstNet team.

More in the Knowledge Hub →
Can I get dedicated GPU capacity?

Yes. InfraAI offers both dedicated and shared GPU options, sized to the workload, with rand-denominated billing.

Can I tour the AI Factory?

Yes. FirstCoreAI runs 60 to 90 minute tours covering the hardware, the model catalogue and a live workload. Book through FirstCoreAI.

See it running on home soil

Your place in the stack

Sovereign AI is layer 5 of 5.

The intelligence layer sits on top of everything FirstNet runs, and can work with data from every layer below it. See the full stack →

  1. Sovereign AI
  2. Voice
  3. Security
  4. Cloud
  5. Connectivity

Bring us a problem.
We’ll show you the stack that solves it.

Ask FirstNet for an instant answer, or leave your details in the chat and the right specialist will contact you.

Call