Short answer
Six production model endpoints are live on the FirstCoreAI AI Factory, which InfraAI delivers as a service: an 80B mixture-of-experts coding model, a 31B general chat model, vision-language, speech-to-text, multilingual embeddings and a retrieval reranker. All are open-weight models hosted in FirstNet's South African data centre.
In detail
The catalogue at a glance:
- Coding: an 80B mixture-of-experts model with a 131K context window.
- General chat: a 31B model with a 128K context window.
- Vision-language: for tasks that combine images and text.
- Speech-to-text: transcription through the audio transcriptions endpoint.
- Multilingual embeddings and a retrieval reranker: the building blocks for search and RAG.
Every model is exposed through an OpenAI-compatible API and is benchmarked on the same hardware against frontier APIs for accuracy, latency, concurrency and cost per unit of work. On the production cluster, FirstCoreAI measured a median time to first token of about 0.6 seconds on the chat and coding models.
Where a workload genuinely needs an overseas frontier model, requests can be routed to a frontier API, and FirstCoreAI gives notice when a request leaves the country.
Source: FirstNet GPU & Inference service page →
Didn’t answer your question?
