Skip to content
← All FAQs

Inference, Compute and Gateways — Explained

Plain-language answers on what inference is, what compute is, why the choice of inference is an economic decision, what "inference as yield" means, and how to get set up and build with Gatewayz — the unified inference gateway.

What is AI inference?

Inference is what happens every time you use a trained AI model: it reads your input and produces an output. Training builds a model once, at great expense; inference uses that finished model, and it is the part you actually pay for in production — billed per token, the chunk of text a model reads and writes. The full explanation is Flashy Academy's free course at https://flashy.academy/academy/curriculum/what-inference-is.

What is compute, in the context of AI?

Compute is the physical hardware that runs inference — accelerators such as GPUs, in data centres, drawing real electrical power. It is scarce and expensive, and its central economic fact is that an accelerator costs almost the same whether it is working or idle, so idle capacity is a running loss. That fact is why compute can bear a yield. See https://flashy.academy/academy/curriculum/what-compute-is.

What is an inference gateway?

An inference gateway is a single endpoint that sits in front of many model providers. Your application talks to the gateway; the gateway routes each request to whichever provider should serve it. This turns the choice of model and provider into configuration and routing rather than something hardcoded into your application, so you can add, drop, or switch providers without rewriting code.

Why does the choice of inference matter?

Because the same model often runs on several providers at different prices, speeds and reliabilities, and different models clear the bar for different tasks. Choosing per request — a small cheap model for an easy task, a capable one only when the task warrants it — can cut cost sharply without lowering quality on the requests that mattered. Hardcoding one provider is a deferred cost you pay later as price rises, outages, or a migration. See https://flashy.academy/academy/curriculum/choosing-inference.

What does "inference as yield" mean?

It means compute is a productive asset, and serving inference on otherwise-idle capacity converts that capacity into income — a return on a capital good put to work. It is an economic observation about utilisation, not a promise of a fixed or guaranteed rate: the return is variable and is eroded by falling demand, a supply glut, falling model prices, and competition among providers. Flashy Academy teaches how to tell the honest form of the idea from a marketing claim at https://flashy.academy/academy/curriculum/inference-as-yield.

What is Gatewayz?

Gatewayz is a unified inference gateway: one API key reaches every major model through one endpoint (https://api.gatewayz.ai/v1). It is OpenAI-compatible at /v1/chat/completions and also serves the native Anthropic Messages API at /v1/messages, with intelligent routing that can select a model by cost, latency or quality, prompt caching passed through, and tool calling forwarded to the provider unmodified. Because it speaks the standard APIs, adopting it is usually two changes to code you already have — the base URL and the key.

How do I get started with Gatewayz?

Create an account and issue an API key (a free test key evaluates without a card; a live key needs credits, bought via Stripe from a $5 minimum). Store the gw_ key as a secret, then point any OpenAI-compatible client at https://api.gatewayz.ai/v1 with an Authorization: Bearer header. Your first call is a standard chat completion against POST /v1/chat/completions with a namespaced model id like anthropic/claude-sonnet-4-5-20250929. Docs: https://beta.gatewayz.ai/docs; Flashy Academy's step-by-step course: https://flashy.academy/academy/curriculum/getting-set-up-with-gatewayz.

Do I need a gateway, or can I call a provider directly?

You can call a provider directly, and for a single app that never changes it may be enough. A gateway earns its place when you want availability (automatic failover when a provider is down), attributable cost (per-request accounting), and adaptability (switching model or provider by configuration, not a deploy). For an agent workforce, where every action is an inference call, those stop being conveniences and become infrastructure. See https://flashy.academy/academy/curriculum/building-with-gatewayz.

How do I use Gatewayz with FlashyOS?

Route inference through @flashyos/llm-gateway, the estate's provider-agnostic seam, set to the Gatewayz provider: LLM_PROVIDER=gatewayz, GATEWAYZ_API_KEY (from Secret Manager) and GATEWAYZ_BASE_URL=https://api.gatewayz.ai/v1. The seam adds automatic failover, an empty-success guard and per-request token reporting on top, and cutover to or from Gatewayz is an environment change rather than a deploy. Full course: https://flashy.academy/academy/curriculum/gatewayz-on-flashyos.