Public alpha · api.tezgrid.com
Inference on machines that already exist.
TezGrid runs AI models on computers people already own, like gaming PCs, Macs and office servers. If your app uses OpenAI, you switch by changing one line of code. If you own the computer, you keep 94.5% of what it earns.
Alpha software. Pricing and features will change. Read the paper
from openai import OpenAI client = OpenAI( base_url="https://api.openai.com/v1", base_url="https://api.tezgrid.com/v1", api_key=os.environ["TEZGRID_API_KEY"],) client.chat.completions.create( model="Qwen2.5-1.5B-Instruct", messages=[{"role": "user", "content": "Hi"}], stream=True,)
01 How it works
One endpoint in front of a fleet you don't have to run.
The gateway handles auth, limits, routing and metering. Operators run the models on their own hardware. Your code only ever sees a standard OpenAI response.
-
Your app sends a normal chat request
Point any OpenAI client at
api.tezgrid.com/v1. Streaming, models list and errors behave the way your SDK expects. -
The gateway picks a node
Candidates are scored on latency, region, trust tier and price. Unhealthy nodes are probed and skipped, and a stalled request fails over to the next one.
-
The node runs the model, the ledger settles
Inference happens on the operator's machine via llama.cpp or Ollama. Every request is metered per token: 94.5% to the operator, 5.5% to the platform.
02 Two sides
A marketplace for idle inference.
Developers bring demand through an API they already know. Operators bring supply from hardware that is otherwise sitting idle.
For developers
Keep your SDK. Change the base URL.
Nothing to rewrite. Usage, keys and spend live in the console.
- OpenAI-compatible chat completions with streaming
- Per-key spend caps and hourly usage in the console
- Optional region, latency and minimum-trust headers
- Idempotency keys, so a retry doesn't charge you twice
For operators
Earn from a machine you already paid for.
Install the CLI, load a model, go online. No public port needed.
- NVIDIA GPUs, Apple Silicon, AMD and desktop or server CPUs
tezgrid doctortells you which models fit before you download- Outbound tunnel, so it works behind home NAT and firewalls
- 94.5% of every settled charge, tracked per request
03 Why it costs less
Fewer hands between the silicon and your request.
Most inference is bought, rented, repackaged and resold before it reaches you, and every layer takes a margin. TezGrid sends requests to machines that are already paid for, where the marginal cost is mostly electricity.
Typical API provider
3 markups stacked on top of the compute
TezGrid
One 5.5% fee. The other 94.5% goes to the operator.
04 Privacy & trust
What we protect, and what we don't.
Prompts carry customer conversations, source code and internal plans. On a network of independent machines, privacy has to come in layers you can pick from, with the limits of each one written down.
No inbound port
Nodes connect out to the gateway over a persistent encrypted channel. Operators never expose an inference port to the internet.
Encrypted per request
Prompts and responses are envelope-encrypted on the tunnel hop between gateway and node, so they don't cross the public internet as plain text.
Trust tiers
Every response tells you which tier served it. Send
X-TezGrid-Min-Trustto refuse anything below the tier you need.Confidential tier
For work the host must not read: only nodes on confidential hardware with an active sealed connection are eligible.
One machine, one account
Each physical host is fingerprinted and bound to one operator account and one location, which keeps duplicate and spoofed nodes out.
Baseline on every call
API keys, rate and spend limits, and request logs that record IDs and timings, not prompt text.
05 API
The API you already use.
Chat completions and models, OpenAI-shaped. TezGrid-specific behaviour is opt-in through headers, so plain OpenAI code keeps working.
- POST /v1/chat/completionsStreaming or not. Errors come back as 401 for a missing key and 402 for an empty wallet.
- GET /v1/modelsModels currently served by online nodes.
- X-TezGrid-Min-TrustRefuse nodes below a trust tier, up to Confidential.
- Idempotency-KeySafe retries: the same key is never billed twice.
- X-TezGrid-Provider-RegionResponse header. Also returned: router latency, TTFT and failover attempts.
Full reference in the API docs.
06 Pricing
Per token, priced by supply.
These are modeled marketplace rates, not a fixed quote. Real prices move with model, region and how much capacity is online.
| Model | Quant | Memory | TezGrid, modeled | Typical list | Difference |
|---|---|---|---|---|---|
| Llama 3.1 8B | Q4_K_M | 6 GB | $0.02 / 1M | $0.20 / 1M | 10× lower |
| DeepSeek R1 Distill Qwen 7B | Q4_K_M | 5 GB | $0.02 / 1M | $0.20 / 1M | 10× lower |
| Llama 3.1 70B | Q4_K_M | 40 GB | $0.12 / 1M | $0.90 / 1M | 7.5× lower |
| Qwen 2.5 72B | Q4_K_M | 42 GB | $0.14 / 1M | $0.95 / 1M | 6.8× lower |
| DeepSeek V3 | MoE, 37B active | 96 GB | $0.22 / 1M | $1.50 / 1M | 6.8× lower |
"Typical list" means published per-token rates for comparable open models from major hosts. During the alpha the gateway settles a flat 1 µ per prompt token and 2 µ per completion token ($1 and $2 per million) on every model. Wallet top-ups are not open yet; settled charges and the operator split are visible per request in the console.
07 Run a node
From idle to online in four commands.
The installer adds the CLI and llama.cpp's llama-server, then scans the machine. Already use Ollama? That works too. You'll know which models fit before you download any weights.
Install
One command per platform. Linux, macOS and Windows. Includes
llama-server; Ollama works too.Check the machine
tezgrid doctorreports CPU, RAM, GPU, disk and runtimes.tezgrid recommendlists models that fit.Install a model
GGUF weights from the catalog, Ollama, or your own URL for closed models.
Go online
tezgrid upregisters the node and keeps a heartbeat.tezgrid downtakes it offline.
Earnings depend on demand in your region. Please read the installer before piping it to a shell. Full setup guide
Earnings estimate
Net monthly after electricity, for the hardware this browser detects.
Using the scenario's modeled probability.
Operator revenue minus power cost
This setup loses money at current assumptions
Power costs more than the expected work pays. Things that would change that:
Estimates come from the gateway's earnings model and your inputs. Actual earnings depend on routing, uptime and real demand, and are not guaranteed.
→ Go deeper
Read how routing, metering and trust actually work.
The research paper explains how a machine is chosen, what happens when one fails mid-answer, how every token is counted, and what we cannot promise yet. The docs cover every endpoint.