Skip to the content.

The compute network

TokenBroker doesn’t rely on a single data center. It routes work to a network of compute nodes — desktops, phones, and personal API keys contributed by members like you — plus cloud-hosted models. This page explains how the network actually works, how to join it in whatever way fits your hardware, and exactly how you get paid for participating. If you want the case for why this exists at all, that’s Take back your compute; this page is the how.

Why a network?

Some models are small enough to run on a laptop. Others need serious GPUs. Some contributors have neither, but do have a spare phone or a free-tier API key sitting unused. A network lets the service take all of it: a node with an RTX-class GPU serves bigger local models, a modest node serves smaller ones, a phone or a GPU-less laptop relays cloud and external-provider work, and everything can fall back to a hosted cloud model when no node fits.

The more nodes join — in whatever form — the more capacity the network has, which keeps the service available around the clock and reduces reliance on any single piece of hardware, any single provider, or any single person’s uptime. That redundancy is the entire point, not an incidental benefit; see the full breakdown for why depending on more than one upstream is a deliberate design choice.

How nodes are chosen

Every request is matched to the best available node for that model:

  1. A node that already has the model pulled and meets its requirements — fastest path.
  2. A node with enough disk space to pull the model on demand.
  3. A cloud-hosted model (e.g. Ollama Cloud) when the request is for a cloud model or no capable local node is online.

Nodes only ever receive work they can actually do — the gateway checks GPU VRAM, system RAM, and storage before assigning anything. A node never gets a job for a model it can’t handle.

Dispatch tiers - which worker gets the job

Joining the network with the worker

The TokenBroker Worker is a small Windows application you download from your dashboard. Highlights:

The worker connects outbound only — no ports are opened on your router.

What the worker needs

Optional: connect your Ollama Cloud account

If you sign in to your free Ollama account in the worker, the node can take on cloud-hosted models too — including large models your hardware could never run. The worker tracks your cloud usage and reports how much remains, and it stops accepting cloud work the moment your free quota is exhausted (it resumes when the quota resets).

Optional: bring your own cloud API key

You don’t need a GPU, or even Ollama, to contribute real capacity. If you hold a free-tier API key from a supported external provider — Google AI Studio today — you can add it to your account on the Compute page, and every worker on that account (desktop or mobile) can serve jobs for that provider’s models directly, calling the provider’s own API with your key. This is a completely separate quota pool from Ollama Cloud, so an account with both contributes more capacity than either alone.

Known-floor capacity comparison: Ollama Cloud vs Google AI Studio

A brand-new worker is credited with a conservative baseline the moment it joins, rather than counting as zero capacity until it has a track record. Ollama’s baseline came from a real, deliberately bounded measurement (a controlled one-hour test run to completion with zero rate-limit errors); Google doesn’t publish a fixed limit at all, so that figure is a careful, conservative estimate instead — the live network status page always favors real measured throughput over either baseline once a worker has one.

Mobile workers (Android & iOS)

The network isn’t limited to desktops. A companion mobile app runs the same relay loop in the background — heartbeat, claim a job, run it, report back — using a foreground service so it keeps working with the screen off. A phone can’t run large local models, so it serves cloud-hosted and external-provider work (Ollama Cloud, or a provider key you’ve added to your account) rather than pulling models locally. Pairing a phone to your account takes a one-time code generated from the dashboard.

How workers earn credits

Life of one job, from creation to credit

Every completed, accepted job pays credits into your account. Quality matters:

Anti-exploit protections

The network is built to resist gaming:

These protections exist so that real users get real work done, and credits mean something.

Private (encrypted) jobs

Requests that need extra privacy can be submitted as private jobs: the hub encodes the prompt into an encrypted glyph, only trusted workers that have private jobs enabled receive it, and the worker decodes it in memory, computes, and returns the answer as an encrypted glyph. The worker never writes the prompt to disk and never displays it.

If you run a worker, you control this with the Priv on / Priv off button in the worker window (or the tray menu). Private jobs are on by default; turn them off and your node will only ever receive public work.

Live network status

Anyone — signed in or not — can see the network’s real-time aggregate capacity, how many independent accounts are contributing, and recent job activity at /network. The page also explains exactly how the numbers are calculated, including the honest caveats: some figures are directly measured, others are conservative estimates for a provider that doesn’t publish its own limits. It never shows per-node or per-person detail — only network-wide totals.

Worker dashboard

Your dashboard’s Compute page shows:

Removing the worker

Uninstall is one double-click. The worker leaves no background services, and your account keeps whatever credits you’ve earned.

Pay-gated (subscription) models

Some Ollama cloud models require a paid ollama.com account (Pro/Max) rather than the free tier — a node without that subscription simply can’t run them. The network handles this automatically and transparently:

  1. Each worker figures out on its own whether its Ollama account can access paid-tier models, and reports that status honestly.
  2. A model that turns out to need a paying account gets remembered per node — a free-tier node stops being offered (and stops advertising) work it already proved it can’t do, instead of failing the same way repeatedly.
  3. Before a job is queued at all, the network checks whether any currently online worker can actually serve that model right now. If none can, you get an immediate, clear error explaining the model may need a paid-tier worker — instead of a job that sits in the queue until a timeout expires and fails anyway.