brain
Developers

Build on the crowd.

Brain speaks the OpenAI Chat Completions format. Change the base URL, set model: "brain/auto", and the router does the rest.

POST/v1/chat/completions
GET/v1/models
Python
from openai import OpenAI

client = OpenAI(
    base_url="https://YOUR_BRAIN_HOST/v1",
    api_key="YOUR_BRAIN_API_KEY",
)

res = client.chat.completions.create(
    model="brain/auto",
    messages=[{"role": "user", "content": "Summarize this audit report."}],
)

print(res.choices[0].message.content)
print(res.model_extra["brain"]["target"])  # BROWSER_NETWORK | CLOUD_FALLBACK | EXTERNAL_MODEL_PROVIDER

Request path

Target status below is read from this server's configuration at request time. Keys stay in server environment variables and are never sent to the browser.

Client
OpenAI SDK / curl
Gateway
auth · validate · rate limit
Router
rank eligible targets
Browser poolV1: verify jobs only
BROWSER_NETWORK
Cloud fallbacknot configured
CLOUD_FALLBACK
External providernot configured
EXTERNAL_MODEL_PROVIDER
Response
503 · routing trace

No target can serve LLM requests on this server right now; the gateway returns 503 with the routing trace instead of a fabricated answer.

Routing

Ineligible targets are filtered out, the rest ranked by a weighted score. Weights are configurable per request class. Every response carries the full decision in brain.routing.

CompatibilityHard filter. Can this target run this model at all?required
CapacityFree capacity on the target right now.0.10
LatencyEstimated time to first full response.0.30
CostUSD per 1M tokens. Unknown cost is scored neutral, not free.0.30
ReliabilityRolling success rate of the target.0.30
Network preferenceBias toward the browser pool when it is eligible.+0.25

Response

Standard OpenAI shape, plus a brain extension that tells you where the request ran and why. SDKs ignore unknown fields.

400invalid_requestMalformed body, bad roles or empty messages.
401invalid_api_keyMissing or unknown API key (when BRAIN_API_KEYS is set).
404model_not_foundModel id is not one of the brain/* models.
413too_largePrompt exceeds the V1 size limit.
429rate_limitedPer-IP fixed window exceeded.
502upstream_failedEvery eligible target failed; attempts are listed.
503no_provider_availableNo target can serve the model; routing trace included.
200 · application/json
{
  "id": "chatcmpl-3f9c…",
  "object": "chat.completion",
  "model": "brain/auto",
  "choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
  "usage": { "prompt_tokens": 412, "completion_tokens": 233, "total_tokens": 645 },
  "brain": {
    "target": "CLOUD_FALLBACK",
    "provider": "cloud-fallback",
    "latencyMs": 1184,
    "routing": { "ranked": [ … ], "selected": { … } },
    "attempts": [{ "providerId": "cloud-fallback", "ok": true }],
    "plan": { "provenance": "simulated", "shards": [ … ] }
  }
}

Node protocol

What a contributing browser does. Session tokens are random, stored only as hashes, and bound to one node.

POST/api/benchmark/challengeServer issues a seeded WGSL workload. Its clock starts now.
POST/api/nodes/registerNode returns the result; server spot-checks secret blocks and scores on its own clock.
POST/api/nodes/joinStandby node goes live and is announced to the network.
POST/api/nodes/heartbeatEvery 4s. Silent for 12s = offline.
POST/api/jobs/nextPull a job. Inputs are a seed; expected outputs never leave the server.
POST/api/jobs/resultSubmit hashes. Verified, reputation updated, compute credited.
POST/api/nodes/leaveRevokes the session token.
GET/api/network/streamServer-Sent Events: joins, leaves, verified jobs.

Contributors are adversarial.

The server never trusts a client's claimed GPU, score, uptime or results. Rewards follow verified work only.

Server-issued challengesENFORCED

Benchmarks and jobs are generated server-side from secret seeds. Clients cannot pick their own work.

Server-clock scoringENFORCED

Compute score = verified ops ÷ server-measured wall time. The client's timing is recorded but never trusted.

Secret spot-checksENFORCED

The server recomputes 6 randomly chosen rows / blocks it never reveals. One mismatch fails the job.

Canary jobsENFORCED

20% of jobs are small canaries with a fully known answer.

Plausibility boundsENFORCED

Results returned faster than physically possible for the workload are rejected.

ReputationENFORCED

EWMA over outcomes; failures weigh double. Below 0.35 the node is banned.

Rate limitingENFORCED

Per hashed IP, per minute: 600 node calls, 60 API calls, 20 playground runs.

Redundant executionINTERFACE

Same unit to N nodes, majority result wins. Policy + comparison implemented; dispatcher wiring lands with LLM shards.