Use the crowd.
One endpoint. The router places every request on the best target it can find: the browser network when it can serve the model, cloud fallback when it can't. Your existing OpenAI SDK works unchanged.
Automatically chooses the cheapest execution path that satisfies the request.
Distributed Qwen-class inference across the browser pool.
Coding-optimized inference on distilled reasoning models.
Distributed embeddings. Small model, highly parallel, browser-native.
EST Prices are published after cost benchmarks against reference providers. No savings are claimed until they are measured.
Playground
Requests go through the same server-side gateway as the public API. Provider keys never leave the server. The execution target and latency are real; the shard plan shows how the browser pool partitions the work and is simulated until distributed LLM execution ships.
Drop-in for your OpenAI client.
Point the base URL at Brain, choose brain/auto, and keep the rest of your code.
import OpenAI from "openai";
const brain = new OpenAI({
baseURL: "https://YOUR_BRAIN_HOST/v1",
apiKey: process.env.BRAIN_API_KEY,
});
const res = await brain.chat.completions.create({
model: "brain/auto",
messages: [{ role: "user", content: "Audit this contract…" }],
});