Chat.
Run a small language model in your browser. Stream the answer, stop it, edit the prompt and try again.
Open workspace →Crypto research. Unfinished ideas. A model that can run right here.
No account. No API key for local models. You choose what leaves your device.
FOLLOW THE MESSAGE
Run a small language model in your browser. Stream the answer, stop it, edit the prompt and try again.
Open workspace →Run the same prompt with two local models. Keep the exact model name, response time and output side by side.
Try a comparison →Export a readable transcript or encrypt a copy with a passphrase. Open it again when you want to continue.
Open the note vault →SMALL MODELS / REAL INFERENCE
Models download only when you load one. Device memory and WebGPU support determine what can run.
YOUR WORKSPACE
SAME QUESTION / TWO PERSPECTIVES
Runs one model at a time to avoid holding two models in GPU memory.
MODEL A
MODEL B
TAKE YOUR NOTES WITH YOU
Export the current conversation. Use a passphrase to encrypt the file, or save a readable transcript. No server vault or account is involved.
AES-GCM encryption with a PBKDF2-derived key. Keep your passphrase separately; the site cannot recover it.
WHAT THIS VERSION DOES
The model runs through WebLLM using your browser’s WebGPU support. Model files come from the model publisher when you choose to load them. Prompts are not sent to a hosted inference server in Local mode.
Models this small can make factual and reasoning errors. Use them for drafts and explanations, then check cryptocurrency contract facts against primary sources.
You can supply an OpenAI-compatible HTTPS endpoint, model ID and key. Requests travel directly from your browser to that provider. The provider can read the request; this mode does not claim a hardware enclave or zero retention.
Keys remain in memory and are excluded from exports. CORS restrictions, provider authentication and provider charges still apply.
Conversations stay in this tab’s memory. Reloading clears an unexported conversation. Cached model files do not contain your chat. Encrypted exports contain a random salt, nonce and ciphertext.
The site does not perform wallet actions, approve tokens or promise investment outcomes.
Every response shows the actual model and runtime mode. Comparison results include elapsed time. Generation failures remain errors; a scripted answer is never substituted.
The model’s weights must be on your device before local inference can run. The page shows real download progress, and subsequent loads can use the browser cache.
WebGPU, compatible graphics drivers and enough device memory are required. Try a current desktop browser or use your own API endpoint. There is no fake local-mode fallback.
This version offers local inference and explicit external API access. It does not operate an anonymous routing network or an attested hardware enclave.