Screenshots from an iPhone, not mockups — swipe the rail, tap one to enlarge.
A ChatGPT alternative you actually own. The same thing your people already know how to use — on every phone and desktop in the building — answered by a 304-billion-parameter — Parameters are the settings a model learns during training. More of them generally means a more capable model — and a bigger machine to run it. model running on your own hardware. No accounts, no rate limits, no per-token bill, and nothing anyone types leaves your network.
Every number on this page was measured on a real deployment — two NVIDIA GB10s running DeepSeek-V4-Flash — rather than taken from a vendor's card.
Whole contracts, whole codebases, a year of incident history — in one prompt. Needle retrieval tested 20/20 out to 847,000 tokens — A standard test: hide 20 specific facts deep inside 847,000 tokens of text, then ask for each one. It found all twenty., so it is depth you can rely on rather than a number on a card.
No chunking, no embeddings, no retrieval pipeline to build, tune and quietly get wrong. The document you paste is the document it reads.
73 tokens per second — A token is roughly three quarters of a word. 73 a second is faster than anyone reads — about a page of answer every seven seconds. on one stream, 88 at peak, and 226 across six people working at once. A cached follow-up on a 110,000-token thread comes back in about 0.6 seconds.
The model is too large for one machine, so it runs across two, joined by a 200 Gb/s direct link — a direct 200-gigabit-per-second cable, which is what lets two machines behave as one that makes them behave as one. No queue, no tier, no four-o'clock rate limit.
Every reply carries real throughput, token counts and time-to-first-token; a dashboard reports the machine itself — energy per token — How much electricity one word of answer costs. It is what turns 'is this expensive to run' into a number instead of an opinion., cache pressure, GPU watts per node, measured and modelled.
Useful when someone asks what the thing costs to run, and answerable with a number rather than an estimate.
Search and page-reading are off until you switch them on, and every single call stops for a human — one prompt per request, with the retrieved text shown before the model is allowed to use it.
Outbound reach is restricted to public addresses, so it cannot be talked into reading an internal service. That is demonstrable, not a policy.
Seven days from ECMWF’s AIFS ensemble — ECMWF is Europe's weather centre. AIFS is their AI forecast model — operational since 1 July 2025, running beside their physics model. They ship 51 runs (50 perturbed plus a control) and this charts the 50: where the runs agree the forecast is reliable, where they scatter it is not, and that scatter is what a one-run weather app hides., fifty runs of the same forecast — so “will it rain Thursday” comes back as a probability instead of one model’s guess.
ECMWF report it beating their physics model by up to 20% on surface temperature. Any city, charted in the chat.
A contract, a manual, a stack of PDFs. With a million tokens the whole document goes in — not the three paragraphs a search step guessed were relevant — so the answer comes from the specification you pasted, not a summary of it.
Text and PDFs with a text layer read in under a second and cost no model at all. Attach several and ask across all of them at once.
An invoice, a whiteboard, a screenshot of an error, a scanned contract. Read on your own hardware by a vision model on your own network — no picture is uploaded anywhere.
Type a question first and it is asked of the picture instead of transcribed. The picture is uploaded to your own server and read on your own hardware — that is the guarantee. Locations are removed from photographs as well, because the file library is shared with everyone on the network.
Conversations live in your browser, which is private and fragile: a phone can clear them, and a private tab discards them when it closes. Set a passphrase and they are also kept on your own server — encrypted on your device first, with a key derived from a phrase that is never transmitted.
What is stored is a blob the server cannot read, and could not read if you asked it to. Type the same phrase on a laptop and your history is there, because the copy is keyed to the phrase rather than to the device.
Paste a link and read what is in it — the subtitles arrive in the conversation with their timestamps, and nothing is downloaded. Ask what was said about a topic and get the answer with the minute it was said at.
This is where a million tokens stops being a number: a two-hour talk is about 36,000 tokens, so it goes in whole. No chunking, no search step guessing which three paragraphs mattered.
Soon: videos with no captions at all — most Serbian ones — need speech transcription on your own hardware. Not built yet.
Five helpers, off until you switch them on. Every single call stops and waits for a human. One card per request, so you can approve a search and refuse a fetch in the same breath.
Retrieved text arrives inside a fence the model is told is data and never instruction, length-capped, with the fence marker stripped from the body so a page cannot forge its own ending.
Outbound reach is restricted to public addresses and re-checked on every redirect hop: it cannot be talked into reading your router's admin page, the server itself, or a cloud metadata endpoint.
Sixteen verified in testing — and the interesting case is the hard one. Serbian works
in both scripts, Latin and Cyrillic, and it keeps rešila and
riješila distinct instead of collapsing ekavian and ijekavian into
whichever it saw more of. Ask in Serbian and the helpers answer in Serbian:
“kako je vreme u Banjoj Luci” reaches the forecast exactly as the English
phrasing does.
Mark any part of an answer and it can be read aloud — useful when the answer is in a language you are learning, with a slower reading at 0.6× for a phrase you want to hear properly. The speaking is done by your own phone or laptop, so the text never leaves the device; the available languages are whichever ones it already has installed, and the app says which is missing and where to add it.
The model is broadly multilingual, so sixteen is what was checked rather than a limit. The interface is English — a translation job, not an architectural one.
A conversation you deliberately pin becomes readable by anyone who can open the app on your network. That is the feature, not an oversight — a pinboard on a home LAN, with no accounts and no per-user visibility.
It is the only thing here that ever leaves your device.
The case for running this inside a business is not that it is cheaper — though it is — but that it removes the conversation every legal and security team has about hosted AI.
Contracts, patient notes, unreleased source, salary spreadsheets, customer data — they go to a machine in your own rack, over your own network, and there is nowhere else for them to go.
No data-processing agreement to negotiate, no sub-processor list to audit, no region to worry about, no vendor policy change to re-review next quarter. Under GDPR, a regulator, or an NDA that simply forbids third-party processing, this is the difference between “we mitigated it” and “it does not apply”.
The whole office shares one box. Nobody rations their own context window because a long document is expensive, and nobody's work stops because a tier limit was reached at four in the afternoon.
The marginal cost of a question is electricity.
Whole contracts, whole codebases, whole incident histories, in one prompt. No retrieval pipeline to build, tune, and quietly get wrong. The specification you paste is the specification it reads.
Every outbound fetch stops and asks a human, one prompt per request, and outbound reach is restricted to public addresses — so the assistant cannot be induced by a web page into reading an internal service.
That is a property you can demonstrate to a security reviewer, not a policy you have to trust.
And the gap, before you find it yourselves: as shipped there are no user accounts. Anyone who can reach the server can use it and read anything pinned. For a household that is the right trade; for a company, put it behind the SSO or reverse proxy already fronting your internal tools, and treat pinned chats as a public noticeboard. Per-user identity is a build, not a redesign.
We deploy it on your hardware, on your network, tuned to the documents your people actually work with — and hand over something your security review can read end to end.
Nothing is welded to one model. Several run side by side on different machines — a 304-billion-parameter model spread across the pair of GB10s, and smaller, faster ones on a single graphics card — and choosing between them is a dropdown that names the machine which will answer. Everything else is unaffected: the same conversations, the same approval gates, the same files. If one machine is asleep the other answers, and the reply says so rather than failing quietly.
The same dropdown can point somewhere else entirely: another machine in your building, a colleague's GPU server, or a hosted API if you would rather not own hardware at all. The harness does not care which — but it does not hide the difference either. Point it at your own machines and the conversation never leaves them. Point it at somebody's API and it does, on their terms and their bill, through this same interface with the same gates and the same shared history.
Both options work; local is the default here for a commercial reason. The questions your people ask are business intelligence — what you are building, what is broken, which client is difficult, what you are bidding on. A quarter of that traffic describes your company more accurately than its annual report does.
Keep it on your own hardware and that record exists in exactly one place: yours. No vendor retention. No training on your data. No terms that change next quarter, and no model replaced mid-project. The hardware is still an asset at year end, and each further question costs electricity rather than a per-token fee.
GB10_API=http://127.0.0.1:8888/v1 # any OpenAI-compatible endpoint GB10_MODEL=deepseek-v4-flash-dspark # the model that endpoint serves
vLLM, SGLang, llama.cpp, Ollama, LM Studio, TGI — or a commercial API, if you want this interface in front of a hosted model. The client assumes nothing about the model, the tokenizer or the vendor. Point a second endpoint at a small local model for cheap tasks and keep the big one for real work.
| Per million tokens | Milutin, on your own box | Hosted frontier API |
|---|---|---|
| Output | ~$0.04 | $25.00 |
| Input | ~$0.005 | $5.00 |
| Rate limits | none | tier-dependent |
| Data leaves your network | no | yes |
Roughly 640× cheaper per output token, on electricity alone at $0.15/kWh. Two GB10s at about $8,000 pay for themselves against that pricing after some 320 million output tokens — about 20 days of continuous six-way use, or 50 days single-stream. Fast payback if the box stays busy, never if it idles. Cheaper again on solar.
Here because a feature list without this section is marketing.
That last one is a deliberate choice, not a consolation. The model was picked because it fits two boxes at a million tokens of context — a different competition, and one it wins.
That compares the wrong two things. Milutin is the harness, not the model — the client, the permission gates, the shared history, the media library, the dashboard. Which model answers is a setting, and it is one line of configuration.
Point it at the machines in your building and the conversation never leaves. Point it at a frontier API and you get the frontier, through this same interface, with the same gates and the same shared history — and the data leaves. Same wheel, different engine. That is the actual choice, and it is not the one the question implies.
| Same harness, behind it… | Your own hardware | A hosted API |
|---|---|---|
| Capability | As much as you bought | The frontier |
| Where your data goes | Nowhere | Their servers |
| Cost per million words out | ~$0.04 | $25 |
| Usage limits | None | Per seat, per tier |
| Long documents | Free to use fully | Priced to discourage |
| The model changing under you | Never | Whenever they ship |
| What you own at the end | The machines | Receipts |
Which makes local capability a dial rather than a verdict: it is as capable as the hardware you put behind it. Two GB10s hold a 304-billion-parameter model at a million tokens of context. More hardware holds more model. The ceiling is a purchase decision you control — and unlike a subscription, the thing you bought is still yours at the end of the year.
For completeness, the part people expect to be argued: on the hardest reasoning benchmarks the frontier hosted models are ahead, and it would be strange if they were not — you are not spending their training budget, and one of their training runs costs several times what this hardware does. That is a fact about models, not about this software, and it is why the harness lets you point at either. Use the hosted one for the work that needs it. Keep everything else in the building.
Milutin Milanković computed 600,000 years of the Earth's orbital cycles by hand, over decades, holding the whole problem in his head because there was nowhere else to put it. The joke is about context length, and nobody has to get it for the name to work.