Introducing Teti Titan 5 and Titan 5 Flash
Two models join the family today. Teti Titan 5, the largest model in the family, is available to Pro, Max, and Quantum users. Teti Titan 5 Flash replaces Aura 3 as the model on the free plan — nobody has to do anything, it is already the default.
The Titans came before the Olympians — the elder generation, named for sheer scale rather than for any one domain. Metis was one of them, which makes the lineage the right way round: our reasoning model carries the name of a single Titaness, and the models that hold an entire body of work at once carry the name of her kind. Flash keeps the name because it is one of them: the same order of model as the flagship, with a different trade between size and speed.
Both are for the problems that are large before they are hard: a whole repository, a year of correspondence, a full case file. Titan 5 is the one for demanding work, and it answers faster than anything else in the family. Titan 5 Flash is the everyday one, at the lowest price in the family — and it takes the free plan from a 31-billion-parameter dense model to a 320-billion-parameter sparse one.
What's new
In both models:
- A million tokens of context. Roughly a 700,000-word window, four times the 262K the free plan had until today: a large codebase, hundreds of documents, or an agent run that goes on for hours, without splitting the work into pieces that lose sight of each other.
- Thinking built in. Reasoning is part of the model, not a separate mode to opt into. On everyday turns they answer directly; when you ask them to think, they think harder rather than switching models.
- Native vision. Screenshots, diagrams and scanned documents are read directly — on the free plan too.
- Up to 65,536 tokens of output. Long enough for a full report, a migration plan, or a substantial file, in one response.
- Served from the European Union. End to end, including when a request fails over to another provider.
Benchmarks
Named public benchmarks — the kind that keep their name and version over time, so a score from today still means something a year from now.
| Benchmark | Titan 5 | Titan 5 Flash | What it measures |
|---|---|---|---|
| Deep SWE v1.1 | 74.2 | 63.4 | resolving real issues in real repositories |
| Terminal-Bench 2.1 | 90.6 | 84.3 | multi-step work in a shell |
| GPQA Diamond | 90.9 | — | graduate-level science questions |
| AIME 2025 | 87.5% | — | competition mathematics |
| ExtractBench, short documents | — | 96.3 | pulling structured data out of a document |
| AutomationBench | — | 48.8 | end-to-end task automation |
The number that matters for professional work is the first one. On Deep SWE, the agentic coding benchmark, 74.2 puts Titan 5 in the same band as GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5, which all land around 74% — models that cost several times as much. Titan 5 Flash scores 63.4 there. That gap is the honest reason the two models cost different amounts, and the reason to step up when the work is engineering rather than everyday.
Where Flash stands out is ExtractBench. Reading a document and returning what you asked for is the single most common thing people do with an attachment, and at 96.3 on short documents this is a model that does it reliably rather than approximately.
Speed is the other half of the trade. Titan 5 sustains several times the output rate of the model it succeeds at the top of the family — a difference you feel on every long answer, and one that comes from how the model is built rather than from how hard we run it.
What the models are made of
Titan 5 runs 552 billion parameters, of which only 8 billion are active on the input side and 16 billion on the output side — a sparse mixture-of-experts with a causal encoder-decoder architecture, alongside a 196-billion-parameter conditional memory that is consulted by lookup rather than run on every token. It was trained from scratch on 45 trillion tokens. That is the whole reason a model this large can answer this fast, and it is also why it costs what it costs: you pay for the fraction that runs, not for the size of the file.
Titan 5 Flash runs 320 billion parameters, of which 18 billion are active per token. It is a sparse mixture-of-experts and the first open frontier model to combine sparse and linear attention, which is what keeps a model of this size affordable to run. It is also the first natively multimodal model of its series, trained on a 30-trillion-token multimodal corpus.
For comparison, the model the free plan ran on until today is dense: 30.7 billion parameters, all of which execute on every token. Flash is ten times the size and activates a little over half as much per token.
The efficient answer is also the cleaner one
Neither model runs all of itself on every word. Sixteen billion parameters out of 552 produce a Titan 5 token, eighteen out of 320 a Flash one, so a request costs a fraction of the computation a dense model of that size would need.
The environmental arithmetic follows from that, without needing a slogan: less computation per token, and far less time occupying a GPU per answer. An answer that arrives in a fifth of the time occupies roughly a fifth of the machine time, and machine time is what electricity is spent on.
Holding the whole problem
A large context window is easy to advertise and harder to use well. The failure mode is a model that technically accepts a million tokens and then attends mostly to the last few thousand, so the middle of your document quietly stops counting.
Titan 5 works the other way round. It favours input-heavy turns, where the prompt is far larger than the answer — long documents, accumulated tool results, retrieved context, an agent's own history. That is the shape most professional work actually has, and it is the shape Titan 5 handles best.
Why European infrastructure matters
Where a model runs is not a detail of the invoice. It decides which jurisdiction can compel access to what passes through it, and which rules govern the people operating the machines. Teti is the private AI; serving our models from Europe is the same commitment expressed in hardware rather than in policy.
This does not make the models European — the weights are open and come from elsewhere, as they do for every model in the family. It makes the processing European, which is the part that touches your data.
Both models are served from data centres inside the European Union. If one endpoint is unavailable, requests fail over to another that is also inside the European Union. No part of the chain sits outside it, so a fail-over cannot move your data out.
And it is a contract, not a promise on a website. Our infrastructure providers are European companies, bound by European law and by data processing agreements with a zero-retention policy written into them: what you send is deleted as soon as the request is served, and it is never used to train, fine-tune, evaluate or improve any model. Narrow exceptions survive for security, abuse and legal orders, as they must — and they are named in the contract rather than left to discretion.
This used to be something you got by paying. With Titan 5 Flash, it is now what the free plan runs on.
Governance
Titan 5 and Titan 5 Flash ship behind the same commitments as the rest of Teti — governed by our Charter and Ethics framework, the public rules that define how we build.
Pricing and availability
| Per million tokens | Titan 5 | Titan 5 Flash |
|---|---|---|
| Input | $1.00 | $0.75 |
| Output | $3.50 | $1.75 |
| Cached input | $0.35 | $0.35 |
Titan 5 is priced below Metis 3 on both ends. Titan 5 Flash carries the same input price as the previous free-plan model, for a model ten times its size and four times its context.
Titan 5 is available today to Pro, Max, and Quantum users across all Teti products. Titan 5 Flash is the default model on the free plan and is available on every plan. On the Cloud API, developers can select teti-titan-5 or teti-titan-flash-5, or try both directly in the console.
Choosing between them
The deciding question is what your work is made of.
Titan 5 Flash is the everyday model: the same million-token window as the flagship, native vision, and the lowest price in the family. Titan 5 answers several times faster and is clearly ahead on agentic coding, which is what you are paying for when the work is engineering, research, or anything that runs for hours. On a paid plan you can move between them per task.
Metis 3 and Aura 3 are on their way out. Both remain available and unchanged for now, and both will be deprecated before long. When they are, requests move to their current equivalents automatically — Aura 3 to Titan 5 Flash, Metis 3 to Titan 5 — so an integration that names the old model keeps working and simply gets served by the newer one. We will announce the dates separately, with a migration window, well before anything changes.
