Skip to main content

Taalas — a Model Etched Into Silicon, and the Ceiling It Actually Breaks

· 15 min read
Baris Bayrak
Software Engineer

A Toronto startup with 24 employees has built a chip whose 53 billion transistors are the weights of a language model rather than a place to put them. On Llama 3.1 8B it serves 17,000 tokens per second to a single user — 48× the 353 tok/s Taalas measured on a Blackwell B200, and 8.6× the 1,981 tok/s that Cerebras, the previous leader, manages on the same model. AMD announced on 6 August 2026 that it has agreed to acquire the company. Taalas's homepage claims Hardcore Models are "1000x more efficient than their software counterparts," which is the kind of number that usually means nothing. Here it means something specific and much narrower than it sounds — and the physics underneath is not only real, it names a ceiling that no amount of GPU engineering gets past.

First, What Is Actually Established​

Everything below distinguishes three tiers of evidence, because the interesting part of this story is how they differ.

Primary, from Taalas (product page, Bajic's launch post): 815 mm², TSMC 6nm, 53 billion transistors, 17k tokens/sec/user on Llama 3.1 8B, 1k-in/1k-out. Also on the record: the quantization is a custom 3-bit base type with 3-bit and 6-bit parameters mixed, which Bajic states "introduces some quality degradations relative to GPU benchmarks."

Primary, from AMD (newsroom): a definitive agreement, "subject to customary closing conditions and regulatory approvals." The release does not state a price, and the deal is not closed — so the shorts are slightly ahead of reality.

Where the comparison numbers live, and a correction. The competitor figures come from a chart image on Taalas's own site, which means they cannot be quoted from page text and have to be read. We read it: Taalas HC1 16,960 · Cerebras 1,981 · SambaNova 932 · Groq 594 · Nvidia B200 353 · Nvidia H200 230 tokens/sec/user. The B200 figure is Taalas-measured; the Groq, SambaNova and Cerebras figures are sourced on that chart to Artificial Analysis. An earlier draft of this post used 350 and 200 for the two Nvidia baselines, taken from secondary coverage rather than the chart. Those were low: the correct baselines are 353 and 230, which makes the head-to-head ratios 48× and 74×, not 48× and 85×. Taalas's own written claim is the more modest "nearly 10X faster than the current state of the art" — which checks out against Cerebras at 8.6×.

Third-party: The Next Platform got the only substantive technical interview. The founder history is worth knowing: Ljubisa Bajic founded Tenstorrent, a general-purpose AI accelerator company, then left to build Taalas around the opposite bet; all three co-founders came from Tenstorrent, and before that AMD.

Not established: there is no independent third-party benchmark. The B200 baseline was run by Taalas.

The Ceiling That Explains Everything​

LLM decoding is not compute-bound. To emit one token you must touch every weight exactly once, and at 16-bit precision that is 16.06 GB of data that has to be hauled out of DRAM per token, forever. This puts a hard ceiling on any conventional accelerator, from any vendor:

ChipHBM bandwidthWeights read per tokenCeilingTaalas's measured figure
H2004.8 TB/s16.06 GB299 tok/s230 (77% of ceiling)
B2008.0 TB/s16.06 GB498 tok/s353 (71% of ceiling)

That table is the whole story, and it cuts both ways. Read one way, it says modern GPUs are nowhere near the theoretical limit. Read the other, and it says they are already at 71–77% of a wall that physics put there — which is a well-engineered chip doing its job, not a chip that can be optimised into 17,000 tok/s. You cannot get to 17,000 by making the memory faster. You have to stop reading the weights.

The self-consistency here is worth noting, and it survives the correction above. Taalas's published Nvidia baselines sit at 71% and 77% of the ceilings computed from published bandwidth — that is, they are generous to Nvidia. A vendor inflating its own ratio publishes low competitor numbers; Taalas published high ones. Their 48× is therefore a more conservative claim than it needed to be.

A Weight That Is Not Data​

The mechanism is a mask ROM. In a mask ROM a bit is not a charge in a cell — it is the presence or absence of a contact at a wire crossing. Value is geometry in a metal layer. Bajic's description of the consequence:

"We have got this scheme for the mask ROM recall fabric — the hard-wired part — where we can store four bits away and do the multiply related to it — everything — with a single transistor. So the density is basically insane. And this is not nuclear physics — it is fully digital. It is just a clever trick that we don't want to broadcast."

The lever is that a weight is a constant, fixed at fab time. Multiplying by a constant is not really arithmetic — it is a lookup. So there is no multiplier array to build (Bajic calls multipliers "the big boy piece of the computer") and no weight movement at all. Only activations move, and only across the die. The architecture pairs that immutable ROM fabric with a mutable SRAM fabric holding the KV cache and LoRA adapters — the model's skeleton is silicon, its mutable state is not.

Attributed where it belongs, and still unconfirmed. Taalas is explicit that it is not disclosing the fabric ("we don't want to broadcast"). The circuit-level story that circulates in secondary coverage — that the chip pre-computes all sixteen products of a 4-bit slice at manufacturing time and lets pass transistors select the correct pre-wired channel — is exactly that: secondary coverage, not a Taalas statement. It is a plausible reading of "one transistor stores four bits and performs the associated multiply," and it is not ours to assert. The functional claim is Bajic's, on the record. The mechanism behind it is not.

Two physical consequences follow, and they are why 200 W is credible. Multiplication by a constant bit decomposes into a conditional wire, so the dominant energy term disappears; and a ROM cell that is not being read does not toggle, so dynamic power tracks only the activations moving through. The resulting power density is 0.25 W/mm² (200 W over 815 mm²) — low enough that standard air cooling suffices, and low precisely because the data is not moving.

Does the Transistor Budget Add Up?​

This is the one part of the design that is externally checkable, so it is worth checking. Take Bajic's own ratio — four bits stored and multiplied per transistor — and Taalas's own quantization statement (3-bit base, 3-bit and 6-bit mixed, so call it ~4 bits per parameter on average):

ComponentDerivationTransistors
Weights, etched in ROM8.03B params × 4 bits ÷ 4 bits per transistor8.0B (15%)
KV cache, adders, muxes, clocking, routing53B − 8.0B~45B (85%)

The arithmetic closes on the stated figure, which is not nothing — a fundamentally different storage scheme that came out three orders of magnitude off the announced transistor count would be a problem. Note what the 15% is and is not: it is the transistor share, not the area share, and ROM cells and logic have different densities, so the two are not interchangeable. What it does say is that the weights are a minority of the silicon: the rest is the compute fabric and the mutable state, and the KV cache has to live in that budget — which is why Taalas's stated flexibility ("configurable context window size and LoRA fine-tuning") is bounded by how much SRAM was sized at tapeout.

A useful contrast: the B200 carries 208 billion transistors. Its die area is roughly 1,600 mm² across its two dies, which puts it near 130 MTr/mm² against HC1's 65 MTr/mm² — about double. That area figure is our estimate, not a published NVIDIA number, and the comparison is order-of-magnitude only. Even loosely, though, it makes the point: the ROM chip is less dense, not more. Density was never the problem being solved. Not moving the data was.

Three Numbers Called "Efficiency"​

This is where the headline claims have to be taken apart, because three quite different figures get flattened into one.

Latency: ~74×, and it is structural. 17,000 tok/s is 58.8 µs per token, which across 32 transformer layers is 1,838 ns per layer — for RMSNorm, QKV projection, attention and the MLP combined. For contrast, take the same ceiling table and ask what one layer costs on an H200: 502 MB of weights per layer at 4.8 TB/s is 105 µs of pure memory time, before a single multiply. That is a 57× gap in memory time alone, and no amount of batching closes it, because it is per-token serial latency either way. End to end against the H200's 230 tok/s the ratio is 74×; against the B200's 353 tok/s, 48×. It is the most defensible number in the whole story.

Energy per token: a few times, not a thousand. HC1 spends 11.8 mJ per token at the stated 200 W per card. An H200 serving one stream at 230 tok/s spends 3,043 mJ per token — 259× more. That is the comparison the headline invites, and it is the wrong one, because batching exists to amortise exactly this. A batched H200 at 20,000 tok/s aggregate is already at 35 mJ per token, only 3× behind. That number is the load-bearing assumption in this post, so here is the sensitivity rather than a single figure:

Batched H200 aggregateH200 mJ/tokenRatio to HC1
5,000 tok/s14012×
10,000 tok/s706×
20,000 tok/s353×
40,000 tok/s17.51.5×

And a second caveat running the other way: 200 W is the per-card figure from the Next Platform interview, while the product page quotes 2.5 kW for a ten-card server — which is 250 W per card including its share of host, NICs and fans. On the higher figure HC1 is 14.7 mJ/token and the honest ratio becomes 2.4×. So the real answer is a range of roughly 2.4× to 3×, and it is a range because both inputs are.

The 1000× is a utilisation figure, and it has a range too. Take the B200 at 353 tok/s. That is 353 × 16.06 GFLOP = 5.67 TFLOP/s. Against a dense FP4 peak of 9,000 TFLOP/s that is 0.063% of peak — invert it and you get roughly 1,600×. But the denominator is the one place this arithmetic can be pushed, because the B200 never executes in FP4 at single-stream: at FP8 the peak is 4,500 TFLOP/s and the ratio is ~800×, and at FP16, the precision actually being executed, the peak is 2,250 TFLOP/s and the ratio is ~400×. So the honest reading of "1000× more efficient" is: it is the reciprocal of GPU utilisation at batch size 1, and it brackets between 400× and 1,600× depending on the precision you pick. That range is a real result — the GPU genuinely runs at a fraction of a percent of peak on a batch-1 decode, because it is starved for bandwidth — and it is not a claim about joules. On joules per FLOP at peak, the B200 is the better chip: 9.00 TFLOP/J against HC1's 1.37 TFLOP/J.

The physical bedrock under all of this is Horowitz's ISSCC 2014 numbers, and they are brutal:

OperationEnergy
32-bit floating-point multiply3.7 pJ
On-chip cache access or functional operation (his prose)10 pJ
DRAM I/Oover 20 pJ per bit
Full DRAM access1–2 nJ

Reading a single weight from DRAM costs many times more than the arithmetic it feeds, and a weight is sixteen bits. Delete the read and most of the bill disappears. That is the entire thesis, and it is a real one.

What It Cannot Do​

  • One model per chip. A new model is a new tapeout. Only two metal layers change, so it is far cheaper and faster than a full design — Taalas quotes two months from an unseen model to deployable cards — but it is a hard discontinuity, not a config file.
  • Quality is knowingly traded. Taalas's own words on the 3-bit/6-bit mix: "some quality degradations." Comparing a degraded 8B model at 17k tok/s to a clean one at 230 is not comparing like with like.
  • Long context is the weak axis. The KV cache grows per token and lives in fixed on-chip SRAM. Window size is configurable within that budget, but the budget itself is fixed at tapeout.
  • Mixture-of-experts inverts the economics. This is our argument rather than a sourced claim. MoE activates a small fraction of weights per token, so a GPU's bandwidth bill scales down with the active fraction. A ROM chip must bake all the experts into silicon to serve any of them, then traverse only some — the full area cost, a fraction of the benefit. It is the architecture that most directly attacks the premise.
  • One die is one die. At 815 mm² the part is already at 95% of the 858 mm² reticle limit — the standard single-exposure field for 193i lithography. Anything larger than a mid-sized model has to be pipelined across chips.
  • No independent verification. Vendor-run benchmarks, as of the launch.

Why AMD Bought It​

Two things are happening at once. The obvious one: The Next Platform reported that Nvidia had just bought Groq for ~$20 billion, and AMD has no equivalent asset and no CUDA moat to defend. Buying a decode-specialised chip is differentiation by a different route, and AMD's stated plan — integrate it alongside Instinct GPUs in Helios racks — is a division of labour rather than a replacement. GPUs keep prefill, training and flexibility; the hardwired part eats the serial-decode latency problem that GPUs cannot structurally solve.

The less obvious one is who is being bought. Bajic was an AMD architect on hybrid CPU-GPU designs before founding Tenstorrent; his co-founders came through AMD and ATI as well. On the economics, the framing from Taalas is that hardening a model costs about one hundredth of training it. That only works if you need to serve that specific model for long enough to amortise the tapeout — a bet that models stop churning every three months. Llama 3.1 8B is still being served today, years after release, and frontier release cadence has been lengthening. It is a bet with evidence behind it, which is not the same as a safe one.

What We Learned​

What worked: the derivation. Almost every dramatic number in this story is reproducible from constants that are published — parameter count, HBM bandwidth, TDP, transistor count, die area — and doing that arithmetic first is what separated the three "efficiency" claims from each other. The 299 tok/s ceiling is not Taalas marketing; it falls out of dividing two public specs, and it is why the 17,000 figure is plausible rather than absurd. Likewise the transistor budget closing at 15% weights and 85% everything-else, and HC1's power density landing low in a way consistent with the claim that nothing is moving.

What didn't — and what an independent audit caught. Our first draft took its Nvidia baselines from secondary coverage instead of from Taalas's own chart: 200 and 350 tok/s where the chart says 230 and 353, which turned a 74× ratio into a claimed 85×. The baselines were not on the page as text; they only exist inside a chart image, and we had not read the image. Three other things did not survive review either: we had attributed the "1000×" claim to the AMD announcement when it is only on Taalas's homepage; we had described the pre-computed-product mechanism as our inference when it is actually circulating in secondary coverage and should have been attributed as such; and we had credited Taalas with a first version built on general-purpose silicon, conflating it with Tenstorrent, which is a different company. All four are now fixed above. The lesson is narrow and repeatable: when a number only exists inside an image, a text extraction will silently substitute someone else's version of it, and every ratio computed from it inherits the error.

The uncomfortable part: the design removes one constraint and promotes another. Once nothing is bandwidth-bound, what is left is per-layer serial latency, power delivery, and the reticle limit — and then the deepest constraint of all, which is that the model is frozen. Whether that is a good trade depends on a prediction about how fast models will keep changing, which is a prediction about the world rather than about silicon. The hardware is sound. The bet is the interesting part.


The arithmetic in this post is re-derived, not transcribed: every figure is recomputed from the published constants above by scripts/check-taalas-post.py, which will fail the build if the numbers drift.