GraniteForCausalLM
GraniteForCausalLM at 40 layers and hidden size 2048. 9 published checkpoints share this shape.
The pass (full forward)
The order of operations for one forward. ×N marks a position that fires once per layer.
17 steps across nested levels, recorded from one complete forward of this shape in the author's independent implementation — the structure is what that run emitted, not a reading of the configuration.
forward
├─ embed
├─ layer ×N
│ ├─ rmsnorm_attn
│ ├─ wq
│ ├─ wk
│ ├─ wv
│ ├─ rope
│ ├─ attn
│ │ ├─ scores
│ │ └─ attn_mix
│ ├─ wo
│ ├─ rmsnorm_mlp
│ ├─ w_gate
│ ├─ w_up
│ └─ w_down
└─ head
Recorded separately from the structure: layer fired 40 times. A ×N run says only that it repeated — how long a run is is not part of the structure, so two models differing only in depth have the same pass.
The same thing in canonical form:
0(1 2*N(3 4 5 6 7 8(9 10) 11 12 13 14 15) 16)
Numbers are positions in the pass, not layer indices. Two passes count as the same structure when these strings match.
Implementing it
One forward, written out, per checkpoint whose arithmetic actually differs. A shape groups checkpoints by class, depth, width and layer stack — none of which decides an activation, a rope base or a routing rule — so where the members of this shape disagree there is a block each, and each names the checkpoint it was generated from.
It is not the recorded pass above, which is captured from a real forward; it is what the published configuration says the arithmetic is. Where the configuration does not say, the line says that instead of guessing.
The names are the ones the published checkpoint uses where a name is published, and canonical otherwise. A checkpoint loader may store them differently — fusing a gate/up pair into one matrix, or splitting one published projection in two — and those are that loader's names, not the model's. Everything spelled here is what the download contains.
From granite-3.0-2b-instruct
Also covers granite-guardian-3.0-2b, granite-3.0-2b-base.
x = embed[ids] * 12 # [T, 2048] <- embedding_multiplier
for i in 0 .. 39:
h = rmsnorm(x, layers[i].input_layernorm, eps=1e-05)
q = h @ layers[i].self_attn.q_proj.T # [T, 32*64]
k = h @ layers[i].self_attn.k_proj.T # [T, 8*64]
v = h @ layers[i].self_attn.v_proj.T # [T, 8*64]
q, k = rope(q, k, theta=10000)
k, v = repeat_kv(k, v, 4) # 32 query heads share 8 key/value heads
a = softmax(q @ k.T * 0.015625, mask=causal) @ v # <- attention_multiplier
a = a @ layers[i].self_attn.o_proj.T
x = x + a * 0.22 # <- residual_multiplier
h = rmsnorm(x, layers[i].post_attention_layernorm)
g = silu(h @ layers[i].mlp.gate_proj.T) # [T, 8192]
u = h @ layers[i].mlp.up_proj.T
y = (g * u) @ layers[i].mlp.down_proj.T
x = x + y * 0.22
x = rmsnorm(x, model.norm)
logits = (x @ embed.T) / 8 # tied to the input embedding; <- logits_scaling
From granite-3.3-2b-instruct
Also covers granite-3.3-2b-base.
x = embed[ids] * 12 # [T, 2048] <- embedding_multiplier
for i in 0 .. 39:
h = rmsnorm(x, layers[i].input_layernorm, eps=1e-05)
q = h @ layers[i].self_attn.q_proj.T # [T, 32*64]
k = h @ layers[i].self_attn.k_proj.T # [T, 8*64]
v = h @ layers[i].self_attn.v_proj.T # [T, 8*64]
q, k = rope(q, k, theta=10000000)
k, v = repeat_kv(k, v, 4) # 32 query heads share 8 key/value heads
a = softmax(q @ k.T * 0.015625, mask=causal) @ v # <- attention_multiplier
a = a @ layers[i].self_attn.o_proj.T
x = x + a * 0.22 # <- residual_multiplier
h = rmsnorm(x, layers[i].post_attention_layernorm)
g = silu(h @ layers[i].mlp.gate_proj.T) # [T, 8192]
u = h @ layers[i].mlp.up_proj.T
y = (g * u) @ layers[i].mlp.down_proj.T
x = x + y * 0.22
x = rmsnorm(x, model.norm)
logits = (x @ embed.T) / 8 # tied to the input embedding; <- logits_scaling
From granite-3.1-2b-instruct
Also covers granite-3.2-2b-instruct, granite-guardian-3.1-2b, granite-3.1-2b-base.
x = embed[ids] * 12 # [T, 2048] <- embedding_multiplier
for i in 0 .. 39:
h = rmsnorm(x, layers[i].input_layernorm, eps=1e-05)
q = h @ layers[i].self_attn.q_proj.T # [T, 32*64]
k = h @ layers[i].self_attn.k_proj.T # [T, 8*64]
v = h @ layers[i].self_attn.v_proj.T # [T, 8*64]
q, k = rope(q, k, theta=5000000)
k, v = repeat_kv(k, v, 4) # 32 query heads share 8 key/value heads
a = softmax(q @ k.T * 0.015625, mask=causal) @ v # <- attention_multiplier
a = a @ layers[i].self_attn.o_proj.T
x = x + a * 0.22 # <- residual_multiplier
h = rmsnorm(x, layers[i].post_attention_layernorm)
g = silu(h @ layers[i].mlp.gate_proj.T) # [T, 8192]
u = h @ layers[i].mlp.up_proj.T
y = (g * u) @ layers[i].mlp.down_proj.T
x = x + y * 0.22
x = rmsnorm(x, model.norm)
logits = (x @ embed.T) / 8 # tied to the input embedding; <- logits_scaling
Notes that apply to more than one block
Stated once here rather than under each block above.
head_dim is not published; 2048 / 32 = 64 is used.
The Granite multipliers are the part with no Llama analogue: embedding_multiplier, attention_multiplier, residual_multiplier, logits_scaling. They are ordinary floats and a port that ignores them still produces fluent text, which is what makes them worth printing where they act.
Geometry
| layers | 40 |
| hidden size | 2,048 |
| attention heads | 32, 8 key/value |
| feed-forward width | 8,192 |
| vocabulary | 49,155 (5) / 49,159 (1) / 49,152 (3) — by checkpoint in the list below |
| trained context | 4,096 tokens (2) / 131,072 tokens (6) / 8,192 tokens (1) — by checkpoint in the list below |
| largest checkpoint | 2.6B |
Weight structure
The tensors one element of the repeating stack holds, by the names the published checkpoint uses.
| repeating stack | depth | what one element holds |
|---|---|---|
model.layers.# | 40 | input_layernorm.weight, mlp.down_proj.weight, mlp.gate_proj.weight, mlp.up_proj.weight, post_attention_layernorm.weight, self_attn.k_proj.weight, self_attn.o_proj.weight, self_attn.q_proj.weight, self_attn.v_proj.weight |
Other shapes of GraniteForCausalLM
Identical across all of them: KV heads 8.
Only the columns that differ are shown; a cell with several values means the checkpoints of that shape disagree.
| shape | layers | width | heads | FFN width | vocabulary | context | checkpoints |
|---|---|---|---|---|---|---|---|
64L x 4,096 | 64 | 4,096 | 32 | 32,768 | 100,352 | 131,072 | 3 |
40L x 4,096 | 40 | 4,096 | 32 | 12,800 | 49,152 / 49,155 / 49,159 / 100,352 | 4,096 / 8,192 / 131,072 | 19 |
40L x 2,560 | 40 | 2,560 | 40 | 8,192 | 100,352 | 131,072 | 3 |
40L x 2,048 (this sheet) | 40 | 2,048 | 32 | 8,192 | 49,152 / 49,155 / 49,159 | 4,096 / 8,192 / 131,072 | 9 |
28L x 4,096 | 28 | 4,096 | 32 | 12,800 | 49,155 | 131,072 | 2 |
How this was checked
implemented, evidence grade shape-parity, per the assessment, for each checkpoint of this shape: a converted checkpoint of this shape was compared against a reference recording; whether it holds this checkpoint's weights is not established. The implementation meant here and below is an unpublished independent inference implementation by the author.
What stands behind the block above, beyond the published configuration it is read from:
- a decoder implemented from the published configuration
- a configuration parser that refuses what it cannot express
- a checkpoint loader, which is where published tensor names are read
- an independent reference implementation of this architecture in Python, driving the published modelling code — an executable statement of what the model should compute, written against the publication rather than against any one implementation of it
Where an independent implementation and the published configuration disagree, the configuration is what this page reports and the disagreement is what it says.
Checkpoints with this architecture
| model | parameters | context | vocabulary |
|---|---|---|---|
granite-3.0-2b-instruct | 2.6B | 4,096 | 49,155 |
granite-3.3-2b-instruct | 2.5B | 131,072 | 49,159 |
granite-3.1-2b-instruct | 2.5B | 131,072 | 49,155 |
granite-3.2-2b-instruct | 2.5B | 131,072 | 49,155 |
granite-guardian-3.0-2b | 2.5B | 8,192 | 49,155 |
granite-guardian-3.1-2b | 2.5B | 131,072 | 49,155 |
granite-3.0-2b-base | 2.5B | 4,096 | 49,152 |
granite-3.1-2b-base | 2.5B | 131,072 | 49,152 |
granite-3.3-2b-base | 2.5B | 131,072 | 49,152 |