The transformer layer set for smol transformer + graph-augmented inference.
Attention supports self-attention (causal) and cross-attention (no causal
mask, K/V from graph embeddings). RoPE included for positional encoding.
👾 Generated with [Letta Code](https://letta.com)
Co-Authored-By: Letta Code <noreply@letta.com>
acc was seeded with @splat(bias) and reduced with .Add, so every
linear output carried 7 extra copies of its bias. conv2d was unaffected
(bias via memset), which is why the trunk matched zig-solver to 5e-6
while fc/head activations diverged — the 81% labels regression.
Also adds a scalar tail for in_n < 8 (pose_fc has in_n=4) and a
known-answer selftest (test_linear).
👾 Generated with [Letta Code](https://letta.com)
Co-Authored-By: Letta Code <noreply@letta.com>