a2ff07234c8f8c2e6b38384dbca66e564992aaa7
acc was seeded with @splat(bias) and reduced with .Add, so every linear output carried 7 extra copies of its bias. conv2d was unaffected (bias via memset), which is why the trunk matched zig-solver to 5e-6 while fc/head activations diverged — the 81% labels regression. Also adds a scalar tail for in_n < 8 (pose_fc has in_n=4) and a known-answer selftest (test_linear). 👾 Generated with [Letta Code](https://letta.com) Co-Authored-By: Letta Code <noreply@letta.com>
crucible
CPU-only inference library for PyTorch .pt state dicts, in pure Zig.
No Python, no venv, no torch — the archive reader and the math both live here.
const crucible = @import("crucible");
var sd = try crucible.loadStateDict(alloc, io, "model.pt");
defer sd.deinit();
const weights = try (sd.get("trunk.0.weight") orelse return error.Missing).toF32(alloc);
try crucible.layers.conv2d(alloc, &input, h, w, in_c, weights, bias, &out, out_c);
What it is
- Restricted pickle VM — decodes the data-loading subset of pickle
(protocol 2) used by
.ptstate dicts. GLOBAL resolution is whitelisted: it reads tensors, storages and OrderedDicts, and refuses everything else. Unliketorch.load, it structurally cannot execute arbitrary Python. - ZIP container reader — store + raw-deflate entries via
std.zip. - Tensor views — dtype (f32/f64/i64/i32/u8), storage offset, sizes, strides; contiguous and strided materialization to f32.
- Layers (CPU, inference) —
conv2d(f32x8 FMA over output width),linear,relu,maxpool2,adaptiveAvgPool2d,softmax,padInput.
What it is not
- Not a training library. No autograd, no CUDA, no NPU.
- Not a full pickle implementation — unsupported opcodes and globals are errors, not imports.
- Not a model format converter. state_dict-style checkpoints only.
Usage
Add to your build.zig.zon:
zig fetch --save git+https://git.chaosmith.systems/pierre/crucible
const crucible = b.dependency("crucible", .{});
exe_mod.addImport("crucible", crucible.module("crucible"));
See examples/stripsolver.zig for a complete CNN: loads a real trained
.pt, runs a 3×conv + pose-conditioned multi-head forward.
Performance
On a 132k-parameter CNN (44×100 input, 3× conv3x3, pose-conditioned
4-head FC): ~5 ms per forward pass with f32x8 FMA (-Dcpu=x86_64_v3),
2.9 MB RSS. Validated at 99.95% digit accuracy against the reference
PyTorch implementation — identical predictions.
Status
v0.1 — working for real state dicts (Conv2d/Linear/MaxPool2d/ AdaptiveAvgPool2d/ReLU/Softmax). Layer set grows on demand. AVX-512 path: when hardware that has it does.
Description
An inference runtime for loading Pytorch weights (.pt) and running efficient inference via native CPU (AVX2+ required)
55 KiB
0 Stars
1 Watchers
0 Forks
Languages
Zig
100%