Files
pierreandLetta Code d401f6b07e crucible v0.1: CPU inference library for PyTorch .pt state dicts, in pure Zig
- restricted pickle VM (whitelisted globals — cannot execute arbitrary
  Python, unlike torch.load)
- zip container reader (store + raw-deflate)
- Tensor views: dtype/offset/sizes/strides, f32 materialization
- layers: conv2d f32x8 FMA, linear, relu, maxpool2, adaptiveAvgPool2d,
  softmax, padInput
- examples/stripsolver: real .pt forward, 3777 @ 1.0

👾 Generated with [Letta Code](https://letta.com)

Co-Authored-By: Letta Code <noreply@letta.com>
2026-09-29 18:41:46 +03:00

2.2 KiB
Raw Permalink Blame History

crucible

CPU-only inference library for PyTorch .pt state dicts, in pure Zig. No Python, no venv, no torch — the archive reader and the math both live here.

const crucible = @import("crucible");

var sd = try crucible.loadStateDict(alloc, io, "model.pt");
defer sd.deinit();

const weights = try (sd.get("trunk.0.weight") orelse return error.Missing).toF32(alloc);
try crucible.layers.conv2d(alloc, &input, h, w, in_c, weights, bias, &out, out_c);

What it is

  • Restricted pickle VM — decodes the data-loading subset of pickle (protocol 2) used by .pt state dicts. GLOBAL resolution is whitelisted: it reads tensors, storages and OrderedDicts, and refuses everything else. Unlike torch.load, it structurally cannot execute arbitrary Python.
  • ZIP container reader — store + raw-deflate entries via std.zip.
  • Tensor views — dtype (f32/f64/i64/i32/u8), storage offset, sizes, strides; contiguous and strided materialization to f32.
  • Layers (CPU, inference) — conv2d (f32x8 FMA over output width), linear, relu, maxpool2, adaptiveAvgPool2d, softmax, padInput.

What it is not

  • Not a training library. No autograd, no CUDA, no NPU.
  • Not a full pickle implementation — unsupported opcodes and globals are errors, not imports.
  • Not a model format converter. state_dict-style checkpoints only.

Usage

Add to your build.zig.zon:

zig fetch --save git+https://git.chaosmith.systems/pierre/crucible
const crucible = b.dependency("crucible", .{});
exe_mod.addImport("crucible", crucible.module("crucible"));

See examples/stripsolver.zig for a complete CNN: loads a real trained .pt, runs a 3×conv + pose-conditioned multi-head forward.

Performance

On a 132k-parameter CNN (44×100 input, 3× conv3x3, pose-conditioned 4-head FC): ~5 ms per forward pass with f32x8 FMA (-Dcpu=x86_64_v3), 2.9 MB RSS. Validated at 99.95% digit accuracy against the reference PyTorch implementation — identical predictions.

Status

v0.1 — working for real state dicts (Conv2d/Linear/MaxPool2d/ AdaptiveAvgPool2d/ReLU/Softmax). Layer set grows on demand. AVX-512 path: when hardware that has it does.