crucible v0.1: CPU inference library for PyTorch .pt state dicts, in pure Zig

- restricted pickle VM (whitelisted globals — cannot execute arbitrary
  Python, unlike torch.load)
- zip container reader (store + raw-deflate)
- Tensor views: dtype/offset/sizes/strides, f32 materialization
- layers: conv2d f32x8 FMA, linear, relu, maxpool2, adaptiveAvgPool2d,
  softmax, padInput
- examples/stripsolver: real .pt forward, 3777 @ 1.0

👾 Generated with [Letta Code](https://letta.com)

Co-Authored-By: Letta Code <noreply@letta.com>
This commit is contained in:
pierreandLetta Code committed 2026-09-29 18:41:46 +03:00
commit d401f6b07e
11 files changed
+1228

No files matched your search

+62
View File
@@ -0,0 +1,62 @@
# crucible
CPU-only inference library for PyTorch `.pt` state dicts, in pure Zig.
No Python, no venv, no torch — the archive reader and the math both live here.
```zig
const crucible = @import("crucible");
var sd = try crucible.loadStateDict(alloc, io, "model.pt");
defer sd.deinit();
const weights = try (sd.get("trunk.0.weight") orelse return error.Missing).toF32(alloc);
try crucible.layers.conv2d(alloc, &input, h, w, in_c, weights, bias, &out, out_c);
```
## What it is
- **Restricted pickle VM** — decodes the data-loading subset of pickle
(protocol 2) used by `.pt` state dicts. GLOBAL resolution is whitelisted:
it reads tensors, storages and OrderedDicts, and refuses everything else.
Unlike `torch.load`, it structurally cannot execute arbitrary Python.
- **ZIP container reader** — store + raw-deflate entries via `std.zip`.
- **Tensor views** — dtype (f32/f64/i64/i32/u8), storage offset, sizes,
strides; contiguous and strided materialization to f32.
- **Layers (CPU, inference)** — `conv2d` (f32x8 FMA over output width),
`linear`, `relu`, `maxpool2`, `adaptiveAvgPool2d`, `softmax`, `padInput`.
## What it is not
- Not a training library. No autograd, no CUDA, no NPU.
- Not a full pickle implementation — unsupported opcodes and globals are
errors, not imports.
- Not a model format converter. state_dict-style checkpoints only.
## Usage
Add to your `build.zig.zon`:
```sh
zig fetch --save git+https://git.chaosmith.systems/pierre/crucible
```
```zig
const crucible = b.dependency("crucible", .{});
exe_mod.addImport("crucible", crucible.module("crucible"));
```
See `examples/stripsolver.zig` for a complete CNN: loads a real trained
`.pt`, runs a 3×conv + pose-conditioned multi-head forward.
## Performance
On a 132k-parameter CNN (44×100 input, 3× conv3x3, pose-conditioned
4-head FC): ~5 ms per forward pass with f32x8 FMA (`-Dcpu=x86_64_v3`),
2.9 MB RSS. Validated at 99.95% digit accuracy against the reference
PyTorch implementation — identical predictions.
## Status
v0.1 — working for real state dicts (Conv2d/Linear/MaxPool2d/
AdaptiveAvgPool2d/ReLU/Softmax). Layer set grows on demand.
AVX-512 path: when hardware that has it does.