
Local inference · September 2026
Kidi: lightweight native AI inference
I built a C++23 inference toolkit for CPU and Apple Metal, with an experimental WebAssembly port. It runs Gemma text generation and RTG translation without PyTorch or Transformers at inference time.
The experimental browser demo is my first attempt at WebAssembly; it's still slow at the moment. CPU-only, with no WebGPU acceleration. Prompts stay on your device. About 2.49 GB of model files, cached after download; current 64-bit Chromium and ample memory recommended.

