Selected work
From GPU kernels to autonomous agents and image compression. Systems I build, test, and operate across infrastructure, security, and product engineering.
LLM inference, down to the metal.
Custom kernels, speculative decoding, and end-to-end serving on AMD gfx906. An independent engineering project with performance measurements, correctness investigations, and documented failures.
Qwen3.6-35B-A3B · DFlash code-generation peak
Two 16-GiB Radeon Pro VII GPUs · Q4_0 / Router257
holdmybyte
A JPEGmini alternative I built around a different contract: choose a perceptual quality target, then verify every output image against its source at full resolution. The system combines CPU/GPU quality measurement, per-image encoder selection, and a search for smaller files that meet the target.
Retested on image sets from five weddings. The whitepaper documents the architecture and an initial 535-image wedding benchmark, including the operating points where JPEGmini remains faster or produces smaller files. Read the whitepaper (PDF) ↗
Local LLM Inference Cluster
Design and operation of a multi-node cluster with 23 AMD MI50 GPUs. Reusable benchmarks and persistent results make experiments comparable across runs. Read the case study →
artemis-mini
An autonomous penetration-testing loop with Docker sandboxing and structured logging. Isolated execution and recorded iterations make its behavior and failure states inspectable.
Prompt Injection Automation
Adversarial-evaluation tooling that generates, executes, and scores prompt-injection and jailbreak test cases against LLM systems in controlled environments.