samankeon.com

Hi, I'm Saman.

I'm an ML engineer at Meta working on LLM efficiency: quantization, inference optimization, and the GPU kernels underneath. Before that I spent a decade on backend and infrastructure work, from SRE and multi-cloud architecture to co-founding a trading company's software stack.

This site collects my writing. The blog covers how modern LLMs actually work under the hood (attention, MoE, RoPE, sampling, batching) and a hands-on series on writing Triton and CUDA kernels, with benchmarks. I also build MintEngine, an educational inference engine for checking model implementations layer by layer against production engines.

Recent posts

All posts →