Sithan Kanna
Distilling a Bigram
12 Sep 2026
Build a Bigram Model in 10 Minutes
30 Aug 2026
Memory to Train a Transformer
Draft · 11 Jul 2026
Estimating Layers in Transformers
27 Jun 2026
Why KV Cache and Not QKV Cache?
12 Jun 2026
Is RL Secretly Supervised Learning?
28 Dec 2025
Currently Reading — LLM-Generated Notes
Estimating the Compute and Memory Needed for Language Models
Draft · 11 Jul 2026
Also:
CEO Hour
— a timed thinking practice you can run on your own.