Skip to content
Mission Ascent · Pre-launch
T–00:00:10
Loading assets0%
Mohit Agarwal
Research

Local LLM Tracing & Replay

See inside a local LLM while it runs: per-layer timing, live attention heatmaps and numerical errors, for any GGUF model, without touching the model.

When
Jun 2026 · GDSC
Crew
Solo
Domain
AI · Systems
Stack
C++17 · llama.cpp · FTXUI · CMake
LT-03 · TelemetrySim
ATTENTION · L12 · H07CAUSAL MASK · KQ_SOFT_MAXLAYER LATENCY · MS● ANOMALY LEDGER · REPLAY READY
5
Model families
0
Model changes
C++17
~2.2k lines

The problem

Running Llama or Mistral locally is easy. Seeing what happens inside, which layer is slow, where attention goes, where a NaN first appears, is not.

What I built

A tracer that hooks into llama.cpp's graph callback, so it works with any GGUF model (Llama, Mistral, Qwen, Gemma, Phi) without changing a line of model code. It times every layer, pulls out attention matrices and shows them as a live heatmap you can switch head by head, and logs NaN, Inf and outlier activations.

Runs can be recorded to a binary trace and replayed later without the model. Everything shows up in a five-panel terminal UI with vim keys.

The hard part

Collecting events from inside inference without slowing it down. I used a lock-free ring buffer so the hot path never waits on the UI.

Result

About 2,200 lines of C++17 in six days, in eleven small commits, building on Linux and Windows.