How llama.cpp GEMV Kernel Works on CUDA
A walkthrough of the FP16 CUDA GEMV kernel used for batch-size-one LLM decoding.
Notes on machine learning, efficient LLMs, and CUDA.
A walkthrough of the FP16 CUDA GEMV kernel used for batch-size-one LLM decoding.