Heterogeneous Computing 1 What Porting llama2.c to an OpenCL Device Taught Me About LLM Inference Aug 12, 2026