Advancing LLM Inference Performance for Edge/Cloud<br />Generative AI introduces significant challenges for system design, primarily due to the memory-bound nature of large language model (LLM) inference. This talk explores LPDDR-based Processing-in-Memor