[ SESSION DETAILS / 세션 주요 내용 ]
Advancing LLM Inference Performance for Edge/Cloud
Generative AI introduces significant challenges for system design, primarily due to the memory-bound nature of large language model (LLM) inference. This talk explores LPDDR-based Processing-in-Memory (PIM) technology, which enables in-memory execution of GEMV operations to substantially improve both performance and energy efficiency.
엣지 및 클라우드를 위한 LLM 추론 성능 향상
생성형 AI는 대규모 언어모델(LLM)의 추론 과정이 메모리 성능에 크게 의존하기 때문에 시스템 설계에 많은 어려움을 가져오고 있습니다.
본 발표에서는 LPDDR 기반 PIM(Processing-in-Memory) 기술을 소개합니다. 이 기술은 GEMV(행렬-벡터 곱) 연산을 메모리 내부에서 수행함으로써
성능과 에너지 효율을 크게 향상시키는 방법을 설명합니다.
BACK TO LIST