퀄컴 CEO, 2028년까지 100B LLM 구동 스마트폰 등장 시사…DRAM 수급·가격은 언급 피함
🇰🇷 한글 번역
스마트폰에서 더 밀도 높은 AI 모델을 실행하는 데 있어 가장 큰 병목 현상은 메모리 용량이며, 현재 DRAM 상황을 볼 때 16GB의 문턱을 곧 넘기는 것은 어려워 보인다. 그러나 퀄컴의 크리스티아노 아몬 CEO는 핸드셋이 1000억 파라미터(100B) 규모의 대규모 언어 모델(LLM)을 지속적으로 실행할 수 있는 시대가 곧 도래할 것이라고 믿으며, 이러한 기능을 달성할 수 있는 스마트폰을 구축 중인 AI 기업들이 존재한다고 확인했다. 더 밀도 높은 모델을 실행하는 실용성에 관해 아몬은 공급이 부족하고 가격이 급등할 때 스마트폰 제조사들이 대량의 RAM을 탑재할 것이라고 언급했다.
2028년까지 스마트폰에서 100B LLM을 실행할 수 있는 유일한 방법은 플래시 메모리를 통해 스트리밍하거나 희소 MoE(전문가 혼합) AI 모델을 실행하는 것이다.
파이어사이드 알파와의 인터뷰에서 크리스티아노 아몬은 100B LLM을 실행할 수 있는 스마트폰을 원하는 AI 기업들의 이름을 공개하지 않았지만, 2028년에 이러한 AI 모델을 실행할 수 있는 기기가 등장한다는 그의 대담한 주장은 확실히 헤라클레스급의 과업이다. 메모리 위기로 인해 수많은 스마트폰 제조사들이 더 높은 RAM 변형 모델을 출시하는 것을 포기했지만, 불과 1년 전만 해도 24GB의 LPDDR5X 메모리를 자랑하는 안드로이드 기기들이 있었다.
계산을 해보더라도 4비트 양자화된 1000억 파라미터 모델을 실행하려면 스마트폰이 최소 50GB의 RAM을 가져야 하며, 운영체제, 앱, 컨텍스트 캐시를 위해 충분한 메모리가 남아있어야 한다. 2비트 양자화된 모델은 여전히 불가능한 과제로, 스마트폰은 100B LLM 전용으로 25GB의 메모리만 탑재해야 하기 때문이다. DRAM의 비용과 공급 장벽도 함께 고려하더라도, 더 밀도 높은 모델을 지속적으로 실행하면 열이 심하게 발생하여 열 스로틀링을 유발한다는 점은 아직 고려하지 않은 부분이다.
애플과 퀄컴은 A20 Pro와 스냅드래곤 8 엘리트 익스트림 제너레이션 6와 같은 칩셋 패키지의 전환을 통해 이 문제를 어느 정도 해결했지만, 100B 모델을 진지하게 실행하기 위한 더 독창적인 접근 방식이 취해지지 않는 한 실제 메모리 병목 현상은 계속될 것이다. 예를 들어, 공격적인 양자화는 메모리 요구 사항을 급격히 줄여 1비트 변형이 DRAM에 맞출 수 있게 하지만, 불행히도 4비트 양자화 이후에는 특히 추론 작업에서 LLM의 품질이 현저히 저하된다.
또 다른 대안은 전문가 오프로딩이 있는 MoE(전문가 혼합)로, 토큰당 소수의 전문가만 활성화되어 RAM에 남아있는 동안 나머지는 플래시 저장소에 위치한다. 플래시 저장소에 관해 말하자면, 애플은 온보드 메모리에서 LLM을 실행하는 연구를 진행해 왔지만, UFS 및 NVMe 저장소가 DRAM보다 느리기 때문에 추론 속도가 떨어질 수 있다.
스마트폰에서 실행될 때 200억에서 300억 파라미터 규모의 증류된 모델은 100B 모델의 품질에 근접할 수 있다. 이는 가장 현실적인 전략이기 때문에 기업들에게 더 마케팅하기 쉬운 전략으로 보인다. 불행히도 퀄컴 CEO가 로드맵을 제시하지 않고 이러한 발언을 하는 것은 스마트폰에서 더 밀도 높은 모델을 실행하는 데 우리가 더 가까워지게 해주지 못하지만, 이러한 기기들이 더 큰 기계들과 동일한 기능을 갖추게 된다는 상상은 여전히 즐거운 생각이다.
구글에서 Wccftech를 팔로우하여 더 많은 뉴스 커버리지를 피드에서 받아보십시오.
📄 원문 (English)
The biggest bottleneck of running denser AI models on smartphones is the memory count, and looking at the DRAM situation, it’s unlikely that we’ll cross the 16GB threshold soon. However, Qualcomm CEO Cristiano Amon believes that we’ll soon be living in an era where handsets can run 100B LLMs continuously and confirms that AI companies are building smartphones that can achieve this feat. As for the practicality of running denser models, Amon has mentioned how phone makers will incorporate hefty amounts of RAM, especially when supply is scarce, and prices are through the roof.
The only ways to run 100B LLMs on smartphones by 2028 are streaming through flash memory or running sparse MoE AI models
Speaking with Fireside Alpha, Cristiano Amon didn’t disclose any names of AI companies wanting smartphones that can run 100B LLMs, but his bold claim that devices being able to run these AI models arriving in 2028 is definitely a Herculean task. Due to the memory crisis, a plethora of smartphone makers have stepped away from introducing higher RAM variants, even though just a year ago, we had Android devices boasting a whopping 24GB of LPDDR5X memory.
Even if we do the math, running a 4-bit-quantized 100-billion-parameter model would require a smartphone to have at least 50GB of RAM, with sufficient memory left for the OS, apps, and context cache. A 2-bit quantized model is still an impossible feat, since smartphones would be required to feature 25GB of memory solely for the 100B LLM. Assuming the DRAM cost and supply obstacles are also scaled, we haven’t even taken into account that running denser models continuously contributes immensely to heat, causing thermal throttling.
Apple and Qualcomm have somewhat addressed this problem by a shift in chipset packaging with the A20 Pro and Snapdragon 8 Elite Extreme Gen 6, but the actual memory bottleneck will continue to exist unless more unique approaches are taken to seriously run 100B models. For example, aggressive quantization can sharply reduce memory requirements, meaning that 1-bit variants can fit in DRAM. Unfortunately, the LLM’s quality degrades significantly beyond 4-bit quantization, especially for reasoning tasks.
Another alternative is MoE (Mixture of Experts) with expert offloading, where only a small share of experts is active per token, so they stay in RAM, while the rest reside on flash storage. Speaking of flash storage, Apple has researched running LLMs on the onboard memory, but inference speeds can drop due to UFS and NVMe storage being slower than DRAM.
A distilled 20B to 30B model can approach the quality of a 100B when running on a smartphone. This appears as the more marketable strategy for companies as it’s the most realistic. Sadly, Qualcomm’s CEO making such statements without outlining a roadmap won’t get us any closer to running denser models on smartphones, but it’s still a nice thought to imagine these devices having the same capabilities as bigger machines.
Follow Wccftech on Google to get more of our news coverage in your feeds.
▶ 칩 라이브 24 홈으로 — chiplive24.com