TESLA LIVE 24.com · 24시간 반도체·SOXL 시세·속보 한국어 실시간

AMD, RTX Spark 출시 예상 앞서 '고르곤 할로' 벤치마크로 선점 시도

Tom's Hardware · 2026.10.06 00:38 · 원문 사이트
🇰🇷 한글 번역
AMD는 예상되는 RTX 스파크(RTX Spark) 출시보다 앞서 고르곤 할로(Gorgon Halo) AI 벤치마크를 통해 선점하려는 움직임을 보이고 있다. 동사는 지금까지 ‘50만 대 이상의 에이전트 PC(agentic PCs)’를 출하했다고 밝혔다. AMD는 에이전트 PC용 대형 SoC(시스템 온 칩)도 제조한다는 점을 상기시키고자 한다. 엔비디아는 지난 주 말 다소 노골적인 티저와 칩의 가을 출시 발표에 이어 이번 주 후반 RTX 스파크 장치를 출시할 것으로 예상된다. 수요일(10월 7일) 출시가 예상되는 마이크로소프트의 행사에 앞서 AMD는 RTX 스파크 장치와 직접 경쟁할 새로운 고르곤 할로 칩의 일부 벤치마크를 공유했다. 이 같은 추가적인 정보는 일부 기기가 7,099달러 이상에 판매되는 등 첫 고르곤 할로 장치 출시 며칠 만에 제공된 것이다. 의심스러운 기크벤치(Geekbench) 유출 몇 건을 제외하면, 우리는 아직 RTX 스파크의 성능 결과를 보지 못했으므로 AMD는 이를 비교 대상으로 사용하지 않는다. 대신 AMD는 라이젠 AI 맥스+ 프로 495(Ryzen AI Max+ Pro 495)를 인텔의 코어 울트라 X9 388H(Core Ultra X9 388H)와 비교한다. 이 칩들은 동일한 범주의 장치가 아니지만, AMD는 자사의 최상위 스택 제품을 인텔의 최상위 스택 제품과 비교한다고 주장할 것이다. AMD는 생성형 AI 성능을 측정하기 위해 ComfyUI를 사용했다. AMD는 다양한 모델의 여러 실행을 평균내어 총 처리량을 인텔의 경쟁 제품과 비교했다. AMD는 라이젠 AI 맥스+ 프로 495의 최고 사양인 192GB 구성을 사용했으며, 64GB 메모리가 탑재된 시스템의 코어 울트라 X9 388H와 비교했다(팬서 레이크는 최대 128GB까지 지원한다). 성능 우위는 1.1배에서 32.2배까지 다양하지만, 최상단 수치는 명확한 아웃라이어(이상치)다. 우리는 허깅 페이스(Hugging Face)에서 ‘유브(Yuve)’를 검색했지만 아무런 결과도 찾지 못했다. 이 격차는 최적화 문제 때문일 수도 있고, 팬서 레이크 머신에서 실행하기에는 단순히 너무 크기 때문일 수도 있다. 우리는 게임이나 일반 애플리케이션 성능에 대한 데이터를 가지고 있지 않지만, 이전 세대 스트릭 할로(Strix Halo) 칩과 비교해 큰 변화가 있을 것으로 예상하지는 않는다. 고르곤 할로는 해당 범주의 대부분을 새로 고친 것이다. 다만 고르곤 할로는 팬서 레이크와 경쟁하는 것은 아니다. 고르곤 할로는 주로 RTX 스파크와 경쟁하며, 그 다음으로는 애플의 대형 M 시리즈 SoC들과 경쟁한다. AMD가 지칭하는 에이전트 PC라는 범주는 확장되고 있는 것으로 보이지만, RTX 스파크에 대한 일부 과장된 홍보가 시사하는 것보다는 여전히 훨씬 작다. AMD는 기자간담회에서 사전 브리핑을 통해 AI PC(노트북을 의미함)를 ‘수천만 대’ 출하했다고 자랑했으나, 이후 ‘50만 대 이상’의 에이전트 PC를 출하했다고 정정했다. 추정컨대 이 수치는 스트릭/고르곤 할로 장치에 대한 것이다. 지금까지 고르곤 할로와 RTX 스파크의 대결은 주로 메모리 용량에 초점을 맞춰져 왔다. RTX 스파크 장치는 DGX 스파크의 GB10(두 칩은 거의 동일함)과 동일한 128GB의 통합 메모리로 최대치를 기록한다. 반면 AMD는 고르곤 할로로 최대 192GB를 지원한다. 더 높은 용량은 더 큰 모델을 로컬에서 실행할 수 있게 해 주지만, 성능 수준은 낮아진다. GLM 5.3 Flash(3,200억 파라미터)에서 AMD는 최고 처리량 1초당 20개의 토큰을 기록했다. 다만 이는 Unsloth의 UD-IQ4_XS 혼합 양자화 형식을 사용했을 때의 수치다. 참고로, 우리는 4비트 양자화를 적용한 GPT-OSS 120B를 실행하는 DGX 스파크에서 최고 토큰 처리량 1초당 64개의 토큰을 기록했으며, 이전 세대 라이젠 AI 맥스+ 395는 1초당 56개의 토큰을 기록했다. (이 부분 번역 실패)
📄 원문 (English)
AMD attempts to get ahead of expected RTX Spark launch with Gorgon Halo AI benchmarks — company says it has shipped 'over half a million agentic PCs' to date AMD would like to remind you that it makes a big SoC for agentic PCs, too. Nvidia is expected to launch RTX Spark devices later this week, following a not-so-subtle tease at the end of last week (and the chip’s announced fall release). Ahead of Microsoft’s event on Wednesday, October 7 (where the launch is expected) AMD has shared some benchmarks for its new Gorgon Halo chips — the range that will directly compete with RTX Spark devices. The extra insight comes a matter of days after the first Gorgon Halo devices launched, some of which cost upwards of $7,099. Short of a few questionable Geekbench leaks, we haven’t seen any performance results for the RTX Spark yet, so AMD isn’t using it as a comparison point. Rather, it’s comparing the Ryzen AI Max+ Pro 495 to Intel’s Core Ultra X9 388H. These chips aren’t in the same class of device, though AMD would argue that it’s comparing its top-of-stack part to Intel’s top-of-stack part. AMD used ComfyUI to measure generative AI performance. AMD averaged multiple runs of various models, comparing total throughput to Intel’s competition. AMD used its top-spec 192GB configuration of the Ryzen AI Max+ Pro 495 and compared it to the Core Ultra X9 388H in a system with 64GB of memory (Panther Lake supports up to 128GB). The performance advantage ranges from 1.1x up to 32.2x, though the end point is a clear outlier. We searched for Yuve on Hugging Face and didn’t find any results. It’s possible this delta comes down to an optimization issue, or that it’s simply too big to run on the Panther Lake machine. We don’t have gaming or general application performance, but we don’t expect a major swing compared to last-gen Strix Halo chips. Gorgon Halo is largely a refresh of that range. Gorgon Halo isn’t getting into the ring with Panther Lake, however. It’s going mainly against the RTX Spark, and to a lesser extent, Apple’s larger M-series SoCs. This category of agentic PCs, as AMD calls it, is seemingly expanding, though it’s still far smaller than some of the hype around the RTX Spark would have you believe. AMD bragged in a prebriefing with the press about shipping “10s of millions” of AI PCs (read: laptops) before clarifying that it had slipped “over half a million” agentic PCs. Presumably, those are numbers for Strix/Gorgon Halo devices. The matchup between Gorgon Halo and the RTX Spark has, up to this point, focused mainly on memory capacity. RTX Spark devices top out at 128GB of unified memory, same as the GB10 in the DGX Spark (the two chips are nearly identical). AMD, on the other hand, supports up to 192GB with Gorgon Halo. Higher capacity means running larger models locally, though at a lower performance level. Get Tom's Hardware's best news and in-depth reviews, straight to your inbox. In the GLM 5.3 Flash with 320 billion parameters, AMD saw peak throughput of 20 tokens per second, though using Unsloth's UD-IQ4_XS mixed-quantization format. For context, we clocked peak token throughput on the DGX Spark running GPT-OSS 120B with 4-bit quantization at 64 tokens per second, and the last-gen Ryzen AI MAx+ 395 at 56 tokens per second. Note that although GLM 5.3 Flash has 320 billion total parameters, only 18 billion are activated for each token. Similarly, GPT-OSS 120B is another mixture-of-experts model with 120 billion total parameters, though only around 5 billion are active per token. AMD also shared Qwen 3.8 Flash Next performance, a multimodal MoE model with 125 billion main parameters, 51 billion embedding parameters, and about 4 billion parameters for multi-token prediction (MTP), with 5 billion parameters active per token. AMD says the Ryzen AI Max+ Pro 495 achieves up to 42 tokens per second with this model, once again using Unsloth’s dynamic 4-bit quantization and MTP. That’s solid performance, but in both cases, “up to” carries a lot on its shoulders. The story of token throughput is told as the context length increases, showcasing what happens when someone actually runs these models locally, not just boots them up cold. Performance drops at higher context lengths, naturally, which could pose some issues for the larger models. If GLM 5.3 Flash provides up to 20 tokens per second, it could very easily decline into unusable territory as the context length increases. We largely know what performance to expect out of the Ryzen AI Max+ Pro 495, and Gorgon Halo more broadly. It’s a refresh of Strix Halo, with notable spec changes being the bump up to 192GB of unified memory from 128GB, as well as a 100 MHz jump on boost clocks for the 495. Otherwise, the range is using identical core counts and microarchitectures as previous-gen Strix Halo chips. With these proxies — Strix Halo for Gorgon Halo, and GB10 for RTX Spark — we can already get a good idea about how these parts will stack up. As you can see in our Ryzen AI Halo review (packing the Ryzen AI Max+ 395), AMD’s part universally underperformed compared to the DGX Spark in both time to first token and tokens per second across three models. More unified memory will allow you to run larger models, but that doesn’t mean those models will run faster. The first Gorgon Halo devices are available for sale now, such as the Minisforum MS-S1 Max-P495. For the top-line configuration, prices sit around $7,000 right now, though we expect a broad range of prices once different devices are available, likely driving above that $7,000 mark. Full presentation View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original View Original Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds. Jake Roach is the Senior CPU Analyst at Tom’s Hardware, writing reviews, news, and features about the latest consumer and workstation processors.
▶ 칩 라이브 24 홈으로 — chiplive24.com

다른 반도체 뉴스

The 552B DeepSeek V4.1-Flash Model Offers A Peak Output Of 494 Tokens/Second When Powered By An At-Home Rig Spanning 4x NVIDIA DGX Spark Units · WccftechSamsung May Report Its First 100 Trillion Won Quarter. Micron Stock Has More to Lose Than to Gain. · Yahoo FinanceAltman-backed Volantis reveals plan to vault the memory wall by baking photonics into AI accelerators · The RegisterBroadcom Is Up More Than 25X in 10 Years. Can AI Drive the Next Chapter? · Yahoo FinanceTSMC stock hits all-time high after Elon Musk confirms early Terafab talks · Yahoo FinanceChip Stocks Pause After Two-Day Pop. AMD Stock Gets Price-Target Hikes. · Yahoo FinanceIntel's Googlebooks might not run some Android apps as well as Qualcomm's — Google says this is because Android apps were designed for Arm chips · Tom's HardwareNvidia Generates $7.2 Million per Employee, Up 500% in 3 Years. This Number Is Even More Astounding. · Yahoo Finance