TESLA LIVE 24.com · 24시간 반도체·SOXL 시세·속보 한국어 실시간

4대 NVIDIA DGX Spark 장비로 구동되는 홈 서버에서 DeepSeek V4.1-Flash 모델의 초당 최대 494 토큰 출력

Wccftech · 2026.10.06 01:46 · 원문 사이트
🇰🇷 한글 번역
자사의 독점 데이터 보안을 최대한의 충실도로 보장하려면, 자체 환경에 강력한 오픈 가중치 모델을 호스팅하는 것이 여전히 최선의 선택입니다. 최적화된 환경은 비교적 큰 모델이라도 비교할 수 없는 처리량을 제공합니다. 실제로 5,520억 파라미터를 가진 DeepSeek V4.1-Flash 모델은 최근 4대의 NVIDIA DGX Spark 유닛으로 구성된 가정용 시스템에서 호스팅될 때 Cerebras급 추론 처리량을 제공할 수 있었습니다. 광학 라우터나 이더넷 스위치가 없는 4노드 NVIDIA DGX Spark 클러스터는 초당 약 500개의 토큰에 가까운 최대 출력 성능을 발휘했습니다. AMD의 베테랑이자 기술 전문가인 패트릭 무어헤드는 4대의 NVIDIA DGX Spark 유닛으로 구성된 자신의 가정용 AI 설정을 공개했는데, 이는 약 512GB의 통합 메모지를 제공하는 유닛들을 하단에 추가 냉각 팬이 달린 컴팩트한 'tec MOJO' 랙에 쌓아 올린 형태입니다. 여기서 사용된 각 NVIDIA DGX Spark 시스템의 구성 요소는 다음과 같습니다. - 20코어 Arm CPU (10개의 고성능 Cortex-X925 코어 + 10개의 효율성 Cortex-A725 코어). 각 코어는 전용 L2 캐시를 가지며, 클러스터당 16MB의 L3 캐시를 공유하므로 총 32MB입니다. - 5세대 Tensor 코어와 DLSS 4 지원, RTX 레이트레이싱 코어를 탑재한 GB10 iGPU. AI 워크로드를 위해 최대 31 TFLOPs의 FP32 연산 성능과 1000 TOPS의 NVFP4(FP4) 연산 성능을 제공합니다. - CPU와 GPU가 공유하는 128GB의 일관성 있는 통합 LPDDR5X 시스템 메모리. 대역폭은 약 273GB/s(256비트 메모리 인터페이스)입니다. - 고속 클러스터링을 위한 듀얼 200Gb/s QSFP 포트를 갖춘 ConnectX-7 스마트NIC, 그리고 10GbE, Wi-Fi 7, 블루투스, HDMI 2.1a, 여러 USB-C 포트를 포함합니다. - 일반적인 저장 용량: 1~4TB NVMe SSD. - 제공된 어댑터 사용 시 전력 소모는 약 140~240W 범위입니다. - 전체 CUDA AI 소프트웨어 스택을 갖춘 NVIDIA DGX OS(Ubuntu 기반). 여기서 주목해야 할 중요한 점은, 4대의 NVIDIA DGX Spark 유닛이 스위치 없는 링 토폴로지를 통해 QSFP 케이블로 서로 연결되었다는 것입니다. - Spark 1은 Spark 2와 Spark 4에 연결됩니다. - Spark 2는 Spark 1과 Spark 3에 연결됩니다. - Spark 3은 Spark 2와 Spark 4에 연결됩니다. - Spark 4는 Spark 3과 Spark 1에 연결됩니다. NVIDIA는 4대의 DGX Spark 유닛을 연결할 때 200GbE 스위치(스타 토폴로지) 사용을 권장합니다. 물론 스위치는 전체 시스템의 비용 증가 요인이 됩니다. 한편, 무어헤드는 현명한 결정을 내려 자신의 4노드 NVIDIA DGX Spark 랙에 고효율인 DeepSeek V4.1-Flash 모델을 호스팅했습니다. 최근 별도의 게시글에서 상세히 설명한 바와 같이, 이 모델은 5,520억 파라미터를 보유하고 있으며, 프리필(prefill) 단계에서는 80억 개의 활성 파라미터, 디코드(decode) 단계에서는 160억 개의 활성 파라미터를 사용합니다. 또한 이 모델은 1개의 공유 전문가(shared expert)와 384개의 라우팅된 전문가(routed experts, 6개 활성)를 갖추고 있습니다. 또한 DeepSeek의 최신 모델은 인과적 인코더-디코더(Causal Encoder-Decoder, CED), 압축 희소 어텐션 2(Compressed Sparse Attention 2, CSA2), FP4 양자화, SWA 경량 재생(SWA Bounded Replay) 등 여러 아키텍처 요소를 통합하여 토큰당 KV 캐시를 단 890바이트로 줄였습니다. 4노드 NVIDIA DGX Spark 시스템과 DeepSeek의 V4.1-Flash 모델을 결합한 결과, 성능은 다음과 같습니다. - 최대 출력: 494 tok/s (코드, 32개 동시 요청). - 산문 최대: 280 tok/s (32개 동시). - 단일 요청: 산문 약 58 tok/s, 코드 약 96 tok/s. - 프롬프트 처리: 약 4,764 tok/s. - 첫 번째 토큰: 약 0.2초 (유휴 상태). - 메모리 사용량: 476GB 이러한 구성의 비용과 관련하여, 현재 128GB NVIDIA DGX Spark 유닛의 소매가는 저장 용량과 가용성에 따라 $5,500에서 $9,000 이상까지 다양합니다. 따라서 이 시스템의 비용은 다음과 같이 구성됩니다. - 128GB 유닛 4대: 일반적으로 $22,000~$36,000 - 링 구성을 위한 QSFP 직접 연결 케이블: 수백 달러 - 랙/선반, 추가 팬, 전력 분배 장치 및 기타 잡다한 항목: 추가로 수백 달러 이를 통해 시스템의 총 비용은 $25,000에서 $40,000 사이에 있을 것으로 예상됩니다. 이는 상당히 비싼 금액이지만, 이 구성은 24시간 내내 가동할 수 있으며 각 유닛이 부하 상태에서 약 400~600W의 전력을 소모하므로 전기 요금은 최소한으로만 발생합니다. 또한 OpenAI나 Anthropic의 구독 서비스와 달리 API 비용이나 사용량 제한에 노출되지 않는다는 장점이 있습니다. 더 많은 뉴스 coverage를 피드에서 확인하려면 Wccftech를 Google에서 팔로우하세요.
📄 원문 (English)
If you want to ensure the security of your proprietary data with near-total fidelity, hosting a capable open-weight model on your own setup is still the best course of action, with optimized setups offering unmatched throughput even for relatively large models. In fact, the 552-billion-parameter DeepSeek V4.1-Flash was recently able to offer a Cerebras-class inference throughput when hosted on an at-home rig made up of 4x NVIDIA DGX Spark units. The 4-node NVIDIA DGX Spark cluster, sans an optical router or an Ethernet switch, was able to offer a peak output of nearly 500 tokens per second The AMD veteran and tech expert, Patrick Moorhead, has just disclosed his at-home AI setup, which consists of 4x NVIDIA DGX Spark units - offering a pooled unified memory of ~512 GB - stacked inside a compact "tec MOJO" rack with additional cooling fans underneath. For the benefit of those who might not be aware, each NVIDIA's DGX Spark system used here consists of: - A 20-core Arm CPU (10 high-performance Cortex-X925 cores + 10 efficiency Cortex-A725 cores, where each core has a private L2 cache and a 16 MB L3 cache per cluster, so 32 MB in total). - The GB10 iGPU that features fifth-gen Tensor Cores with DLSS 4 support and RTX Ray Tracing cores, offering up to 31 TFLOPs of FP32 and 1000 TOPS of NVFP4 (FP4) compute for AI workloads. - A 128 GB coherent unified LPDDR5X system memory (shared between CPU and GPU) with ~273 GB/s bandwidth (256-bit memory interface). - ConnectX-7 SmartNIC with dual 200 Gb/s QSFP ports for high-speed clustering, plus 10 GbE, Wi-Fi 7, Bluetooth, HDMI 2.1a, and multiple USB-C ports. - Typical storage of 1–4 TB NVMe SSD. - Power draw in the ~140–240 W range with the included adapter. - NVIDIA DGX OS (Ubuntu-based) with the full CUDA AI software stack. In what is a critical point to note here, all four NVIDIA DGX Spark units are connected to each other via the switchless ring topology using the QSFP cables: - Spark 1 connects to Spark 2 and Spark 4 - Spark 2 connects to Spark 1 and Spark 3 - Spark 3 connects to Spark 2 and Spark 4 - Spark 4 connects to Spark 3 and Spark 1 Do note that NVIDIA recommends using a 200 GbE switch (star topology) to connect 4x DGX Spark units. Of course, the switch adds to the overall cost of the rig. Meanwhile, in what was a smart decision, Moorhead chose to host the highly efficient DeepSeek V4.1-Flash model on his 4-node NVIDIA DGX Spark rack. As we detailed in a dedicated post recently, the model spans 552 billion parameters, with just 8 billion active parameters active at the prefill stage and 16 billion at the decode stage. Additionally, the model has 1 shared expert and 384 routed experts (6 active). Also, DeepSeek's latest model incorporates several architectural elements - including Causal Encoder-Decoder (CED), Compressed Sparse Attention 2 (CSA2), FP4 Quantization, and SWA Bounded Replay - that reduce its overall KV cache to just 890 bytes per token! Given this pairing of a 4-node NVIDIA DGX Spark rig with DeepSeek's V4.1-Flash model, the results speak for themselves: - Peak output: 494 tok/s (code, 32 concurrent requests). - Prose peak: 280 tok/s (32 concurrent). - Single-request: ~58 tok/s prose, ~96 tok/s code. - Prompt processing: ~4,764 tok/s. - First token: ~0.2 s (idle). - Memory footprint: 476 GB As for the cost of this setup, each 128 GB NVIDIA DGX Spark unit currently retails for anywhere between $5,500 and $9,000+ depending on the storage capacity and availability. As such, the rig's cost includes: - Four 128 GB units, typically $22,000–$36,000. - QSFP direct-attach cables for a ring: a few hundred dollars. - Rack/shelf, extra fans, power distribution, and miscellaneous: a few hundred dollars more. This means that the rig's total cost would fall anywhere between $25,000 and $40,000. While this is quite pricey, you could run this setup 24/7, incurring only minimal electricity charges (each unit consumes roughly 400–600 W under load), and without being subject to API costs and load limits, as would be the case with a subscription from OpenAI or Anthropic. Follow Wccftech on Google to get more of our news coverage in your feeds.
▶ 칩 라이브 24 홈으로 — chiplive24.com

다른 반도체 뉴스

Why Isn't Micron's Stock Taking Off After Reporting Strong Q4 Numbers? · Yahoo FinanceTaiwan Semiconductor Gains 2% to Record as Musk Confirms Early Terafab Talks; Broadcom Rises 2%, Qualcomm Pulls Back · Yahoo FinanceNvidia Is 'Tip Of The Spear' For AI Trade, Says Dan Niles — Warns 'At A Certain Point Either The Bond Market's Wrong Or Stock Market Is Wrong' · Yahoo FinanceNvidia Is 'Tip Of The Spear' For AI Trade, Says Dan Niles — Warns 'At A Certain Point Either The Bond Market's Wrong Or Stock Market Is Wrong' · TradingViewNvidia Is 'Tip Of The Spear' For AI Trade, Says Dan Niles — Warns 'At A Certain Point Either The Bond Market's Wrong Or Stock Market Is Wrong' · StocktwitsNvidia Spent $20 Billion on Buybacks Last Quarter. Here's What That Means for Your Shares. · Yahoo FinanceSamsung May Report Its First 100 Trillion Won Quarter. Micron Stock Has More to Lose Than to Gain. · Yahoo FinanceAltman-backed Volantis reveals plan to vault the memory wall by baking photonics into AI accelerators · The Register