TESLA LIVE 24.com · 24시간 반도체·SOXL 시세·속보 한국어 실시간

엔비디아 지원 Reflection, 4주간 10,500개 GB300 GPU와 일일 4,640만 개 샌드박스로 효율적인 5,010억 파라미터 오픈 가중치 모델 Beam 학습

Wccftech · 2026.10.06 09:32 · 원문 사이트
🇰🇷 한글 번역
서양의 AI 연구소들이 이제야 효율적인 오픈 가중치 모델 분야에서 중국의 딥시크(DeepSeek)와 즈이푸(Z.ai)와 경쟁하기 시작하고 있다. 그 한 예로 리플렉션(Reflection)의 빔(Beam)을 들 수 있는데, 이는 5,010억 개의 파라미터를 가진 희소 전문가 혼합(MoE) 모델로, 압도적인 효율성을 자랑한다. 리플렉션은 빔이 “GLM 5.2보다 3~4배 더 효율적이며, 주요 서양산 오픈 모델보다 4배 이상 효율적”이라고 주장한다. 여러분 중 이 사실을 잘 모르는 분들을 위해 설명하자면, 리플렉션의 빔은 5,010억 개의 파라미터를 갖추고 있지만, 동시에 활성화되는 파라미터는 단 230억 개에 불과하다. 전문가 혼합(MoE)을 이해하기 위해, 수많은 전문 셰프(라우팅된 전문가)를 보유한 대형 주방을 상상해 보라. 하지만 모델은 필요할 때만 전문 셰프들을 활성화한다. 예를 들어, 수플레를 준비하려면 모델은 동양 요리 전문가가 아니라 디저트 전문가만 활성화한다. 더욱이, AI 모델은 기본적으로 두 가지 핵심 요소로 구성된 일련의 트랜스포머 블록들로 이루어져 있다. - 어텐션 레이어(Attention Layer)는 문장 내 단어들이 서로 어떻게 연결되는지를 살펴본다. 예를 들어, 모델에 “파리는 프랑스의 수도이다”라는 프롬프트를 입력하면, 어텐션 레이어는 ‘수도’라는 단어를 ‘프랑스’와 연결하고 ‘파리’를 정답으로 식별한다. 그럼에도 불구하고 이 레이어는 프랑스나 파리가 무엇을 의미하는지는 알지 못하지만, 이러한 정보를 KV 캐시 형태로 저장한다. 컨텍스트가 증가함에 따라 KV 캐시도 함께 커진다. - 순방향 네트워크(Feed-Forward Network, FFN)는 모델의 총체적 지식을 가중치 형태로 보유하고 있으며, 이 레이어가 각 단어 뒤에 숨겨진 의미를 찾아내기 위해 심층 수학 공식을 사용한다. 예를 들어, 프랑스라는 단어를 보면 해당 레이어는 국가 백과사전 기능을 활성화하는 식이다. 하지만 KV 캐시가 너무 커지는 것을 방지하고 효율성을 높이기 위해 빔은 여러 가지 트릭을 사용한다. - 균일한 전문가 활용 - 모든 전문 전문가(셰프)들이 거의 동일한 양의 작업을 수행한다. 이는 단일 전문가가 지나치게 피로해지거나 너무 오랫동안 휴식하지 않음을 의미한다. 리플렉션에 따르면, 빔에서 가장 바쁜 전문가는 평균보다 불과 1.04배 더 바빴으며, 이는 거의 완벽한 균형을 이룬다. - 깊이 기반 잔차 스케일링 및 제어된 잔차 스트림 - 모델은 일반적으로 개별 하위 레이어(예: 어텐션 및 순방향 네트워크)의 수학적 출력이 지속적으로 누적되는 수십 개의 연속된 레이어를 통해 토큰을 처리한다. 그러나 모델이 매우 커질수록 이러한 누적 더하기는 일반적으로 막대한 활성화 성장과 수치적 이상치를 유발하여 신호 왜곡이나 학습 불안정성을 초래한다. 이를 극복하기 위해 빔은 토큰이 네트워크를 더 깊게 진행함에 따라 각 하위 레이어 출력의 크기를 점진적으로 감소시키는 스케일링 인자를 사용하여, 정보가 깨끗하고 예측 가능한 방식으로 흐르도록 보장한다. - 샌드위치 정규화(SandwichNorm) - 모델은 일반적으로 토큰이 하위 레이어에 진입하기 전이나 퇴출한 후에 토큰을 정규화한다. 이는 토큰 값이 엄격하고 예측 가능한 범위 내에 머물도록 한다. 반면 빔은 각 하위 레이어의 진입점과 퇴출점(샌드위치) 모두에서 토큰을 정규화하여, 활성화(여러 트랜스포머 레이어를 통해 토큰이 처리될 때 모델이 생성하는 내부 수학 표현으로, 은닉 상태라고도 함)가 흐트러지거나 극단적인 수치적 이상치를 형성하지 않도록 한다. - 어텐션 게이트팅 - 모델이 어텐션(키-값 캐시 생성)을 계산한 후, 실제로 그 어텐션 결과를 얼마나 사용할지 결정하는 작은 ‘볼륨 노브’(게이트)가 작동합니다. 이를 통해 모델은 각 부분에 대해 신호를 증폭하거나 감쇠시켜 더 선택적이고 안정적으로 동작하도록 만듭니다. - FP32 잔차 누적 - 모델은 속도와 메모리를 절약하기 위해 저렴하고 정밀도가 낮은 숫자로 실행되지만, 중요한 잔차(스킵 연결) 더하기 연산은 여전히 완전한 고정밀 32비트 숫자로 수행됩니다. 이는 미세한 반올림 오류가 누적되어 학습이 깨지는 것을 방지합니다. 이제 Reflection의 Beam이 얼마나 효율적인지 구성 요소들을 살펴보았으니, 이제 모델의 학습 과정으로 넘어가겠습니다. 일반적으로 모델이 사전 학습을 거친 후, AI 연구소들은 코딩과 같은 실용적인 기술을 가르치기 위해 모델을 자체 샌드박스 내의 다양한 에이전트와 연결하는 강화 학습(RL)이라는 과정을 수행합니다. Reflection은 전용 블로그에서 다음과 같이 밝혔습니다. “우리는 Beam을 성공적인 솔루션에는 보상을 제공하고 불필요한 토큰은 억제하는 제어 가능한 길이 페널티와 함께 훈련했습니다. RL 초기에는 완성 길이가 줄어드는 동안에도 성능이 향상되었습니다: 모델은 더 적은 추론으로 작업을 더 효과적으로 해결하는 방법을 배웠습니다. 이후 Beam이 더 강력한 에이전트 능력을 갖추면서 완성 길이가 다시 늘어났지만, 이러한 추가 토큰들은 성능의 추가 향상을 뒷받침했습니다.” 흥미롭게도 Reflection은 Beam을 4주(28일) 동안 훈련했으며, 10,500대의 NVIDIA GB300 GPU에 배포된 13억 개의 샌드박스를 사용했다고 선언합니다. 이는 하루에 4,640만 개의 샌드박스에 해당하는 수치입니다. 이 강화 학습은 100만 개의 코딩, 에이전트, STEM 환경을 포괄했으며, 1억 번의 롤아웃(에이전트가 특정 환경과 상호작용할 때 생성되는 결과와 그에 따른 보상)이 포함되었습니다. 여기에 주목할 점은 DeepSeek의 독점 DSec 유닛이 300만 개의 샌드박스를 처리한다는 것입니다. Reflection의 규모와 동등한 수준을 달성하려면 DeepSeek는 단위당 656대의 GB300(총 약 10,500개 GPU)을 갖춘 약 16개의 DSec 유닛을 운영해야 합니다. 다시 본론으로 돌아가, 일반적으로 강화 학습에서는 이전 세션에서 학습된 결과로 모델 가중치를 업데이트할 때 훈련이 일시 중단되어야 합니다. 그러나 Reflection은 파이프라인을 분리했습니다. - 롤아웃 워커는 토큰 응답을 지속적으로 생성합니다. - 러너 엔진은 해당 응답을 사용하여 모델 가중치를 지속적으로 업데이트합니다. 러너 엔진이 롤아웃 워커가 여전히 텍스트를 생성하는 동안 모델 가중치를 업데이트하기 때문에, 데이터가 비동기화(비동기 정책 경사)되어 정책의 노후화라는 문제가 발생합니다. Beam은 샌드위치 정규화(SandwichNorm), 어텐션 게이트팅, 깊이 기반 잔차 스케일링(모두 앞서 설명함)이라는 핵심 안정화 구성 요소를 통해 발생하는 노이즈와 수치적 불일치를 상쇄합니다. 이러한 다양한 기법을 활용함으로써 Reflection은 Beam의 훈련 속도를 극적으로 가속화할 수 있었습니다. 그 결과, Beam의 성능은 GLM-5.2와 거의 맞먹으면서도 추론 관련 연산량을 약 3~4배 덜 사용합니다. 물론 Beam은 여전히 3.8 Max보다 낮은 순위에 있지만, 그럼에도 불구하고 그 크기 대비 매우 강력한 모델입니다. 더 많은 뉴스 커버리지를 피드에서 확인하려면 Wccftech를 Google에서 팔로우하세요.
📄 원문 (English)
Western AI labs are finally trying to compete with China's DeepSeek and Zhipu (Z.ai) when it comes to efficient open-weight models. As a case in point, look no further than Reflection's Beam, a 501-billion-parameter, sparse Mixture of Experts (MoE) model that is insanely efficient. Reflection claims Beam is "3-4x more efficient than GLM 5.2 and over 4x more efficient than leading Western open models" For the benefit of those who might not be aware, Reflection's Beam sports 501 billion parameters, but only 23 billion are activated at any given time. To understand MoE, imagine a large kitchen that has a large number of specialized chefs (routed experts). The model, however, only activates specialized chefs when needed. For instance, to prepare a soufflé, the model would only activate its dessert specialists, and not those who excel in preparing oriental dishes. What's more, An AI model basically consists of a series of transformer blocks that are made up of two key elements: - The Attention Layer looks at how words in a sentence connect with each other. For instance, if you give the model a prompt that says Paris is the capital of France, the attention layer will link the word 'capital' to 'France' and identify 'Paris' as the answer. Even so, this layer does not know what is meant by France or Paris, but would save these notes in the form of KV cache. As context increases, so does KV cache. - The Feed-Forward Network (FFN) holds the sum knowledge of a model in the form of weights, and it is this layer that uses deep mathematical formulas to unearth the meaning behind each word. For instance, it would look at France and activate its country encyclopedia, and so on. However, to prevent the KV cache from growing too big and boost efficiency, Beam uses a number of tricks: - Uniform expert utilization - All the specialist experts/chefs get roughly the same amount of work. This means no single expert is exhausted or left idle for too long. According to Reflection, Beam's busiest expert was only 1.04 times busier than the average, which is almost perfectly even. - Depth-based residual scaling and controlled residual stream - Models typically process tokens through dozens of consecutive layers, where the mathematical outputs from individual sublayers (like attention and feed-forward networks) continuously accumulate. As models get very large, however, this cumulative addition typically triggers massive activation growth and numerical outliers, causing signal distortion or training instability. To combat this, Beam uses scaling factors that gradually dampen the magnitude of each sublayer's outputs as the token progresses deeper through the network, ensuring that information flows in a clean, predictable manner. - SandwichNorm - Models typically normalize tokens either before they enter a sub-layer or after they exit it. This ensures that token values remain within a strict, predictable range. Beam, however, normalizes tokens at each sub-layer entry as well as exit points (sandwich), ensuring that activations (the internal mathematical representations that a model generates as tokens are processed by multiple transformer layers, also called hidden state) never drift or form extreme numerical outliers. - Attention gating - After the model calculates attention (creates KV cache), there’s a little "volume knob" (a gate) that decides how much of that attention result actually gets used. It can turn the signal up or down for different parts, making the model more selective and stable. - FP32 residual accumulation - While the model runs with cheaper, lower-precision numbers to save speed and memory, the important residual (skip-connection) additions are still done in full high-precision 32-bit numbers. This stops tiny rounding errors from piling up and breaking training. Now that we've looked over elements that make Reflection's Beam quite efficient, let's go over the model's training. After a model undergoes a pre-train, AI labs typically perform what's known as Reinforcement Learning (RL), where the model is connected with a multitude of agents in their own sandboxes to teach it practical skills such as coding. Reflection noted in its dedicated blog: "We trained Beam with a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens. Early in RL, performance improved even as completion lengths fell: the model learned to solve tasks more effectively with less reasoning. Later, as Beam developed stronger agentic capabilities, completion lengths grew again, but those additional tokens supported further gains in performance." Interestingly, Reflection declares that it trained Beam over a period of 4 weeks (28 days) and used 1.3 billion sandboxes, deployed over 10,500 NVIDIA GB300 GPUs. This equates to 46.4 million sandboxes per day. The RL spanned 1 million coding, agentic, and STEM environments, and involved 100 million rollouts (the outcome and the attendant rewards generated as an agent interacts with a given environment). Do note that DeepSeek's proprietary DSec unit handles 3 million sandboxes. To achieve parity with Reflection's scale, DeepSeek would have to run around 16 DSec units, spanning 656 GB300s per unit (~10,500 GPUs in total). Coming back, typically in Reinforcement Learning, training has to pause as model weights are updated with learned outcomes from previous sessions. Reflection, however, decoupled the pipeline: - Rollout Workers continuously generate token responses. - Learner Engines continuously update the model weights using those responses Because the learner engines update the model while the rollout workers are still generating text, the data becomes out-of-sync (asynchronous policy gradients), leading to what is known as policy staleness. Beam then counteracts the resulting noise and numerical mismatch through its core stabilization cohort: SandwichNorm, attention gating, and depth-based residual scaling (all explained earlier). By employing all of these nifty tricks, Reflection was able to dramatically speed up Beam's training. As a result, Beam's performance largely matches that of GLM-5.2, while requiring around 3-4x less inference-related compute. Of course, Beam is still outranked by Qwen 3.8 Max, but the model is nonetheless extremely strong for its size. Follow Wccftech on Google to get more of our news coverage in your feeds.
▶ 칩 라이브 24 홈으로 — chiplive24.com

다른 반도체 뉴스

Nvidia vs. Taiwan Semiconductor Manufacturing: Which Technology Stock Is a Better Buy in 2026? · Yahoo FinanceSchneider Electric drops $22.6B on PTC as datacenter boom rains money on infra companies · The RegisterReducing Contact Resistance Pushes 2D Transistors Toward Advanced CMOS (HUST, PolyU, UCSB, NUS) · SemiEngineeringSouth Korea's Q3 Semiconductor ETF Performance Diverges Sharply: Large-Cap Funds Plunge While Materials/Parts/Equipment Funds Hold Up · finance.biggo.comElon Musk Confirmed Terafab Talks With TSMC. Intel Has Been Its Only Named Chip Partner. · Yahoo FinanceJim Cramer Makes His Feelings About Taiwan Semiconductor (TSM) Clear · Yahoo FinanceSPCX Stock Ends Higher On Starship Fuel Plans, Analyst Optimism And TSMC Talks · Yahoo FinanceMicron (MU): AI Memory Boom Could Outlast Your Expectations · Yahoo Finance