ServeTheHomeAlchipTSMCSamsungSK HynixReutersLam ResearchGartnerBloombergDeloitteAMDGroqMicronNVIDIATrendForceNvidiaObjective Analysisd-Matrix

d-Matrix unveils Raptor 3D-DRAM inference chip at Hot Chips 2026, claiming 10x HBM4 bandwidth to break AI memory wall; Nvidia hikes server prices 15%

At Hot Chips 2026, d-Matrix introduced its Raptor 3D-DRAM inference chip, boasting 10x the bandwidth of HBM4, addressing the AI memory wall. Meanwhile, Nvidia announced a 15% price increase for its AI servers, driven by rising memory costs and supply chain challenges.

Startup Fortune Startup Fortune+1 source24 August 2026 · 07:52 UTC
CuriousCats Full Story

d-Matrix unveiled its Raptor 3D-DRAM inference chip at Hot Chips 2026, claiming to deliver 10x the bandwidth of HBM4. This innovation aims to tackle the AI memory wall, a critical issue as demand for memory bandwidth surges with the growth of large language models (LLMs).1

The company’s 3DIMC technology is designed to enhance inference performance, particularly during the decode phase, which is memory-bandwidth-bound. d-Matrix asserts that improving decode bandwidth is essential for overall inference efficiency, as most inference time is spent in this phase.3

Raptor utilizes a TSMC N4 logic die stacked on a 3D DRAM die, achieving a bandwidth of 32.6 GB/s per mm², significantly outperforming HBM parts at 1.5 GB/s. This results in a 13.5x improvement in energy efficiency, consuming 2.96 mW per GB/s compared to HBM's 40 mW.

In parallel, Nvidia announced a 15% price hike for its AI servers, attributed to rising memory costs and supply chain constraints. This increase affects major configurations like Vera Rubin and Grace Blackwell, with some customers facing significant price adjustments for systems shipping early next year.2

As the demand for AI capabilities escalates, the competition between d-Matrix and established memory providers like Samsung, SK Hynix, and Micron intensifies, highlighting the urgent need for innovative solutions to overcome the memory bandwidth challenges in AI applications.

Key Insight
“The memory wall is driving Nvidia to raise AI server prices by over 15% for systems shipping early next year, including Vera Rubin and Grace Blackwell configurations. Meanwhile, TrendForce projects conventional DRAM contract prices to climb 58-63% quarter-over-quarter in Q2 2026 as suppliers shift capacity toward HBM.”
CuriousCats studied:
1
Startup FortuneStartup Fortune
“At Hot Chips 2026, d-Matrix unveiled Raptor, a 3D-DRAM inference chip claiming 10x the bandwidth of HBM4, while Samsung, SK Hynix and Micron detailed the memory wall now driving Nvidia's 15% AI server price hikes.”
Startup Fortune →
2
ServeTheHomeServeTheHome
“Model weights keep growing, and the KV cache scales with context length multiplied by batch size. So 64 users at 1M context can mean roughly 935 GB of KV cache. Weights and cache together create a problem that is both a capacity problem and a bandwidth problem, and both sides keep growing.”
ServeTheHome →
Ask CuriousCats
What is the Raptor 3D-DRAM chip?
Why is there a memory wall in AI?
How does HBM4 compare to previous technologies?
Are Nvidia's server prices the highest in the market?
How does TrendForce's projection compare with past quarters?
Get your CIA-level briefing,
in real time.
CuriousCats monitors the internet every minute for you and brings you the most personalized brief of videos, social media posts, news and more.
Download the App
If you liked this, you’ll love your CuriousCats brief.
News, videos, opinions and more — without the noise.
Get CuriousCats