The AMD Threadripper Halo Station is a liquid-cooled workstation built around the 96-core Threadripper PRO 9995WX CPU—codenamed “Shimada Peak”—and up to four AMD Instinct MI350P accelerators, according to announcements made by AMD during the opening keynote at IFA. Jack Huynh, AMD’s Senior Vice President and General Manager of Computing and Graphics, described the system during the Berlin event as a new class of workstation designed to run artificial intelligence models with more than one trillion parameters locally without a cloud connection. The hardware unveiling marked a milestone at the trade show, representing the first time in the 102-year history of IFA that a silicon company held the opening keynote slot, as reported by TechTimes.
Hardware Specifications and Architectural Design
The Threadripper PRO 9995WX Zen 5 central processing unit features 96 cores and 192 threads, a 5.4 GHz boost frequency, and support for up to 2 terabytes of DDR5 system memory, according to AMD’s presentation details. To handle heavy local AI workloads, the system incorporates AMD Instinct MI350P PCIe accelerators. Each MI350P card features 144 gigabytes of HBM3E memory running at 4 terabytes per second, delivering a bandwidth fourteen times higher than any LPDDR5X variant. At IFA, AMD showcased a configuration utilizing two MI350P cards, yielding a total of 288 gigabytes of HBM3E memory. Users can scale the workstation up to four cards to reach 576 gigabytes of accelerator memory. Each MI350P card carries a Total Board Power of up to 600 watts and utilizes independent liquid cooling alongside the CPU.
Did you know?
According to TechTimes coverage of the IFA keynote, the Threadripper Halo Station was introduced alongside the Ryzen AI Max Pro 400 platform (“Kraken Halo”), which features commercial system support from partners including Lenovo and HP.
Memory Bandwidth and Local AI Limitations
Traditional desktop architectures separate memory into distinct system RAM and discrete GPU VRAM pools, creating a hard ceiling for local inference. According to analysis reported by TechTimes, running a 300-billion-parameter language model in FP16 precision requires approximately 600 gigabytes, while compressed FP4 formats need roughly 150 gigabytes—surpassing the VRAM limits of standard consumer and enterprise single-GPU setups. By utilizing unified memory approaches across platforms like Kraken Halo—which pairs Zen 5 cores, an RDNA 3.5 graphics engine, and an XDNA 2 neural processing unit over a 256-bit LPDDR5X interface—AMD aims to bridge the gap between cloud infrastructure and local developer hardware. However, complete system specifications, pricing, and exact availability dates for the Threadripper Halo Station have not yet been announced by AMD.
Frequently Asked Questions
What CPU powers the AMD Threadripper Halo Station?
The system is built around the 96-core, 192-thread AMD Threadripper PRO 9995WX Zen 5 processor codenamed “Shimada Peak,” featuring a 5.4 GHz boost frequency.
How much memory do the Instinct MI350P accelerators provide?
According to AMD, each MI350P card includes 144 gigabytes of HBM3E memory running at 4 terabytes per second. Configured with up to four cards, the workstation can reach 576 gigabytes of accelerator memory.

Can the Threadripper Halo Station run large language models locally?
Yes. AMD states that the accelerator memory capacity allows the workstation to hold a trillion-parameter AI model entirely on-device without relying on a cloud connection.
Have pricing and availability details been released?
No. Complete system specifications, pricing, and availability dates were not announced during the IFA keynote.
What are your thoughts on running trillion-parameter AI models locally? Let us know in the comments below, share this article with your engineering network, or subscribe to our newsletter for more hardware updates.
Worth a look