Skip to main content
TechnologyJun 11, 2026· 3 min read

d-Matrix Launches Production of Corsair: The AI Accelerator Aiming to Surpass GPU Limitations in Inference

The acceleration of artificial intelligence in data centers is entering a new phase.

d-Matrix has announced that Corsair, a solution specifically designed for low-latency AI inference workloads, has entered full production, and the first significant volume supplies will reach hyperscalers, next-generation cloud operators, and laboratories engaged in developing the most advanced AI models in the coming months.

According to the company, the increasing prevalence of agent-based AI applications, coding assistance tools, and real-time voice assistants is putting pressure on infrastructures that rely solely on GPUs. In these scenarios, the determining factor is not only the overall computing power but especially the speed at which the system can generate tokens and respond to user requests.

For this reason, an increasingly widespread approach is emerging in large data centers: the disaggregation of workloads. Instead of entrusting all operations to GPUs, infrastructures combine CPUs, GPUs, and specialized accelerators, assigning each component tasks for which it is most efficient. In this framework, GPUs continue to primarily manage the prefill phase, characterized by high computational intensity, while Corsair is used to accelerate the decode phase, which is the token generation stage.

The uniqueness of d-Matrix's approach lies in the Digital In-Memory Compute (DIMC) architecture, developed to bring processing and memory as close together as possible. Unlike traditional GPUs, which mainly rely on external HBM memory, Corsair utilizes SRAM memory integrated directly into the compute complex. This is a significant architectural choice. The HBM memories currently used in NVIDIA and AMD products offer bandwidths on the order of 1.2 TB/s per stack for HBM3 and up to about 2 TB/s with HBM4. Integrated SRAM, while being more expensive and traditionally limited to the internal caches of processors, can achieve significantly higher values. d-Matrix claims an internal memory bandwidth of 150 TB/s for Corsair, a value the company indicates is about an order of magnitude higher than the HBM solutions available today.

According to the technical documentation provided by the company, a single Corsair PCIe card integrates 2 GB of SRAM-based "Performance Memory" with a bandwidth of 150 TB/s and up to 256 GB of LPDDR5 memory used as "Capacity Memory." The card also features 2,048 DIMC cores, offers up to 2,400 TFLOPS in 8-bit operations, and reaches 9,600 TFLOPS in dense 4-bit processing. In the dual card configuration, these values double, reaching 4,096 DIMC cores and an internal memory bandwidth of 300 TB/s.

The company claims this configuration allows for up to ten times better interactive performance compared to competing solutions based solely on GPUs, as well as up to three times improvement in energy efficiency and cost-performance ratio. However, these are claims made by the manufacturer and are based on preliminary estimates.

On the infrastructural front, Corsair is the foundation of the SquadRack platform, developed alongside Arista, Broadcom, and Supermicro. The system has been designed to be installed in traditional data centers without requiring liquid cooling, instead utilizing air-cooled PCIe cards with a TDP of 600 watts per unit.

Another point highlighted by d-Matrix concerns the production chain. Corsair is manufactured by TSMC using the N6 process and adopts a chiplet architecture based on organic substrates. The absence of advanced CoWoS packaging and large quantities of HBM memory is presented by the company as an advantage in terms of component availability and production predictability, a particularly relevant aspect in a market where the demand for AI accelerators continues to grow rapidly.

With production now underway and the first deliveries expected during the summer, d-Matrix aims to carve out a space in the high-performance AI inference segment by proposing a model complementary to traditional GPUs rather than a direct replacement. This follows in the footsteps of what NVIDIA did at the last GTC with Groq. The goal is to meet the needs of applications that are increasingly sensitive to latency, which represent one of the major challenges of the new generation of AI-based services.