AWS Makes Graviton5 Available: Up to 25% More Performance for the Era of Agent AI
AWS has announced the general availability of the new Amazon EC2 M9g and M9gd instances based on Graviton5, the latest generation of its internally developed Arm processor family. First previewed during AWS re:Invent 2025, the new CPU is designed to meet the growing demands of workloads related to agent-based artificial intelligence, a paradigm requiring real-time reasoning capabilities, code generation, and orchestration of complex distributed processes.
The new chip represents a significant evolution from the previous generation. AWS claims performance increases of up to 25% in general computing power, with benefits reaching 35% in web applications and machine learning inference workloads, while database performance improves by up to 30%.
One of the distinctive elements of Graviton5 is the presence of 192 cores within a single package, the highest currently available among Amazon EC2 instances. The architecture has been optimized to reduce the data distance traveled between various cores, enabling a reduction in internal communication latency of up to 33% and an increase in available bandwidth. This approach is particularly advantageous for applications that require high levels of parallelism, such as high-performance databases, data analytics platforms, application servers, online video games, and EDA software for electronic design.
AWS has also significantly expanded the integrated L3 cache. Compared to Graviton4, the new generation offers an overall capacity five times higher and provides approximately 2.6 times more cache to each core. The goal is to reduce the number of accesses to main memory and improve application responsiveness, especially when handling large datasets.
The memory subsystem also receives significant updates with support for DDR5-8800 modules, which AWS defines as the fastest DDR5 memory currently available in the public cloud. Additionally, support for PCI Express Gen 6 has been added, allowing for further increases in communication speed with accelerators and storage devices.
On the I/O front, the new M9g instances offer up to 15% additional network bandwidth and an average increase of 20% in bandwidth dedicated to Amazon Elastic Block Store. In high-end models, network connectivity improvements can reach up to double compared to the previous generation. The M9gd instances, designed for workloads requiring high-performance local storage, incorporate up to 11.4 TB of NVMe SSDs and promise up to 30% more IOPS.
AWS also emphasizes the benefits of energy efficiency. Graviton5 is manufactured using a 3-nanometer production process and benefits from the vertical integration characterizing the company's approach from chip design through server architecture. This strategy has allowed for specific optimizations, including advanced cooling solutions and more efficient management of hardware resources.
Another new feature concerns security. The new instances are based on the sixth generation of the AWS Nitro System and introduce the Nitro Isolation Engine, a component developed through formal verification techniques. According to AWS, the system employs mathematical models to ensure isolation between virtual machines and prevent unauthorized access to customer data, further extending the previously adopted zero-operator access model used in the EC2 infrastructure.
Market interest in the new platform already appears significant. AWS announced in April that Meta plans to use tens of millions of Graviton cores for agent-based AI initiatives, while companies like Uber and Snowflake are adopting the platform for specific workloads. Overall, over 120,000 customers are already using Graviton family processors.
Several partners have shared preliminary results obtained with Graviton5. Airbnb reports improvements of up to 25% compared to other architectures of the same generation and up to 20% compared to Graviton4 in its search systems. Atlassian notes a 30% increase in Jira performance accompanied by a 20% reduction in latency. SAP has observed improvements ranging from 35% to 60% in OLTP queries of SAP HANA Cloud, while Siemens and Synopsys report accelerations in the execution of their EDA tools.