AWS and NVIDIA announced on August 26 that they plan to install two million additional GPUs in 2027–2028. The order covers Blackwell Ultra, Rubin and Rubin Ultra. It adds to the previously announced million accelerators whose deployment starts in 2026.
A GPU performs many calculations in parallel. These processors train models, generate text and process images. Amazon installs them in its data centres, and customers access the computing resources remotely without buying or maintaining the equipment.
Two million accelerators does not mean two million identical servers
A server can contain several GPUs, and training a large model can require a group of servers connected by a fast network. Some customers rent whole groups. Others need a single server to run a trained model. The total accelerator count does not reveal how many machines of each type an individual customer will be able to use.
AWS has not published the regional allocation, detailed rollout schedule or tariffs for this batch. The announcement gives the total volume and installation period: 2027–2028.
The equipment will also sit behind managed services
The same AWS–NVIDIA agreement covers Nemotron models in Bedrock and SageMaker, accelerated data processing and search, and robotics. The expansion reaches several AWS services beyond GPU rental.
Consider document search. A team can rent a server and install a model itself, or send requests to a hosted model through an API, the interface an application uses to request and receive an answer. Both approaches use remote servers. With the hosted model, the service provider selects and maintains them.
Before a model can answer, the archive also needs preparation: extracting text from files and building an index that helps the application find relevant passages. That is a separate computing task. Accelerating it can help when a large archive changes frequently, even if the answering model stays the same.
NVIDIA will also contribute to Amazon’s own chips
Another part of the agreement concerns Trainium, AWS’s own model accelerators. Amazon and NVIDIA are working on integration with NVLink Fusion, a technology for fast chip-to-chip communication, and access to NVIDIA’s new NVHBM memory.
Memory holds the data an accelerator works with, while interconnects move it between computing units. Performance therefore depends on more than the individual chip. Slow data transfer can leave an accelerator waiting even when it has enough computing power.
The companies describe a future architecture that can combine Trainium and GPUs at rack level. Their collaboration also reaches Amazon equipment using its own accelerators, where NVIDIA memory and interconnect technologies are planned.



