IBM is preparing to deploy a large NVIDIA-based AI computing cluster on IBM Cloud under a multiyear agreement with Together AI worth $240 million. The cluster is expected to become available in the first quarter of 2027 and will use NVIDIA HGX B300 systems connected through NVIDIA Spectrum-X Ethernet networking. Together AI plans to use […]
The post IBM Cloud and Together AI expand AI infrastructure with NVIDIA appeared first on Cloud Computing News.
IBM is preparing to deploy a large NVIDIA-based AI computing cluster on IBM Cloud under a multiyear agreement with Together AI worth $240 million.
The cluster is expected to become available in the first quarter of 2027 and will use NVIDIA HGX B300 systems connected through NVIDIA Spectrum-X Ethernet networking. Together AI plans to use the infrastructure to run inference workloads for open-source AI models.
The initial deployment will include about 2,000 NVIDIA Blackwell 300 GPUs and will be located in the US, Together AI chief revenue officer Kai Mak told Reuters. Mak said Together AI expects the capacity to be fully committed two to three months before it becomes available.
The deployment will be IBM Cloud’s first dedicated large-scale inference cluster built around HGX B300 systems. According to NVIDIA, the HGX B300 and Spectrum-X configuration is built to deliver 30 times more AI factory output than previous generations.
The cluster is being built primarily for inference, where trained models process requests and generate outputs. Reuters reported that inference has become one of the largest drivers of demand for computing capacity, prompting cloud providers and chipmakers to expand AI infrastructure.
Together AI provides infrastructure and software for AI inference, training, fine-tuning, and agent-based workloads. The company recently raised $800 million in a Series C funding round that valued it at $8.3 billion.
Together AI is also securing computing resources outside the IBM agreement. Alongside the funding round, the company said it had secured commitments for more than 500 MW of compute capacity, which will be financed independently by new investors to support its expected infrastructure requirements.
Together AI says its inference service currently processes more than 400 trillion tokens each month. The company said in July that monthly token volume across its APIs had increased from 30 billion to more than 400 trillion in nine months.
The company has also introduced a different way for customers to reserve inference capacity. Together AI launched a Provisioned Throughput service in July that allows customers to reserve a defined rate of model processing measured in tokens per minute, rather than managing the underlying GPU capacity themselves.
Customers purchase Provisioned Throughput Units, or PTUs, with each unit representing a fixed slice of guaranteed throughput for a selected model or model family. Together AI manages the supporting infrastructure, while customers receive reserved token-processing capacity through the same API used for its other inference services.
The IBM agreement adds another dedicated pool of computing resources to that infrastructure. Together AI selected IBM and NVIDIA based partly on the available GPU capacity and their respective infrastructure roadmaps.
“Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale,” Together AI CEO Vipul Ved Prakash said. “Working alongside IBM with NVIDIA gives us that foundation. This cluster lets us bring production-grade inference to more companies, faster, and it’s a big step in our push to make open-source AI the obvious choice for enterprises.”
Together AI’s platform supports open models including DeepSeek, MiniMax, and Kimi. Reuters reported that businesses using open models have been weighing both AI costs and cybersecurity concerns associated with proprietary model services.
The deployment brings together different parts of the AI infrastructure stack. IBM will provide the cloud environment and deploy the cluster, NVIDIA will supply the HGX B300 computing systems and Spectrum-X networking, and Together AI will use that capacity to operate inference services for open models.
“Enterprises are in a race to adopt agentic AI at scale to drive real business outcomes,” IBM Cloud general manager Alan Peacock said. “IBM and NVIDIA are delivering scalable, economical, enterprise-grade AI infrastructure that can help Together AI accelerate innovation for the next generation of AI infrastructure.”
IBM expands its NVIDIA infrastructure partnershipThe agreement extends IBM’s existing work with NVIDIA rather than establishing a new infrastructure relationship between the companies.
IBM expanded that collaboration in March 2026, when it announced plans to make NVIDIA Blackwell Ultra GPUs available through IBM Cloud. IBM said the infrastructure would support workloads including large-scale model training, high-throughput inference, and AI reasoning.
The two companies are also working across other parts of IBM’s infrastructure portfolio. NVIDIA selected IBM Storage Scale System 6000 to provide 10 PB of high-performance storage for its GPU-based analytics systems, while IBM Storage Scale 6000 has been certified and validated for NVIDIA DGX platforms.
IBM and NVIDIA have also been exploring the integration of IBM Sovereign Core with NVIDIA infrastructure and Nemotron models for GPU-intensive AI workloads that need to operate within regional boundaries because of data residency or regulatory requirements. Their wider work also covers Red Hat AI infrastructure and enterprise consulting.
The Together AI deployment adds another workload to that existing relationship, this time centred on dedicated inference capacity running on IBM Cloud.
IBM is not the only cloud provider combining internally developed infrastructure with NVIDIA technology. AWS agreed this year to buy one million NVIDIA GPUs by the end of 2027 and plans to deploy NVIDIA ConnectX and Spectrum-X networking equipment in its data centres while continuing to develop its own processors and networking hardware.
AI capacity is contracted ahead of deploymentTogether AI expects the IBM cluster to be committed before it enters service. Other AI infrastructure providers are also reporting large commitments for capacity that is still being deployed.
CoreWeave reported a $104.2 billion revenue backlog for the second quarter, up from $99.4 billion in the first quarter. The company also said it had secured more than $25 billion in net new customer commitments so far in the current quarter.
Nebius separately reported more than $40 billion in customer commitments and four AI cloud agreements averaging more than $1 billion each.
The figures are not directly comparable with Together AI’s $240 million IBM contract because the companies use different commercial models, but they provide additional examples of computing capacity being sold or committed ahead of full deployment.
Nebius also said it expects more than $9 billion in customer prepayments during 2026 and has raised its contracted power target for the year to 5 GW. It plans to deploy more than 1 GW of capacity annually beginning in 2027.
Together AI’s agreement differs from those contracts in its focus on dedicated inference capacity. The company will use the NVIDIA systems deployed on IBM Cloud to provide open-model inference to customers when the cluster becomes available in 2027.
“AI factories are becoming essential enterprise infrastructure—like electricity and telecommunications—turning compute and data into intelligence,” NVIDIA senior director Dion Harris said. “With NVIDIA HGX B300 systems and NVIDIA Spectrum-X Ethernet networking on IBM Cloud, IBM and Together AI will deliver an accelerated computing platform to help enterprises deploy open-source AI with the performance, efficiency and scale required for real-time AI services.”
(Photo by Carson Masterson)
See also: IBM moves to buy Confluent in an $11 billion cloud and AI deal

Want to learn more about Cloud Computing from industry leaders? Check out Cyber Security & Cloud Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.
Cloud Computing News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | Nvidia in talks to back OpenAI’s $500bn Ohio data centre | 0 | 7.71 | 28-07-2026 |
| 2 | Deploying NVIDIA Run:ai on VMware Cloud Foundation | 0 | 13.63 | 31-07-2026 |
| 3 | Volta lands $10B AI cloud deal as GPU clouds expand | 0 | 11.33 | 05-08-2026 |
| 4 | NVIDIA Rallies $500 Billion Wall Street Alliance To Fuel AI Expansion | 0 | 10.98 | 11-08-2026 |
| 5 | Рег.облако подключил новый ЦОД в Москве под интенсивный рост спроса бизнеса на ИИ-вычисления | 5 | 7 | 17-03-2026 |
| 6 | Shares In AI Cloud Firm Nebius Soar On Nvidia Investment | 0 | 5.76 | 12-03-2026 |
| 7 | Thinking Machines Lab inks massive compute deal with Nvidia | 0 | 10.41 | 10-03-2026 |
| 8 | Vultr says its Nvidia-powered AI infrastructure costs 50% to 90% less than hyperscalers | 0 | 19.81 | 03-04-2026 |
| 9 | Крупный и средний бизнес за полтора года вчетверо увеличил траты на GPU-серверы — данные Рег.облака | 0 | 7 | 04-06-2026 |
| 10 | FERC orders faster grid access for AI data centres | 0 | 7 | 19-06-2026 |