The GPU market can be complex and challenging to navigate. If you’ve spent hours searching for information on the H100 market, you’ve likely encountered countless offers with similar messages: “Talk to our sales.” All GPU infrastructure options might seem alike, but you know they’re not — especially if you’ve heard stories about stability issues and hidden challenges in managing GPU clusters. So, where do you turn?
If any of this sounds familiar, you’re in the right place. We recognize the lack of a comprehensive guide to the GPU market, and we’re here to share our insights, particularly about the flagship H100 GPU. As of late August 2024, this guide will cover:
- What are the general prices for various options in the market?
- How can I ensure that GPUs are reliable?
- Do hardware specifications beyond GPUs matter?
- Where are the GPUs? Does location matter?
Who are we? We’re the team behind Caffe, ONNX, PyTorch, and etcd. We are now building an AI cloud at Lepton AI. We run a fleet of GPU resources across all major IaaS providers and on-premises resources for clients. In our previous roles, we’ve operated AI infrastructure for some of the world’s largest tech companies, including Meta, Uber, and Alibaba. This extensive experience has given us deep knowledge of both the technology and the market, enabling us to offer the most cost-effective and reliable solutions for both training and inference.
Whether you’re looking to buy or rent H100 GPUs, you can contact us here. We look forward to hearing from you and helping you along your journey!
Getting GPUs: Pricing
First of all, price. A short summary is, the H100 price is going to hold for a while but will eventually come down, and rental terms are getting shorter, allowing a more flexible capacity planning. We’ll cover the comparison between renting and buying GPUs, breaking down the price by key parameters.
Renting: expectations from short to long terms
Press enter or click to view image in full size
Market rewards predictability
So far, the most common way to access H100 GPUs is reserved computation capacities. This is because GPUs are expensive, and vendors do not fancy the idea of idle GPUs. Reservation provides predictability, and you as a user get rewarded in turn with a better price.
Reservations typically start with a minimum commitment of 6 months. For small-scale clusters (between 16 and 512 GPUs), the current fair baseline pricing is approximately $2.60/h for a 6-month commitment, $2.40/h for a 12-month commitment, and less than $2.20/h for longer terms. These prices include fully loaded configurations, featuring server-grade CPUs, 2 TB of memory, 40 TB of local NVMe storage, and InfiniBand/RoCE interconnects (we’ll cover hardware specs later). Pricing may vary outside the US, depending on the specific location, with some regions offering lower or higher rates based on local conditions. But in general, if you see the price deviating significantly from the above, you might inquire about the underlying reason to make sure you are with a good provider.
For large-scale clusters (more than 512 GPUs), pricing is highly variable and depends on several factors, making it difficult to establish a standard baseline.
On-demand is becoming a thing
In late 2023 and early 2024, on-demand access to GPUs was nearly impossible. Although Lambda Labs offered on-demand H100 GPUs at around $3.50 per hour, availability was extremely limited, often requiring significant luck to secure a machine. We see changes coming to on-demand land as well. Currently, on-demand H100 GPUs are more accessible, with prices around $3 — $3.5 per hour across several providers, including Lambda Labs, Voltage Park, DigitalOcean, Runpod, and CoreWeave. However, these on-demand H100 GPUs generally come with limitations, such as the lack of high-bandwidth GPU fabric (like InfiniBand or RoCE), and a wait time longer than normal CPU-based on-demand resources.
A100 as a benchmark for H100 pricing trends
The history of the A100 GPU gives us useful insights into the current trends in the H100 market. When the A100 was first released in 2020, it cost $2.4 per hour to rent. By 2023, this price had dropped to $1.8 per hour, and by 2024, it had decreased further to about $1.4 per hour. This drop in price came alongside a significant increase in availability. For instance, Azure now offers A100 spot instances with reasonable availability and attractive discounts.
The H100 seems to be following a similar path, with prices falling by around 20% over the past year. Lead times, which used to be up to 6 months, have now reduced to just a few weeks or even less. However, the pricing of the H100 remains unclear and is heavily influenced by marketing tactics.
Looking ahead, predicting future H100 prices is difficult due to uncertainties around the upcoming Blackwell GPUs and changing demand. Even so, we expect that supply will keep improving and pricing will become clearer. As availability becomes less of an issue, factors like reliability, support, and software capabilities will become more important in choosing between different GPUs.
Buying: the cost breakdown
Another option is to buy the machines upfront. “I’ve heard that cloud providers charge a high premium. Should I just build my own GPU cluster?” If you would like to do so, or if you just want to have a quick look under the hood at the cost breakdown, read on.
We’ll make a few accounting assumptions: the machine and other parts will be linearly deprecated over 4 years; we assume the price is all the same over the 4 years; and we ignore financial factors (buying requires a higher upfront payment). These are, of course, gross simplifications, but they help make the math clear. We’ll also convert everything to a “per GPU per hour” price for an easier comparison.
We break down the cost to the following parts: compute hardware, network hardware, power and other IDC cost, and spare parts.
Press enter or click to view image in full size
Compute Hardware: $1.1/h
The primary cost factor is undoubtedly the hardware. A Dell HGX system with 8 H100 GPUs, typically configured with 2TB of memory and 40TB of NVMe storage, is priced around $280,000. A similar configuration from Supermicro is slightly less expensive, at approximately $270,000. Factoring in sales tax, this translates to about $1.10 per GPU-hour.
Network Hardware: $0.2/h
Usually, one would also like to place a high-performance network if building a mini-cluster for large scale distributed computation. Normally, this involves two parts: network devices on the server and switches and cables for connecting the services. Depending on the specs and the size of the cluster, expect that the network cost is going to take more than 20% of the machine cost. You can read more about network specs in the later part of the article. We’ll approximately estimate it at around $0.2 per GPU-hour.
Power and other IDC costs: $0.3/h
Each fully utilized H100 GPU consumes approximately 800 watts of power. With the average cost of electricity and IDC rent in the U.S. being about $200 per kilowatt per month, this results in a monthly energy cost of around $160 per GPU. Additionally, onsite maintenance service charges (sometimes called “smart hands”) for each machine typically require one hour per three months, at a rate of $150 per hour. Therefore, the total monthly cost for operating an H100 GPU is roughly $0.3 per card-hour.
Spare parts: $0.1/h
To achieve 99.9% uptime, it’s generally necessary to maintain a reserve of 3–5% spare parts. However, even with this precaution, some unavoidable hardware or network issues might still cause occasional further downtime. For this discussion, we are optimistically using a 5% reserve as our baseline. We also factor in the potential downtime associated with replacing spare parts, estimating a total cost of around $0.1 per hour.
What is not covered
The above cost breakdown leads to a total of about $1.7/h over 4 years. Note that this only involves BOM costs, and it is likely that you will invest a certain amount of people resources to get your own cluster up and running. This ranges from “managed by one’s own researchers” to “dedicated SRE task force”, depending on the scale and the complexity of your dedicated GPU cluster. There are surely going to be complexities like a grumpy research team or non trivial operational cost, but these might be hard to analyze in a standard fashion.
Get Lepton AI’s stories in your inbox
Join Medium for free to get updates from this writer.
If you are a startup, you might want to ask — is it worth it? Which leads us to the next question.
Renting or buying?
This is a hard question, but the years of observation lead us to slightly lean towards renting as Lepton’s recommendation, with the reasons being:
- Price drops. 4 years is a long time. H100 rental prices will surely drop consistently, so the benefit of buying may decay faster than one expects.
- Lower upfront cost: each HGX machine cost approximately 6 Tesla Model Ys. Renting gives you a better cash flow, especially if you are spending expensive VC money.
- Better flexibility: if you need to scale up or down the computation resources, renting is relatively easier to do, at least at the end of each renting term. Scaling the cluster you own is going to be trickier; it is not uncommon that your datacenter runs out of rack spaces, and you’ll have to drive your cluster around for another datacenter.
Of course, there are arguments for buying, such as:
- Data security and related considerations are paramount, and you do need an air-gapped system.
- You are absolutely committed to a fixed amount of computation for a determined couple of years of usage.
Unlike raw GPU infrastructure providers, Lepton helps our users to efficiently operate the computation, whether it be rented GPU computation capacity or dedicated clusters. Lepton’s clients make full use of the capacity by having a fully cloud-native platform to manage GPUs, orchestrate training jobs and inference services, and run AI workloads in an efficient manner.
Using GPUs: Reliability
GPUs are fast and loose. Generally, they experience failures more frequently than traditional CPU machines. If you’re managing around 100 GPU cards, it’s reasonable to anticipate at least one card failure each month. In a recent blog post introducing GPUd, we discussed GPU reliability and recovery. Similar observations have been made in leading GenAI companies, such as Meta which had about 8.6 job interruptions per day during a 54-day period of Llama 3 405B pre-training.
Press enter or click to view image in full size
As a result, whether you rent or buy GPUs, it is crucial to ask your provider if they do the following:
- Do they do thorough pre-production and burn-in testing before delivery?
- Do they do extensive active monitoring during operation of the cluster?
- Are there IDC onsite staffing support, and what are the SLA guarantees?
Thorough pre-production and burn-in testing
GPUs tend to have a higher failure rate when they are first turned on. One of the most crucial steps before deploying GPU machines is to eliminate misconfigurations and dead-on-arrival components. Unlike CPU-only systems, GPU machines have more complex components that require comprehensive testing beyond standard CPU benchmarks. Common tests include GPU burn-in and NCCL all-reduce tests. Such tests ensure that the fundamental hardware specs are met.
We go a step further at Lepton. Before delivering H100 GPU machines to our clients, we conduct not only standard burn-in tests but also run popular training frameworks such as torch.distributed and DeepSpeed, to verify end-to-end training performance. We identify hidden issues such as potential ECC errors, slowdown due to GPU-GPU communication, Ethernet or storage, and network throughput between the GPU cluster and external internet. As a result, our clients receive machines that are fully optimized and ready for demanding workloads from day 1.
Extensive active monitoring
Continuous monitoring is essential. No one wants to have GPUs broken at 3am only to be identified at 9am. At a minimum, your infra provider should have IPMI or a similar tool configured to track basic hardware status and PCIe state. This detects common issues like GPUs disconnecting from the PCIe bus or NVMe disk failures. Some advanced IaaS (infrastructure as a service) providers offer additional comprehensive GPU-GPU and data center network monitoring, but we find this level of service to be rare. At the time of this blog, we haven’t seen any providers, including AWS or GCP, that offer a comprehensive, active GPU health monitoring (ECC errors, power issues, NVLink status, etc.). These are normally left as the responsibility of the users.
At Lepton, we believe users deserve better reliability. And that’s why we open-sourced GPUd: to ensure GPU efficiency and reliability by actively monitoring GPUs and effectively managing AI/ML workloads.
We deploy GPUd on every machine on Lepton to ensure complete monitoring coverage of GPU and related components. This allows us to detect early signs of potential failures and conduct detailed diagnostics to identify the root cause of issues. By doing so, we can determine the most effective way to restore the cluster to a healthy state. This also helps ensure that SLAs are met and enables you to secure appropriate refunds from the provider when necessary.
A typical example is ECC errors: while “software ECC errors” are normally considered “correctable”, we find them to lead to uncorrectable hardware failures consistently in a few days of intensive use. As a result, we can proactively taint nodes and make predictive maintenance. This eliminates unnecessary disruptions to our customers, both training and inference. For our vendor partners, this also helps improve SLAs and an overall satisfaction end to end.
IDC onsite staffing and SLA guarantees
Most infrastructure providers advertise 24/7 onsite support, but the quality of service can vary significantly depending on staff availability. Moreover, a 24/7 onsite guarantee doesn’t always translate to quick resolution of machine failures, especially when the issues go beyond simple reboots. Common data center tasks include resetting GPU cards, replacing or reconnecting network cables, and troubleshooting power devices. However, identifying these issues can take extra time.
To tackle this challenge, Lepton developed GPUd, a proactive tool that monitors machine status and helps quickly pinpoint the root cause of issues. Once identified, onsite staff can usually resolve problems within a few hours, minimizing the risk of extended downtime.
Service Level Agreements (SLAs) hold providers accountable for their performance, offering compensation if standards aren’t met. However, it’s important to note that the SLA for a GPU cluster often differs from that of regular CPU clusters. While most providers offer a high SLA guarantee for the control plane, the uptime guarantee for GPUs is typically lower due to their higher failure rates and the greater difficulty of migration and virtualization compared to CPUs. Together, these measures ensure smooth operations, prompt issue resolution, and protection against significant business disruptions.
Hardware Specs: a datacenter view
Guess not, GPUs cannot operate without peripherals, and the infrastructure around it is rather complex and convoluted. You are not buying a single card. Such surrounding infrastructure is as important as the GPUs to ensure that you are getting the best system performance. We’ll give you a bit more information about the choice of the GPU servers, network, CPU and memory, storage, etc.
Press enter or click to view image in full size
GPU Servers
The Nvidia H100 HGX system is supplied by several major vendors, including but not limited to Dell, Gigabyte, and Supermicro. While the pricing across these vendors is relatively similar and not a significant differentiating factor, there are vendor-specific considerations to keep in mind, such as PSU, cooling, and other hardware-related issues. However, these are typically addressed through firmware and part updates as the H100 platform matures.
We have had positive experiences with Supermicro in terms of delivery, installation, customer support, and incident responses. Other players in the market may prefer Dell due to its broad availability and support network. Overall, GPU machines themselves are a relatively standard product as H100 matures. This might change with the new NVidia NVL36/NVL72 with Blackwell GPUs, so we keep a constant eye on them. In addition, if you are building a cluster, you might want to work with system integrators, such as AMAX, who will provide more end-to-end solutions in addition to individual servers.
GPU Network
GPU networking focuses on high-performance, low-latency interconnects between GPUs so you can do distributed training efficiently. There are normally two choices: InfiniBand (sometimes called IB) and Remote Direct Memory Access over Converged Ethernet (RoCE). Both choices can deliver high network bandwidth, with InfiniBand offering slightly lower port-to-port latency (~200 ns) and RoCE being more of an open standard.
A common misconception is that InfiniBand is essential for training. However, RoCE has developed its maturity over the years, its capability demonstrated by leading models such as LLAMA 3.1 with ultra-scale training infrastructure. In general, InfiniBand is more worry-free and comes with mature commercial fabric management solutions, such as NVidia UFM. RoCE generally provides better availability and is arguably more scalable, although that needs expertise.
Based on our experience, the difference between RoCE and InfiniBand is minimal for training clusters with fewer than 1,000 GPUs. InfiniBand may offer some operational advantages if you’re building and managing the cluster yourself. However, RoCE is an equally viable option, often at a more competitive price point.
In both cases, available bandwidth is a design factor worth mentioning. A popular choice is an 8-way Infiniband or RoCE network card, each delivering 400GB throughput. That is why you often hear “3.2T interconnect” in vendors. In practice, you may also reduce it down to 4-way or 2-way and still train most models efficiently. We do find many providers to offer 8-way for the sake of future proof, though.
CPU/Memory
Given the high cost of the H100 GPU, the expense of the CPU and memory becomes a relatively small portion of the total cost, typically around 10% of the entire machine. Most providers fully equip their systems with CPU and memory to maximize performance. For a machine with 8 H100 GPUs, it’s common to see configurations with over 96 physical cores or 192 vCPUs, often using Intel Xeon Platinum Sapphire Rapids or AMD 9004 series processors. Memory configurations usually include 2TB, although 1TB is also a viable option.
Lepton generally recommends fully utilizing the available CPU and memory resources to ensure they do not become bottlenecks during training or inference.
Storage
Fast local storage is essential for both training and inference tasks in AI workloads. Ideally, a machine should be equipped with at least 20 TB of NVMe storage, though high-performance systems often provide 40 TB or more. The local disk should be large enough to accommodate the entire dataset needed by the server during training or inference, reducing network dependency and maximizing the performance of local GPUs.
However, for tasks involving large datasets like image, video, and audio training, the storage demands often exceed local capacity, requiring the use of remote storage solutions. The most common approach is using NFS, with commercial alternatives like Lustre, VAST, and Weka also being popular. Additionally, object storage options such as S3, Minio, or Ceph are viable, though POSIX file systems are generally more familiar to researchers. At a minimum, each GPU should have a read throughput of 200MB/s, with the system supporting a machine-wide write throughput of 1GB/s for checkpointing.
Beyond storage capacity, other challenges often arise, such as handling numerous small files. For instance, each small image or video might only be a few kilobytes, and researchers often prefer random access directly from serverless storage.
To address these needs, we at Lepton developed a general-purpose storage solution tailored for AI training. Our POSIX-compatible distributed file system persists data on remote storage or object stores, while caching it on local NVMe disks in a peer-to-peer, serverless manner. This approach offers the scalability of remote storage while maintaining the performance and simplicity of local disk access.
Location, location, location
GPU vendors today are distributed across the globe. North America is the most popular choice for hosting due to its lower power costs, affordable network rates, and better availability of parts. Europe ranks second, while Asia Pacific generally has a higher cost of bandwidth and power. But GPU rental prices also depend on demand, so the price does not relate linearly.
For training, location is less critical as long as you can move the chunk of training data (normally at the terabyte or low petabyte scale) in and out of the cluster once. For inference, latency and reliability are significant factors. It’s important to position your infrastructure close to the majority of your customer base. Additionally, distributing your capacity across multiple locations helps prevent single points of failure and enhances network robustness. A rule of thumb: US east to US west adds around 60 milliseconds latency, and US to Asia Pacific adds about 150 milliseconds.
Lepton operates a global supply chain to maximize GPU availability. We also adopt a good amount of Point of Presence (POP) nodes so that latency between the GPUs and client is minimized.
Conclusion
Whether one builds a cluster of one’s own or rents GPUs from an IaaS provider, there is still a long way to go between the raw computation power and a fully up-and-running, high-performance training job. Compute, storage, networking, and model-specific optimizations are all involved to make everything efficient.
The article is only a peek into the vast complexity of the GPU market. The Lepton team has extensive experience in building software and hardware solutions at 10,000s of GPUs scale. In our career, we’ve served autonomous driving, AI for science, and, needless to say, large-scale LLM and other GenAI training and inference. As a result, we have invested considerable effort in understanding the GPU supply landscape, and collaborate with most major GPU providers.
Lepton as a one-stop solution for your AI needs: not only do we help you find GPU resources, we build you a full AI cloud so your engineers and researchers can maximize the productivity. Not all GPUs are the same, and we make them work best for you: expertise knowledge about operating GPU clusters, a reliable software stack, and efficiency in training and deploying your own AI models?
Whether you’re looking to rent or buy H100 GPUs, you can contact us here, or shoot an email to info@lepton.ai. We look forward to hearing from you and helping you along your AI journey!