Your table has 10,000 RCU provisioned and 40% utilization. Your users are still getting throttled. Here is why.
It is late night and your on-call pager fires. The leaderboard table is throttling reads. You check CloudWatch — consumed capacity is well below provisioned. You double the RCU. The throttling continues. You switch to on-demand. Still throttling. You have just discovered that DynamoDB’s most important limit is not on your table. It is on a single partition you cannot see.
The Incident That Taught Me Partitions Are Real
Our table had 10,000 RCU provisioned. CloudWatch showed 40% utilization. Everything looked healthy. But users were reporting timeouts on a specific workflow, and our logs were filling up with ProvisionedThroughputExceededException buried inside SDK retry chains.
The confusing part was the delay between cause and symptom. The AWS SDK retries throttled requests with exponential backoff by default. So what users experienced first was not an error but a slow response. P99 latency crept from 50ms to 800ms over two days before anyone connected it to DynamoDB. By the time explicit errors started surfacing in our application logs, the SDK had already been silently retrying for a while. The dashboard said we were fine. The users said we were not.