Caching Layers in Front of S3
GPU training clusters bottleneck on S3 latency, not compute power — add a caching layer to fix it.
Reporter
Nadia Prasad covers s3 performance and features for Storage Letter.
9 stories
GPU training clusters bottleneck on S3 latency, not compute power — add a caching layer to fix it.
Splitting S3 downloads into parallel byte ranges cuts per-chunk retry time from minutes to seconds.
Geography and object size determine whether S3TA's cost justifies its gains.
EBS hits sub-millisecond, S3 Standard tops 100ms—pick wrong and your GPUs wait for storage.
Spread your S3 keys across more prefixes to multiply your throughput ceiling almost arbitrarily.
Parallel range requests amortize S3's fixed latency cost across simultaneous fetches.
Each storage service excels in different scenarios: know which one fits your architecture.
AWS excels with breadth and existing expertise; GCP wins on egress costs and simplicity.