Skip to content

Amazon SageMaker HyperPod now supports model caching for faster inference autoscaling and reduced cold starts

UpdateVerifiedAdded Sep 22, 2026

Amazon SageMaker HyperPod now supports model caching, an inference optimization that pre-loads model weights and container images onto cluster nodes so pods start in seconds instead of minutes. When running LLM inference at scale for workloads like chat assistants, agentic pipelines, RAG, and document analysis, cold start is a real bottleneck.

Read AWS's release notes

https://aws.amazon.com/about-aws/whats-new/2026/09/sgm-hyperpod-model-caching-inf/

Summaries of vendors' own notes. Product names and logos belong to their owners; logos via logo.dev.

More Amazon SageMaker releases

Also shipped on Sep 11, 2026

SnowflakeSnowflake

Automations in Snowflake CoWork (General availability)

GA
DatabricksLakeflow

Session restore for serverless jobs is in Public Preview

Preview

Weekly: the week's data and AI releases, Tuesday mornings.