Amazon SageMaker HyperPod now supports model caching for faster inference autoscaling and reduced cold starts
Amazon SageMaker HyperPod now supports model caching, an inference optimization that pre-loads model weights and container images onto cluster nodes so pods start in seconds instead of minutes. When running LLM inference at scale for workloads like chat assistants, agentic pipelines, RAG, and document analysis, cold start is a real bottleneck.
Read AWS's release noteshttps://aws.amazon.com/about-aws/whats-new/2026/09/sgm-hyperpod-model-caching-inf/
Summaries of vendors' own notes. Product names and logos belong to their owners; logos via logo.dev.
More Amazon SageMaker releases
Qwen3.6-35B-A3B-NVFP4 and Wan2.1-T2V-1.3B-Diffusers models now available on Amazon SageMaker JumpStart
UpdateAI agentsData integrationPerformance
Gemma-4-31B-it-assistant and Gemma-4-31B-IT-NVFP4 models now available on Amazon SageMaker JumpStart
UpdateAI agentsPricingPerformance
granite-speech-4.1-2b, kanana-2-30b-a3b-instruct, and OpenFold3 models now available on Amazon SageMaker JumpStart
UpdateAI agentsStreamingPerformance
Amazon SageMaker Feature Store now supports individual feature updates to lower write latency
UpdateStreamingPricingData integration
Amazon SageMaker AI now supports instance preference lists for training and processing jobs
UpdatePricingDeveloper toolsPerformance
Amazon SageMaker AI now supports serverless model customization for NVIDIA Nemotron 3.5 Lightning
UpdatePricingDeveloper toolsPerformance
Also shipped on Sep 11, 2026
Sharing foreign Iceberg tables with OpenSharing is now generally available
GAGovernanceApache IcebergObservability