tipsSEP 9 22:26 UTC
AWS guide covers deploying Qwen3.8-2.4T-A95B on SageMaker HyperPod with vLLM
Amazon published a walkthrough for running the open-weight Qwen3.8-2.4T-A95B model, which has 2.4 trillion parameters, on its SageMaker HyperPod service using the vLLM inference engine. The guide covers setting up the cluster, applying NVFP4 quantization, and exposing an OpenAI-compatible endpoint. It also notes support for tool calling, reasoning, and multi-token prediction speculative decoding.