Open-weight reasoning models changed what a self-hosted endpoint has to do. A classic instruction-tuned model returns an answer; a reasoning model such as gpt-oss or a DeepSeek-R1 distill first emits a long chain of thought and then the answer, in a separate channel. If you deploy one with the same container settings, the same timeouts, and the same parsing code you used for a 7B chat model, you will get truncated responses, blown latency budgets, and reasoning text leaking into your product UI.
This tutorial deploys an open-weight reasoning model to a SageMaker AI real-time endpoint using the Large Model Inference (LMI) container with the vLLM backend, then covers the three things that are specific to reasoning models: budgeting output tokens, separating reasoning from the final answer, and enforcing structured output.
1. Pick the model and the instance together
Reasoning models spend tokens, not just parameters. Throughput planning has to assume a long generation — frequently 1,000 to 8,000 output tokens per request — on top of your prompt.
| Model class | Typical fit | Starting instance |
|---|---|---|
| ~20B open-weight reasoning (e.g. gpt-oss-20b), MXFP4/4-bit | single GPU | ml.g6e.2xlarge (1x L40S, 48 GB) |
| 32B dense reasoning distill, BF16 | single large GPU or 2-way TP | ml.g6e.12xlarge / ml.p4d.24xlarge |
| ~120B MoE reasoning (e.g. gpt-oss-120b) | 1x H100-class or 4-way TP on L40S | ml.p5.48xlarge or ml.g6e.48xlarge |
Two sizing rules that matter more than parameter count:
- KV cache dominates. Long reasoning traces at high concurrency need headroom. Leave at least 30–40% of GPU memory for KV cache after weights, or vLLM will queue requests and your p95 will fall off a cliff.
- Check quota before you design.
ml.p5.48xlargequota is not granted by default. Request it in Service Quotas (for the endpoint usage dimension, which is separate from training) before you promise anyone a date.
2. Deploy with the LMI container and vLLM
The LMI container ships vLLM as its default rolling-batch engine, so most of the configuration is environment variables rather than code. Deploy from the Hugging Face Hub:
import sagemaker
from sagemaker.djl_inference import DJLModel
role = sagemaker.get_execution_role()
session = sagemaker.Session()
model = DJLModel(
model_id="openai/gpt-oss-20b",
role=role,
env={
"OPTION_ROLLING_BATCH": "vllm",
"OPTION_TENSOR_PARALLEL_DEGREE": "max",
"OPTION_MAX_MODEL_LEN": "32768",
"OPTION_MAX_ROLLING_BATCH_SIZE": "16",
"OPTION_GPU_MEMORY_UTILIZATION": "0.90",
"OPTION_ENABLE_PREFIX_CACHING": "true",
"OPTION_ENABLE_REASONING": "true",
"OPTION_REASONING_PARSER": "openai_gptoss",
"SAGEMAKER_MODEL_SERVER_TIMEOUT": "900",
},
)
predictor = model.deploy(
instance_type="ml.g6e.2xlarge",
initial_instance_count=1,
endpoint_name="gpt-oss-20b-reasoning",
container_startup_health_check_timeout=1800,
)
Notes on the settings people get wrong:
container_startup_health_check_timeoutmust cover weight download and load. A 120B MoE pulled from the Hub can take 20+ minutes on first boot; endpoint creation fails silently-ish (as "ping failed") if you leave the default. Staging weights in S3 and passingmodel_datawithS3DataType="S3Prefix"cuts this substantially and is what you want in production.OPTION_MAX_MODEL_LENis the context the engine reserves for. Setting it to the model maximum when you only need 32k wastes KV cache and lowers concurrency.OPTION_ENABLE_PREFIX_CACHINGis close to free and is a large win for agent loops and RAG, where every request repeats a long system prompt.- The reasoning parser is what turns raw channel-tagged output into a structured field. Pick the parser that matches your model family (
openai_gptossfor gpt-oss,deepseek_r1for R1-style distills); the wrong one silently returns everything as content.
3. Invoke it, and keep reasoning out of your UI
The LMI container exposes an OpenAI-compatible chat schema. With a reasoning parser enabled, the response splits the chain of thought into reasoning_content and the user-facing answer into content:
import json, boto3
rt = boto3.client("sagemaker-runtime")
payload = {
"messages": [
{"role": "system", "content": "You are a logistics planner. Reasoning: medium."},
{"role": "user", "content": "Three warehouses, two trucks, demand below... what ships first?"},
],
"max_tokens": 4096,
"temperature": 0.6,
}
resp = rt.invoke_endpoint(
EndpointName="gpt-oss-20b-reasoning",
ContentType="application/json",
Body=json.dumps(payload),
)
msg = json.loads(resp["Body"].read())["choices"][0]["message"]
answer = msg["content"] # ship this to the user
trace = msg.get("reasoning_content", "") # log it, do not display it
Three rules we apply on every reasoning deployment:
- Never render
reasoning_contentto end users. It is unfiltered, frequently wrong-then-corrected, and it is the fastest way to leak system-prompt contents into a screenshot. - Do log it, with a retention policy. Reasoning traces are the best debugging artifact you will ever get from an LLM, and they are also text you have to treat as data under your privacy rules. Capture them to S3 with a lifecycle rule rather than into application logs that live forever.
- Budget
max_tokensgenerously but finitely. A reasoning model that hits the cap mid-thought returns nothing useful — the final answer never arrives. If a task truncates, raise the cap or lower the reasoning effort; do not retry blindly at the same setting.
Most gpt-oss-style models accept a reasoning effort hint (low, medium, high) in the system prompt. Treat it as a cost dial: it is the single biggest lever on tokens per request, and low is correct for routing, classification, and extraction.
4. Streaming so the latency is survivable
Time-to-first-useful-token on a reasoning model can be tens of seconds, which is unacceptable in an interactive product if you just block. Use invoke_endpoint_with_response_stream and render a "thinking" state while reasoning deltas arrive, then switch to the answer when the content channel opens:
stream = rt.invoke_endpoint_with_response_stream(
EndpointName="gpt-oss-20b-reasoning",
ContentType="application/json",
Body=json.dumps({**payload, "stream": True}),
)
for event in stream["Body"]:
chunk = event.get("PayloadPart", {}).get("Bytes", b"").decode()
# parse SSE lines; delta may carry reasoning_content or content
For batch or background work, do not use a real-time endpoint at all — an asynchronous endpoint removes the 60-second invocation ceiling and scales to zero between jobs. See choosing an inference type and our async long-running inference walkthrough.
5. Structured output, because reasoning models ramble
If a downstream system consumes the output, constrain it. vLLM in the LMI container supports guided decoding, so you can require a JSON Schema and get a parseable answer after the reasoning channel closes:
payload["response_format"] = {
"type": "json_schema",
"json_schema": {
"name": "dispatch_plan",
"schema": {
"type": "object",
"properties": {
"first_shipment": {"type": "string"},
"confidence": {"type": "number"},
},
"required": ["first_shipment", "confidence"],
"additionalProperties": False,
},
},
}
Guided decoding costs a little throughput and removes an entire class of parse-failure retries. It pairs well with an evaluation gate in CI — see automated LLM evaluation gates with fmeval — so a model swap that starts violating the schema fails the pipeline instead of production.
6. What to watch after launch
ModelLatencyp95 and the output-token histogram together. Latency regressions on reasoning models are almost always token-count regressions, not infrastructure ones. Emit tokens-per-request as a custom CloudWatch metric from the container or from your client.- Queue depth / pending requests. The honest autoscaling signal for a rolling-batch engine;
InvocationsPerInstanceunder-reacts when each invocation is 10x longer than the last. - Cost per resolved task, not cost per token. A reasoning model at
higheffort can cost 5–10x amediumrun for a few points of accuracy. Measure both and let the task decide. - Safety at the boundary. Chain of thought bypasses nothing, but it does produce more text to screen. Putting a guardrail in front of the endpoint is cheap insurance — see Bedrock ApplyGuardrail in front of a SageMaker AI endpoint.
Tear down when you are finished testing:
predictor.delete_endpoint(delete_endpoint_config=True)
predictor.delete_model()
Reasoning models are worth self-hosting when you need the trace, data residency, or a fixed cost per hour rather than per token. They are not a drop-in replacement for a chat model, and the gap shows up in token budgets and output parsing long before it shows up in benchmark scores.
Sizing a reasoning-model endpoint, or deciding whether to host one at all? Talk to us.