+1 (726) 227-3241

Close the Loop: Ground Truth Labeling and Human Review for LLM Outputs on SageMaker AI

Teams spend weeks on retrieval quality, quantization, and endpoint autoscaling, and then ship a generative feature with no mechanism for finding out whether its answers are any good. Automated evaluation gets you part of the way — we covered that in the fmeval walkthrough — but automated graders are calibrated against human judgements, and you cannot calibrate against judgements you never collected.

This tutorial builds the missing half of the loop: a private human labeling and review workflow on AWS, wired so that the records humans touch become (a) a regression suite for your evaluation gate and (b) a preference or instruction dataset for the next fine-tune. Everything here runs alongside an existing SageMaker AI endpoint; you do not need to change the serving path to start collecting.

The shape of the loop

inference request ─▶ endpoint ─▶ response
        │                            │
        └──── data capture ──────────┘
                     │
            triage (rules + judge)
                     │
        ┌────────────┴────────────┐
     auto-accept            human review queue
                                  │
                        reviewed records in S3
                            │            │
                    eval regression   SFT / DPO
                        suite           dataset

Three decisions determine whether this is cheap or miserable: what fraction of traffic you route to humans, how long a single review takes, and whether reviewers see enough context to be decisive. Aim for reviews that take under 45 seconds and route 1–5% of traffic plus 100% of anything that tripped a guardrail or scored badly.

Step 1: capture the raw material

If you already enabled data capture for drift monitoring, you have what you need. If not, turn it on when you create or update the endpoint config:

from sagemaker.model_monitor import DataCaptureConfig

capture = DataCaptureConfig(
    enable_capture=True,
    sampling_percentage=100,
    destination_s3_uri="s3://my-ml-bucket/capture/gen-endpoint",
    capture_options=["REQUEST", "RESPONSE"],
)

predictor.update_data_capture_config(data_capture_config=capture)

Captured records land as newline-delimited JSON under a date-partitioned prefix. For generative workloads, capture is rarely enough on its own: you also want the retrieval context, the prompt template version, and the model/adapter identifier. Emit those yourself from the application as a structured log line, keyed by a request ID that also goes into the endpoint call's InferenceId, so you can join them later.

Step 2: triage — decide what a human sees

Sending everything to humans is how HITL programs die. Build a small triage function that runs on a schedule over the previous hour of capture and writes only the interesting rows to a review manifest.

import json, boto3

s3 = boto3.client("s3")

def needs_human(rec):
    if rec.get("guardrail_action") == "INTERVENED":
        return True                      # always review blocked/edited output
    if rec.get("retrieved_docs", 0) == 0:
        return True                      # RAG answered with no context
    if rec.get("judge_score", 1.0) < 0.6:
        return True                      # weak automated grade
    if rec.get("user_feedback") == "thumbs_down":
        return True
    return rec["rand"] < 0.02            # 2% random baseline sample


def write_manifest(records, key):
    body = "\n".join(
        json.dumps({"source-ref-id": r["request_id"], "source": json.dumps(r)})
        for r in records
    )
    s3.put_object(Bucket="my-ml-bucket", Key=key, Body=body.encode())

The random baseline sample matters more than it looks. Without it, your reviewed data only ever describes failures, and you lose the ability to estimate your true success rate or to detect that a change made ordinary answers slightly worse.

Step 3: a private worker team

SageMaker Ground Truth labeling jobs need a workforce. For anything involving customer text, use a private workforce backed by Amazon Cognito — your own staff or the client's subject-matter experts — rather than a public marketplace workforce.

aws sagemaker create-workteam \
  --workteam-name domain-reviewers \
  --member-definitions 'CognitoMemberDefinition={UserPool=us-east-1_xxxxxxx,UserGroup=reviewers,ClientId=xxxxxxxxxxxx}' \
  --description "Internal SMEs reviewing generated answers"

Creating the first private workteam from the console is easier, because it will create the Cognito user pool and the worker portal URL for you; script the second and subsequent ones. Note the returned WorkteamArn and the portal sign-in URL — reviewers get an email invitation and work entirely inside that portal, which means they never need AWS console access. That property is what makes it acceptable to hand this to a client's operations team.

Step 4: a custom task template

The built-in task types cover classification, bounding boxes, NER, and text ranking. Reviewing a generated answer is usually a custom task: you want the prompt, the retrieved context, the model output, and two or three questions. Ground Truth custom templates are HTML with Liquid variables substituted from each manifest record and crowd-* web components for the widgets.

<script src="https://assets.crowd.aws/crowd-html-elements.js"></script>

<crowd-form>
  {% assign rec = task.input.source | parse_json %}

  <h3>User question</h3>
  <p>{{ rec.prompt }}</p>

  <h3>Retrieved context</h3>
  <pre style="max-height:240px;overflow:auto">{{ rec.context }}</pre>

  <h3>Model answer</h3>
  <p>{{ rec.completion }}</p>

  <crowd-radio-group>
    <crowd-radio-button name="verdict" value="correct">Correct and supported</crowd-radio-button>
    <crowd-radio-button name="verdict" value="unsupported">Plausible but not supported by context</crowd-radio-button>
    <crowd-radio-button name="verdict" value="wrong">Wrong</crowd-radio-button>
    <crowd-radio-button name="verdict" value="unsafe">Unsafe or policy violation</crowd-radio-button>
  </crowd-radio-group>

  <crowd-text-area name="corrected" rows="6"
    label="If not correct, write the answer it should have given"></crowd-text-area>

  <short-instructions>
    Judge only against the retrieved context. Outside knowledge does not count as support.
  </short-instructions>
  <full-instructions header="Review guidelines">
    <p>Mark <b>unsupported</b> when the answer may be true but the context does not say so. This is the
    most common failure and the one we most need labelled correctly.</p>
  </full-instructions>
</crowd-form>

The free-text correction field is the highest-value widget on the page. A verdict gives you an eval label; a correction gives you a training target.

Launch the job against the manifest from step 2 with create-labeling-job, pointing HumanTaskConfig at your workteam ARN, your template in S3, and the pre-/post-processing Lambda ARNs for the custom task type (AWS publishes PRE-Custom and ACS-Custom ARNs per region). Set NumberOfHumanWorkersPerDataObject to 1 for routine triage and 3 for the golden set you will use to measure reviewer agreement — annotation consolidation will merge the three into one label.

Step 5: continuous review instead of batch jobs

Labeling jobs are batch-shaped: manifest in, output manifest out. For an always-on queue, Amazon Augmented AI (A2I) is the streaming equivalent. You define a flow definition once (same workteam, same custom template), then call start-human-loop from the application whenever triage fires:

a2i = boto3.client("sagemaker-a2i-runtime")

a2i.start_human_loop(
    HumanLoopName=f"review-{request_id}",
    FlowDefinitionArn=FLOW_ARN,
    HumanLoopInput={"InputContent": json.dumps({"source": json.dumps(record)})},
)

Completed loops write JSON to the flow definition's S3 output path and emit an EventBridge event, which is the clean place to hang a Lambda that appends the result to your datasets. Use batch labeling jobs for backfills and bootstrapping, and A2I for the steady state; both read the same template, so you maintain one reviewer UI.

Step 6: turn verdicts into two datasets

Every completed review should fan out to two places.

The regression suite. Any record marked wrong, unsupported, or unsafe, plus its corrected answer, becomes a test case. Store as JSONL with prompt, reference, and a tag for the failure class, and point your fmeval or custom evaluation step at it inside the pipeline. This is the only eval set that is guaranteed to describe your real failure modes, and it grows for free once the loop is running.

The training set. Corrections are supervised fine-tuning pairs. Verdict pairs on the same prompt — the original answer marked wrong, the correction marked right — are preference pairs, ready for DPO or for reward modelling in a GRPO run. Keep provenance columns (reviewer ID, timestamp, model version, prompt template version) on every row; when a fine-tune later behaves strangely, the first question is always which slice of data it learned from.

def to_datasets(review):
    rec = json.loads(review["inputContent"])["source"]
    ans = review["humanAnswers"][0]["answerContent"]
    if ans["verdict"] == "correct":
        return {"eval": {"prompt": rec["prompt"], "reference": rec["completion"], "tag": "pass"}}
    return {
        "eval": {"prompt": rec["prompt"], "reference": ans["corrected"], "tag": ans["verdict"]},
        "sft": {"messages": [{"role": "user", "content": rec["prompt"]},
                              {"role": "assistant", "content": ans["corrected"]}]},
        "dpo": {"prompt": rec["prompt"], "chosen": ans["corrected"], "rejected": rec["completion"]},
    }

Governance and cost notes

  • Least privilege on the reviewer path. The labeling job's execution role needs read access to exactly one S3 prefix. Encrypt review buckets with a dedicated KMS key and keep them out of any general-purpose analytics lake.
  • PII. If captured traffic can contain personal data, run a redaction pass (Amazon Comprehend PII detection, or a regex pass for known identifiers) before writing the review manifest. Reviewers should see the minimum needed to judge.
  • Measure the reviewers. Re-serve 5% of items to a second reviewer and track agreement. Agreement below roughly 0.7 Cohen's kappa means your guidelines are ambiguous, not that your reviewers are bad — rewrite full-instructions and re-measure before trusting the labels.
  • Cost. With a private workforce, you pay for SageMaker labeling object charges plus your people's time; the AWS line item is almost always the small half. Routing 2% of traffic is a budget decision you can tune with one constant in the triage function.

Where to start

Do not build the streaming version first. Export one hour of captured traffic, hand-build a 200-row manifest, run a single Ground Truth labeling job with two reviewers, and read the results yourself. You will learn more about your system from those 200 rows than from another week of prompt tuning — and you will discover, before you automate anything, whether your reviewers can actually agree on what a good answer looks like.

If you want this loop stood up against an existing SageMaker AI deployment, with the triage rules, reviewer templates, and dataset plumbing tuned to your domain, get in touch — it is the kind of engagement we can scope in a week.