Gen AI Evaluation Toolkit on AWS
v2.0.0
Copyright © 2026 Amazon Web Services, Inc. and/or its affiliates. All rights reserved. Amazon’s trademarks and trade dress may not be used in connection with any product or service that is not Amazon’s, in any manner that is likely to cause confusion among customers, or in any manner that disparages or discredits Amazon. All other trademarks not owned by Amazon are the property of their respective owners, who may or may not be affiliated with, connected to, or sponsored by Amazon.
Table of contents
@awssolutions/gen-ai-evaluation-toolkit-ts-client
@awssolutions/gen-ai-evaluation-toolkit-ts-client / CreateTestCaseRequest
Interface: CreateTestCaseRequest¶
Shared core fields of a test case / test result. Composed by both the datastore TestCaseData and the evaluator ResultData so the two surfaces stay in lockstep: a test result is a test case that has been run, so it carries the same core fields plus the execution artifacts (output, appMetrics, error). Add fields here only when they belong on BOTH a test case and a test result — this coupling is deliberate and keeps datastore and evaluator integrating without field mapping.
Extended by¶
Properties¶
context?¶
optionalcontext?:DocumentType
Additional context for the test case. Can contain metadata such as flags indicating human review needed, artifacts providing evidence for agent-as-judge evaluation, or any other structured data needed for evaluation.
datasetId¶
datasetId:
string
Unique identifier of the dataset
expected?¶
optionalexpected?:DocumentType
Expected output of the Gen AI Application for the given input. Can be any structured data including simple text responses, expected tool calls, expected topics, or any combination of expected behaviors.
input?¶
optionalinput?:DocumentType
The input to the Gen AI Application. Can be any structured data including prompts, context, and parameters.
metadata?¶
optionalmetadata?:DocumentType
Optional free-form metadata set by the caller. Carried through verbatim to results and ignored by the pipeline — use it to tag test cases (source, split, tenant, run label, etc.) for grouping and debugging.